You Get What You Sample: Evaluating Sampling Strategies for Web Security Measurements Xuenan Zhang
Yuqing Yang
Giancarlo Pellegrino
[email protected] CISPA Helmholtz Center for Information Security Saarbrücken, Germany
[email protected] CISPA Helmholtz Center for Information Security Saarbrücken, Germany
[email protected] CISPA Helmholtz Center for Information Security Saarbrücken, Germany
arXiv:2609.11218v1 [cs.CR] 10 Sep 2026
Abstract Web measurement studies rely on domain datasets such as Tranco to quantify the prevalence and impact of security issues at scale, but exhaustively analyzing these datasets is often infeasible because of the cost of advanced analysis techniques, requiring the use of sampling. Despite its widespread use, sampling remains largely guided by convention—most commonly Top 𝑁 domain selection—rather than evidence, and its influence on the validity and generalizability of security findings has received little systematic evaluation. Consequently, it remains unclear whether common sampling strategies introduce systematic bias, distort observed vulnerability rates, or limit comparability across studies. In this work, we undertake, to the best of our knowledge, the first comprehensive investigation into how sampling methodologies affect the measurements and the conclusions. Through a comprehensive literature review and large-scale measurements of 500k Tranco and 24.8M Common Crawl hosts, we perform a comparative evaluation of datasets and sampling strategies. We show that, while Top 𝑁 sampling may be a rational strategy, the researchers have to bear in mind that Top 𝑁 does not reflect the overall distribution of the web. Instead, probability-based strategies yield stable, unbiased estimates for prevalence and many impact objectives. Hybrid sampling provides no advantages over pure probability sampling, as its deterministic prefix consistently contributes negatively to accuracy. Building on these results, we provide data-backed guidance for future studies, proposing to use an adaptive probability-based sampling strategy that remains effective even when the prevalence of the target issue is unknown. 1
CCS Concepts • Security and privacy → Web application security; • General and reference → Measurement.
Keywords Web Measurements; Large-scale Security Analysis; Sampling Strategies 1 This is an extended version of the ACM CCS paper (https://doi.org/10.1145/3830454.
3846754).
This work is licensed under a Creative Commons Attribution 4.0 International License. CCS ’26, The Hague, Netherlands. © 2026 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-2871-6/2026/11 https://doi.org/10.1145/3830454.3846754
ACM Reference Format: Xuenan Zhang, Yuqing Yang, and Giancarlo Pellegrino. 2026. You Get What You Sample: Evaluating Sampling Strategies for Web Security Measurements. In Proceedings of the 2026 ACM SIGSAC Conference on Computer and Communications Security (CCS ’26), November 15–19, 2026, The Hague, Netherlands. ACM, New York, NY, USA, 20 pages. https://doi.org/10.1145/3830454.3846754
1
Introduction
Web measurement studies play a crucial role in understanding the prevalence and impact of security issues across the Web, relying on seed datasets of millions of domains, such as Tranco [41] and CrUX [17], or public crawled web archives such as Common Crawl [3]. Analyzing these datasets in full is often infeasible due to high computational cost, as measurements may require crawling and applying costly program analysis techniques such as dynamic taint tracking [9], SMT solvers [64], or static analysis [36–38]. As a result, web measurements commonly rely on sampling, analyzing only subsets of domains to reduce evaluation cost. However, the extent to which measurement findings generalize to the broader Web depends on two closely related design choices: which population of domains is used as the measurement basis, and how domains are sampled from that population. The first choice concerns the seed dataset. Popularity-ranked lists such as Tranco [41] emphasize highly visited domains and are therefore a natural fit for impact-oriented studies, whereas broad web archives such as Common Crawl [3] aim to cover the Web more comprehensively and are better tailored with prevalence-oriented analyses. The second choice concerns the sampling strategy applied to that dataset. In current web measurement practice, the dominant approach is Top 𝑁 sampling, in which only the top 1k–10k domains are analyzed, for example in studies of domain takeover [61], deployment of web security mechanisms [38, 54, 58], or fingerprinting attacks [67, 73]. Other strategies, such as random, stratified, and hybrid sampling [6, 48], appear far less frequently. Consequently, key measurement decisions are often made by convention rather than by evidence, even though mismatches between dataset, sampling strategy, and research objective can introduce systematic bias and distort conclusions about security issues on the Web. As such, evaluating methodological choices in empirical research is crucial to distinguish real security phenomena from artifacts introduced by measurement design. Prior work has focused on multiple parts of the web measurement pipeline, including ranking construction [41, 60, 71], browser configurations when crawling [26], crawling strategies [62], and tool behavior [4], yet sampling remains comparatively underexplored, despite determining which parts of the Web are observed at all. To date, we lack evidence on
CCS ’26, November 15–19, 2026, The Hague, Netherlands.
whether common sampling methodologies systematically over- or underrepresent different parts of the Web. If they do, reported vulnerability rates may be skewed, security mechanisms may appear more—or less—deployed than they truly are, and findings across studies may not be meaningfully comparable. More fundamentally, we still lack a systematic understanding of which sampling strategies are reliable, when they fail, and how dataset choice interacts with sampling to affect security measurements. This leaves open a fundamental question: how much do sampling methodologies affect the measurements we report and the conclusions we ultimately draw? In this paper, we address this question through systematic investigations on sampling methodologies in web security measurements. First, we map the landscape of sampling methodologies used in web security measurement by conducting a comprehensive review of prior work and by incorporating established sampling techniques from statistics, resulting in eight distinct strategies. Second, we perform a large-scale, ground-truth evaluation of these strategies covering multiple classes of security issues including client-side XSS, HTTP security header misconfigurations, and TLS configuration errors, using a dataset of 500k domains drawn from the Tranco ranking [41] and 24.8M hosts from the Common Crawl archives [3]. Third, we analyze how and why sampling strategies behave as they do, decomposing hybrid approaches into their constituent components, and evaluating sampling effectiveness along multiple dimensions—prevalence, impact, error magnitude, stability, and sensitivity to the underlying distribution of vulnerabilities. Our findings reveal clear and consistent patterns. Top 𝑁 sampling, despite being the dominant strategy in prior work, deviates from the overall distribution of the web: in our impact measurements, it can over- or under-estimate the true value by up to 10.6 percentage points in datasets with smaller sizes of 10k domains, and simply sampling more domains does not remove this bias. Probability-based methods behave very differently: on the 500k-domain Tranco dataset, Random, Systematic, and Stratified sampling stay within a narrow margin of the true value, making their estimates effectively stable for the measurement goals. Hybrid strategies do not solve the Top 𝑁 sampling problem: they inherit its bias and spend much of their remaining sample budget correcting it. We further show that sampling correctly from the wrong population still yields the wrong answer. For web-wide prevalence, Tranco is the wrong dataset: Random sampling on Tranco remains up to 19.6 percentage points away from the Common Crawl estimate, with a systematic average underestimation rate of 30% even when sampling size increases. Finally, we propose a new practical sampling strategy, the Adaptive Probability Sampling: begin with a 1–2% probability-sampling pilot and expand adaptively, as most non-sparse issues stabilize by 5–10% sampling. Overall, our study highlights the limitations of the most popular Top 𝑁 sampling strategy, and provides practical guidance for researchers to select sampling strategies. Contributions. This paper makes the following contributions:
• We present, to the best of our knowledge, the first empirical study to systematically evaluate how sampling strategies affect security measurements on the Web.
Zhang et al.
• We identify eight sampling strategies (six used in prior web security research and two drawn from statistical sampling theory) and evaluate both their basic and hybrid variants. • We conduct a large-scale ground-truth measurement of 500k domains from Tranco and 24M from Common Crawl to quantify sampling error across multiple security issues, measurement objectives (prevalence and impact), and evaluation metrics. • We provide a decomposition of hybrid sampling using Shapley analysis, isolating the contributions of deterministic and probabilistic components and revealing structural biases in widely used designs. • We deliver data-backed, practical guidance for future researchers, including recommendations on which sampling strategies to use, how large samples should be, and how to make these decisions when the prevalence of the target issue is unknown. • We distill six lessons learned about sampling behavior that generalize across strategies, datasets, and measurement goals.
2
Background
Before presenting our research questions, we provide essential background, beginning with an overview of sampling strategies, followed by representative problems studied in web security measurement, and concluding with common measurement objectives.
2.1
Sampling Strategies
We can broadly distinguish three families of methods [11, 29]: 1) Probability Sampling. A set of sampling methods in which every unit in the population (e.g., a domain in a ranking such as Tranco) has a known, non-zero probability of selection prior to sampling, and units are selected through a randomized process respecting those probabilities. Common methods include Simple Random Sampling, equivalent to the Random N strategy used in web measurement research. Other examples include Systematic Sampling, where samples are selected at fixed intervals (e.g., every i-th element) after a random starting point, and Stratified Random Sampling, where the population is first partitioned into subgroups (strata) and samples are randomly drawn from each stratum proportional to its size. 2) Non-Probability Sampling. A family of methods where a sample is not drawn through a fully randomized mechanism that guarantees every unit a measurable chance of selection. One example in web measurement is Top 𝑁 selection, where only the highest-ranked 𝑁 domains are chosen deterministically, leaving all other domains with zero probability of being sampled. 3) Hybrid Sampling. Hybrid Sampling combines elements of both probability and non-probability methods, e.g., selecting a deterministic Top N prefix and then applying a randomized sampling method to the remaining population.
2.2
Commonly-studied Web Security Issues
A wide range of security issues have been measured on the web, spanning web vulnerabilities, deployment of defenses, and insecure configurations. In this section, we present three representative categories that are frequently analyzed in the wild: (i) client-side vulnerability detection, exemplified by client-side XSS; (ii) HTTP header configurations, where improper or missing headers weaken
You Get What You Sample: Evaluating Sampling Strategies for Web Security Measurements
browser defenses; and (iii) TLS configuration issues, where protocol or certificate misconfigurations undermine secure communication. Client-Side Vulnerability Detection. Modern websites depend on complex client-side logic and third-party scripts, making browserside vulnerabilities an interesting target for large-scale web measurements, such as client-side XSS [43, 47, 57, 63], DOM clobbering [38], and client-side request hijacking [36]. Client-side XSS is one of the most studied classes, and it occurs when untrusted data reaches JavaScript execution sinks (e.g., DOM APIs) without sanitization. Unlike server-side XSS, which requires injecting payloads into server contexts and raises ethical risks, client-side XSS can be detected purely via browser-side instrumentation, enabling safer, scalable measurements. HTTP Security Header Configurations. Another interesting subject for web measurements are HTTP security headers. Headers are convenient as they let websites enforce browser-side protections without modifying client code, mitigating a wide range of threats such as XSS, cookie theft, and protocol downgrades. Prior measurements focused on script and resource controls (Content-Security-Policy [15]), framing defenses (X-Frame-Opt ions) [33], MIME type security (X-Content-Type-Options) [10], secure cookies (Secure, HttpOnly, and SameSite) [40], and HTTPS enforcement (Strict-Transport-Security) [28]. TLS Misconfigurations. TLS is another core pillar of web security and is a major focus of web measurements. Prior work analyzes common deployment flaws, including invalid or expired certificates [18], domain mismatches [5], weak or deprecated protocol versions (e.g., TLS 1.0/1.1) [45], insecure cipher suites [14], and improper key exchange or handshake configurations [12].
2.3
Objectives of Measurement Studies
Web security measurements are a key method for studying security issues at Internet scale rather than through isolated case studies. Common goals include quantifying prevalence, identifying emerging threats, validating assumptions, comparing defenses, and tracking ecosystem evolution. In this paper, we discuss two common measurement objectives, i.e., prevalence and impact, while noting that empirical studies often serve both simultaneously. Rather than proposing a taxonomy, we use these objectives as a conceptual framework to analyze how sampling decisions affect the validity, implications, and generality of web measurements. Prevalence Measurement. Prevalence studies estimate how widespread a security property or vulnerability is across a typically large population of websites. This objective answers questions about adoption, exposure, or susceptibility rates, drawing conclusions about the broader web from a sampled set of domains. For example, Khodayari et al. [38] measure the prevalence of DOM clobbering vulnerabilities and the effectiveness of deployed defenses across the Tranco Top 5K. Impact Measurement. Impact studies quantify the real-world consequences of a security issue beyond its mere presence. This includes estimating affected users, data exposure, behavioral influence, or ecosystem-level effects from the popularity of the domains considered. Unlike prevalence, which measures how many sites are affected,
CCS ’26, November 15–19, 2026, The Hague, Netherlands.
impact measurements emphasize how much the issue matters, often prioritizing smaller sets of high-profile domains (e.g., top-ranked sites) where security failures have disproportionate reach and user impact.
3
Problem Statement
The overarching goal of this paper is to understand how sampling choices shape the results of web security measurements and, ultimately, the scientific conclusions drawn from them. Although sampling is a fundamental step in large-scale measurement, the assumptions behind existing strategies, and their impact on accuracy and bias, are often left implicit. Our work aims to make these assumptions explicit, quantify their effects, and provide practical guidance for researchers who must select sampling methods under real resource constraints. To structure our analysis, we articulate the following research questions: RQ1: What sampling strategies are used in practice, and what alternatives should be considered? We begin by identifying the sampling strategies that appear in the web measurement literature, characterizing both their conceptual design and their practical motivations. This includes deterministic strategies such as Top N, probability-based methods such as Random, Systematic, and Stratified sampling, and hybrid designs combining rankingbased and probabilistic components. By surveying existing practice and formalizing the space of possible approaches, we establish the landscape of strategies that must be evaluated (Section 4). RQ2: How do different sampling strategies behave when applied to real-world web security measurements? The core contribution of our study is an empirical, large-scale comparison of sampling strategies. We examine not only their accuracy but also their stability, convergence behavior, and the structural assumptions embedded in their design. This includes understanding when strategies fail, whether their errors shrink with more data, and how hybrid methods decompose into the contributions of their deterministic and probabilistic components. These analyses allow us to quantify how sampling choices influence measurement outcomes and reveal previously unexamined limitations in widely used approaches (Section 6). RQ3: How do different datasets impact sampling results when measuring security impacts or prevalence? In this study, we further examine how sampling from different datasets may affect the results of security measurements. In security measurement, the goal can be categorized into two types: measuring how many top-visited domains are affected by a security issue (impact), and measuring how frequently a security issue occurs on the Web at large (prevalence). In particular, we evaluate the deviation of sampling results by applying the same sampling strategy on different datasets, namely Tranco (ranked based on popularity) and Common Crawl (not ranked). We analyze the overall distribution of both datasets and evaluate how random sampling from these datasets may lead to over- or under-estimation, and what is the level of expected deviation. This analysis allows us to understand how the choice of dataset may impact the results of security measurements, and how sampling strategies perform on different datasets (Section 7).
CCS ’26, November 15–19, 2026, The Hague, Netherlands.
RQ4: How should researchers choose sampling strategies and sample sizes when the prevalence of the issue is unknown? In practice, researchers rarely know how common a security issue is when designing a study. They must choose a sampling method and a sampling budget under uncertainty, while avoiding both overshooting (expending excessive resources on a problem that would require less) and undershooting (collecting too little data to observe rare events). Our final research question focuses on deriving practical guidance from our results: identifying which sampling functions are robust in the absence of prior knowledge, how sample size requirements scale with prevalence, and how to design sampling plans that remain cost-effective across a wide range of possible scenarios (Section 8).
4
Literature Review for Sampling Strategies
Literature Review. To answer RQ1, we collected all papers published in six top security and measurement venues, i.e., IMC, CCS, USENIX Security, NDSS, IEEE S&P, and WWW. We selected these papers between 2020 and 2024 for two reasons. First, we consider the shifting landscape of domain popularity lists, as Tranco (since late 2019) and CrUX have become common alternatives after Alexa [7] retired in May 2022. Second, web domains emerge and may become outdated. Therefore, capturing recent works reduces the risk of domains being taken down or inaccessible. The analysis consists of two phases. First, we filter out the papers unrelated to web measurement by reading the titles and abstracts, resulting in 185 papers. Then, we analyze these papers in detail, resulting in 107 papers. These papers sample from a variety of datasets, with Tranco (61) and Alexa (41) being the most dominant datasets, followed by CrUX (9), Common Crawl (7), Cisco Umbrella (3) [19], and SecRank (2) [71]. For Quantcast [55], Curlie [22], Open PageRank [50], Majestic [44], and Cloudflare Radar [21], only one paper is associated. In addition, three hostname-enumeration datasets (CAIDA DNS Names [13], Censys CT [16], and Rapid7 FDNS [56]) each appear in one paper. Based on this analysis, Tranco is the most popular domain list, and Common Crawl is the most popular non-popularity-based domain list. As such, we use these two lists as the source of our sampling experiments. On top of that, we utilize the recent Tranco list to improve representativeness, as CrUX and Cloudflare Radar rankings have been integrated into the default Tranco list since August 1, 2023 [69]. In addition to the six strategies identified via the literature review on web security measurements, we also included sampling strategies commonly used in statistics and applied mathematics to identify possible strategies that have been overlooked by the existing web measurement literature but still bear potential to perform representative sampling that can be advocated. Our source for this part of the survey is Etikan et al. [29]2 . We thus identified two additional sampling strategies. The results of our survey are in Table 1. There are two basic types: Top N and random sampling, and 4 hybrid strategies used by related papers by combining two strategies in different ways. We observe that the Top N is mostly used in related literature, whereas hybrid strategies generally are 2 This source is the most cited paper when searching “sampling methods” and “sampling
strategy” on Google Scholar.
Zhang et al.
less popular. For comprehensiveness, we will evaluate all these 8 strategies in the rest of the paper. Sampling Parameters. For strategies requiring explicit parameters, we used configurations drawn from prior work whenever possible. While most methods only need a target sample size, stratified, bucket, and systematic sampling derive their sample size from other parameters. For stratified and bucket sampling, we adopted the setups in [27] and [32] respectively. As our survey found no prior use of systematic sampling, we set its interval heuristically as the ratio between the desired sample size and the total population size.
5 Datasets 5.1 Data Collection We collected data from two sources, the Tranco list [68] and Common Crawl [3], using a 13-server cluster with 128 physical cores and 2 TB RAM per machine. We chose these sources to support two distinct measurement goals. Tranco provides a popularity-ranked domain list, which is a common basis for impact-oriented analyses. Common Crawl aims to cover as much of the Web as possible and is therefore better suited for prevalence analyses. As a website may contain multiple pages, to determine the best configuration to cover as many pages as possible while maintaining reasonable processing performance, we performed a test experiment on a small dataset comprising 10,000 domains. Our experiment shows that depth plays an important role in coverage. When increasing the depth from 1 to 2, the coverage significantly increases, but when the depth increases to 3, the benefit is marginal. As such, we crawl each domain with a Foxhound-based crawler [9] with depth of 2 and a maximum of 250 pages per domain, with a timeout of 30 seconds. Tranco List. For impact-oriented measurements, we used the Tranco list [68] (version 8LZ3V, 24 September 2025) and selected the top 500k domains. We resolved them via Google Public DNS [31] and Cloudflare DNS [20], using batched queries to avoid stressing resolvers. This yielded 435,868 resolvable domains. We then partitioned the dataset into 50 shuffled buckets of 10,000 domains for parallel crawling across our cluster. The crawl ran from 24 September to 8 October 2025. For each visited page, we recorded HTTP headers, tainted flows relevant to client-side XSS, and TLS issues. We retained only same-origin, document-type responses with status code 200, and marked a domain as vulnerable if any such page exhibited a misconfiguration. Additionally, if there is a page with security issues, then we label the exact domain as having corresponding security issues. Overall, 194,794 domains showed at least one security issue. We also logged collection failures. The most common were DNS NXDOMAIN and server unresponsiveness. Across 10k buckets, the share of successfully acquired domains ranged from 72% to 96%. Appendix D presents our error analysis. Common Crawl. For prevalence-oriented measurements, we used Common Crawl [3], which captures web content at broad scale rather than ranking domains by popularity. We retrieved HTTP responses from the Web ARChive (WARC) files of the October 2025 snapshot (ID: CC-MAIN-2025-43) [2]. Such a snapshot contains
You Get What You Sample: Evaluating Sampling Strategies for Web Security Measurements
CCS ’26, November 15–19, 2026, The Hague, Netherlands.
Name
Description
Type
No. Papers
Top N
Top N sampling obtains a ranking list based on metrics such as popularity and samples the first 𝑁 websites to form the dataset. Among the 107 sampling papers, 87 use this strategy, making it the most common. Randomly samples multiple websites from the total dataset, regardless of popularity. Used by 7 papers. First selects the 𝑁 most popular websites, then randomly selects 𝑀 websites from the remaining population, forming a dataset of 𝑁 + 𝑀 websites. Used by 4 papers. First selects the 𝑁 most popular websites, divides the remaining websites into 𝑡 strata of varying sizes (e.g., ranks 100–1,000, 1,000–10,000), and randomly selects 𝑀 websites from each stratum, forming a dataset of 𝑁 + 𝑡𝑀 websites. Used by 7 papers. First selects the 𝑁 most popular websites, divides the remaining websites into 𝑡 equal-sized buckets, and randomly selects 𝑀 websites from each bucket, forming a dataset of 𝑁 + 𝑡𝑀 websites. It differs from stratified sampling only in that the buckets are equal-sized. Used by 1 paper. Randomly selects 𝐾 websites from the top 𝑁 websites, forming a dataset of size 𝐾. Used by 1 paper.
Basic
87
Basic
7
Hybrid
4
Hybrid
7
Hybrid
1
Hybrid
1
Selects samples at fixed intervals (incl. fractions), e.g., every 5th website or from an ordered list, usually starting from a random point. Divides the population into strata (subgroups), often based on website rankings, where lower-ranked websites form larger strata. Randomly samples within each stratum.
Basic
[29]
Basic
[29]
Random N Top N plus Random Top N plus Stratified
Top N plus Bucket
Random K from Top N Systematic sampling Stratified random sampling
Table 1: Sampling strategies from our literature review.
2.6 billion web pages from 47 million hosts, comprising 468 TB of uncompressed content. As such, we prioritize the feasibility of the processing by distributing WARC files to 50 parallel workers. Also, as each WARC file comprises varying amount of domains, a worker is configured to process at most 25 million URLs. In the end, we processed a total of 1.2B unique URLs and 24,834,442 hosts, covering approximately 53% of the reported hosts. From these records, we extracted HTTP response headers for large-scale analysis. As an archival corpus, Common Crawl supports only analyses over static artifacts. We therefore use it for header-based prevalence measurements and reserve runtime analyses, such as taint-tracked XSS detection and TLS error collection, for the Tranco crawl. Overall, 13,594,429 of the 24,834,442 (i.e. 54.7%) hosts contained at least one misconfigured header. Ethics. We followed standard practices to minimize real-world impact; details are provided in Appendix A.
5.2
Security Analyses
We next describe the security analyses applied to the collected data. 5.2.1 Client-Side XSS Analysis. Client-side XSS occurs when tainted JavaScript data flows from a client-side source (e.g., an API reading the navigation URL) to an execution sink (e.g., eval). We consider both reflected and stored variants and detect them via taint analysis. We collect dynamic data flows with Foxhound [9], a Firefoxbased browser with dynamic taint tracking, driven through Playwright [53]. Because tainted flows alone do not establish exploitability, we generate and validate attack payloads using the exploit generator of Steffens et al. [63], following prior methodology [43, 47, 63]. For reflected XSS, we inject generated payloads into URLs and test whether they execute. For stored XSS, we load the page, write the payload into browser storage, reload, and observe whether execution occurs, indicating insertion into an unsafe sink such as
innerHTML or eval without sanitization. We record all confirmed vulnerabilities together with their exploitable taint flows. 5.2.2 Security Headers Analysis. To analyze security-header configurations at scale, we instrumented our Playwright-based crawler with a network-event listener that records all request and response headers during page loads. For each domain, we consider only same-origin, document-type responses with status code 200 so that evaluations reflect the domain’s own configuration rather than that of third-party or redirected hosts. For every qualifying response, we extract security-relevant headers and check them against predefined correctness criteria. A domain is marked as misconfigured if any valid page violates these rules. We analyze six classes of security-sensitive headers: content injection protection, clickjacking protection, cookie security, HSTS, CORS, and content-type/MIME handling. Detailed definitions are given in Appendix C. 5.2.3 TLS Misconfiguration Analysis. To identify TLS misconfigurations, we use Foxhound’s error logs, which capture connection failures and certificate-validation errors during page loads. For each domain, we load its landing page in Foxhound and record any TLS errors raised during the handshake or certificate checks. We focus on six common categories: domain mismatch, untrusted certificate authority, expired certificates, revoked certificates, unsupported protocol versions, and disabled signature algorithms. If Foxhound reports any such error, we flag the domain as TLS-misconfigured. This yields a consistent, browser-validated view of TLS correctness without requiring active probing beyond a standard page load. In Table 2, we present an overview of the number of domains affected by each security issue in our two datasets. For Tranco, domains are grouped by ranking range to show how issue counts vary across popularity levels. For Common Crawl, we report the corresponding overall counts on our dataset.
CCS ’26, November 15–19, 2026, The Hague, Netherlands.
Content type Tranco Ranking
Count
%
1–10K 10,001–50K 50,001–100K 100,001–200K 200,001–350K 350,001–500K
74 201 240 424 576 680
0.74% 0.50% 0.48% 0.42% 0.38% 0.45%
C. Crawl Tot.
Clickjacking Count
Zhang et al.
CORS
% Count
1,184 11.84% 4,401 11.00% 5,420 10.84% 10,556 10.56% 13,825 9.22% 13,077 8.72%
Cookie %
96 0.96% 291 0.73% 372 0.74% 700 0.70% 749 0.50% 686 0.46%
Count
HSTS %
2,335 23.35% 9,144 22.86% 11,863 23.73% 25,768 25.77% 35,769 23.85% 31,469 20.98%
Count
Injection %
1,383 13.83% 5,624 14.06% 7,324 14.65% 15,900 15.90% 22,005 14.67% 20,529 13.69%
Count
CXSS % Count
2,849 28.49% 10,906 27.27% 13,702 27.40% 26,654 26.65% 35,009 23.34% 30,826 20.55%
4,079,855 16.43% 4,588,855 18.48% 69,546 0.28% 7,767,521 31.28% 6,710,461 27.02% 3,647,031 14.69%
TLS % Count
Total %
20 0.20% 47 0.47% 79 0.20% 263 0.66% 84 0.17% 485 0.97% 162 0.16% 759 0.76% 226 0.15% 1,372 0.91% 235 0.16% 1,494 1.00% n.a.
n.a.
n.a.
Count
%
3,926 39.26% 15,966 39.92% 20,778 41.56% 42,662 42.66% 58,568 39.05% 52,894 35.26%
n.a. 13,594,429 54.74%
Table 2: Vulnerability Metrics by Domain Ranking Range
Sampling
Stability Run-wise Majority ZCR Peak RMS Over- Under- Over- Under- Tie
Baseline = 500k Top N 1 Systematic 47 Random 53 48 Stratified
8.659 2.841 3.159 6.741
2.619 0.234 0.242 0.270
95.0% 50.6% 51.5% 49.8%
5.0% 49.4% 51.5% 48.5% 55.4% 50.2% 41.6%
42.6% 5.9% 42.6% 2.0% 48.5% 9.9%
Table 3: Stability and approximation analysis for impact estimation on the full 500k baseline.
6
Assessing Impact
We evaluate how sampling strategies estimate impact, defined here as the fraction of vulnerable domains within a chosen Tranco baseline. For each strategy, we compare the estimate obtained from the sample against the corresponding baseline value, which we refer to as the true impact. We study three questions: how the basic sampling strategies behave on the full list and on Top 𝐾 prefixes, how hybrid strategies behave, and how probability sampling changes across vulnerability classes with different prevalence. We measure accuracy as deviation from the true impact. For randomized strategies, each sampling ratio is evaluated over 50 independent runs. In the figures, we report the median together with the central 50% interval. We use this interval for readability: wider intervals substantially obscure comparisons between overlapping strategies. We do not use it to characterize worst-case behavior. Instead, worst-case spread is captured separately through the Peak metric, which reports the maximum observed deviation across runs. Implementation details for the individual samplers, including our treatment of systematic and stratified sampling, are provided in Appendix F. We measure stability by analyzing how the error evolves across sampling using: (i) the Zero Crossing Rate (ZCR), which quantifies how often a method switches between overestimation and underestimation where a zero crossing occurs when the signed error changes from positive to negative, or vice versa, and a higher ZCR indicates more frequent oscillation around the baseline; (ii) Peak deviation, which captures the worst-case drift from the baseline; and (iii) Root Mean Square (RMS) error, which indicates the typical magnitude of deviation with greater emphasis on larger errors. Runwise Over/Under: We summarize all individual estimates across all sample sizes and replicate runs and calculate the proportions above and below the baseline. Majority Over/Under/Tie: For each sample
size, we compare the numbers of overestimations and underestimations across 50 runs and then count how many sample sizes have a majority of overestimations, a majority of underestimations, or a tie.
6.1
Basic Sampling Methods
6.1.1 Full-list Baseline (500k). Figure 1 and Table 3 show a clear split between Top 𝑁 and probability-based sampling. Top 𝑁 remains far from the true impact across the sampling range, whereas Random, Systematic, and Stratified sampling stay tightly centered around it. Top N.. Top 𝑁 is persistently biased as an estimator of the fulllist baseline. Its deviation reaches nearly ±9 percentage points at small sampling ratios and remains large as the sample grows. This is reflected in its peak deviation (8.659), RMS (2.619), and near-zero crossing behavior (ZCR = 1). The direction of error is also highly asymmetric: Top 𝑁 over-estimates the true impact in 95% of runs. The apparent smoothness of its curve therefore does not indicate stability in the statistical sense; it reflects a consistent structural bias induced by the ranking prefix. Probability Sampling. Random, Systematic, and Stratified sampling behave differently. Their deviations remain small, oscillate around the true impact. Also, deviations of random and stratified sampling narrow as the sample grows. Importantly, we find that the RMS of the probabilistic sampling strategies are below 0.3, way lower compared to the Top N sampling of 2.6. The peak deviation is also lower compared to Top N sampling. Takeaway. For the full 500k baseline, the main result is not subtle: Top 𝑁 is biased, while probability-based strategies are accurate and stable. Systematic sampling provides the strongest overall performance in this setting. 6.1.2 Top 𝐾 Baselines (10k, 20k, 50k, 100k). The full 500k baseline captures impact over the entire ranked population, but many prior studies do not sample from the full ranking. Instead, they first restrict the sampling universe to a prefix and then draw samples within that prefix. For example, Lee et al. [42] evaluated their proposed approach on a 10k random sample from the Top 100K. We model this design choice by repeating the analysis on Top 10K, 20k, 50k and 100K baselines. For each prefix, we recompute the true impact within that baseline and evaluate the samplers against that value, with the minimal sample size of 5% of the target baseline sample size. We exclude Stratified sampling here, as stratification
Deviation from baseline (%)
You Get What You Sample: Evaluating Sampling Strategies for Web Security Measurements
CCS ’26, November 15–19, 2026, The Hague, Netherlands.
15.00 5.00 1.50 0.40 0.12 0.10 0.08 0.06 0.04 0.02 0.00 -0.02 -0.04 -0.06 -0.08 Top N -0.10 Systematic (median, 50% interval) -0.12 -0.40 Random (median, 50% interval) -1.50 Stratified (median, 50% interval) -5.00 -15.00 0.000 0.025 0.050 0.075 0.100 0.125 0.150 0.175 0.200 0.225 0.250 0.275 0.300 0.325 0.350 0.375 0.400 0.425 0.450 0.475 0.500
Sample size as a fraction of full website population
Figure 1: Deviation from the true impact across increasing sampling ratios for the full 500k baseline (5-window smoothing; Y axis on log scale).
Sampling ZCR
Stability Run-wise Majority Peak RMS Over- Under- Over- Under- Tie
Baseline = 10k Top N 0 10.550 3.767 461 8.317 0.724 Random Systematic 491 6.145 0.705
0 50.2 48.9
100 49.8 50.2
45.3 46.3
43.7 11.0 50.1 3.6
Baseline = 20k 9 10.635 3.269 Top N Random 463 5.233 0.514 Systematic 486 4.003 0.497
20.5 49.6 51.0
79.5 50.4 49.0
43.0 50.2
48.5 47.7
Baseline = 50k Top N 5 Random 482 Systematic 507
2.842 1.088 3.031 0.323 2.698 0.333
74.5 50.0 48.8
25.5 50.0 51.2
43.8 46.0
42.8 13.4 51.0 3.0
Baseline = 100k Top N 4 478 Random Systematic 501
3.728 1.120 2.420 0.228 2.199 0.234
10.9 49.9 49.0
89.1 50.1 51.0
43.6 45.1
46.2 10.2 50.3 4.6
8.5 2.1
Table 4: Stability and approximation analysis for impact estimation on Top 𝐾 baselines.
is designed for broader ranked populations; on narrow prefixes, proportional sampling from deeper strata would not match the measurement objective. Figure 2 and Table 4 show the same qualitative pattern as in the full-list experiment, but with higher early variance. On smaller prefixes, vulnerable domains are sparser in absolute terms, so early samples are more sensitive to whether they include one of the few positives. This increases volatility at low sampling ratios especially for Top 10K and Top 20K, and the perturbation range in deviation narrows as baseline size grows from 10K to 100K. However, it does not change the overall character of the estimators. Probability Sampling. Random and Systematic sampling remain centered around the true impact across all Top 𝐾 baselines. Their zero-crossing rates are high because the estimates oscillate frequently around the baseline, not because they drift far from it. The more informative quantities here are Peak and RMS: both methods
improve steadily as the baseline grows, with Random achieving the lower RMS in every setting. By Top 100K, both estimators are already tightly concentrated, with RMS values of 0.228 and 0.234, respectively. Top N.. Top 𝑁 again behaves as a biased estimator rather than a noisy one. Its errors remain large across all prefixes, but the sign depends on the baseline: it under-estimates for some prefixes and over-estimates for others. This shift in direction underscores the core problem with Top 𝑁 : its error is driven by the structure of the ranking prefix, not by ordinary sampling variance. Moreover, Top 𝑁 demonstrates significantly low ZCR close to 0 and largely imbalanced distribution, demonstrating an unstable performance. This stability issue has also contributed to the peak deviation being slightly below the Random sampling by 0.2%, only for the 50k baseline (2.84% vs. 3.03%). However, the RMS of top 𝑁 is thrice as big as that of Random sampling, with 74.5% run-wise over-estimation versus fully balanced performance of Random sampling, indicating that overall bias is still in effect. Takeaway. Across top 𝐾 baselines, probability sampling remains reliable even when early estimates are noisy. The extra volatility at small sampling ratios is a consequence of sparse positives, not estimator bias. Top 𝑁 , in contrast, continues to produce large structural errors as an estimator of the full-baseline.
6.2
Hybrid Sampling
Hybrid samplers combine a deterministic Top 𝑁 prefix with a probability-based sample drawn from the remainder of the ranking. The appeal is straightforward: include prominent domains by construction, then broaden coverage through randomization. We evaluate both whether this improves impact estimation and which component drives the final result. 6.2.1 Impact Estimation. Figure 3 shows that hybrid estimators inherit the bias of their Top 𝑁 prefix and spend the remainder of the sample budget correcting it. The initial jump at the Top 𝑁 cutoff reproduces the standalone behavior of Top 𝑁 : large for Top 5K, smaller for Top 10K, but non-negligible in both cases. Beyond
CCS ’26, November 15–19, 2026, The Hague, Netherlands.
Top N Systematic (median, 50% interval) Random (median, 50% interval) 0.08 0.16 0.24 0.32 0.40 0.48 0.56 0.64 0.72 0.80 0.88 0.96
20.00 10.00 5.00 2.00 1.00 0.50 0.40 0.30 0.20 0.10 0.00 -0.10 -0.20 -0.30 -0.40 -0.50 -1.00 -2.00 -5.00 -10.00 -20.00
Deviation from baseline (%)
Deviation from baseline (%)
20.00 10.00 5.00 2.00 1.00 0.50 0.40 0.30 0.20 0.10 0.00 -0.10 -0.20 -0.30 -0.40 -0.50 -1.00 -2.00 -5.00 -10.00 -20.00
Zhang et al.
Sample size as fraction of top 10,000 websites (a) Top 10K baseline
Top N Systematic (median, 50% interval) Random (median, 50% interval) 0.08 0.16 0.24 0.32 0.40 0.48 0.56 0.64 0.72 0.80 0.88 0.96
Sample size as fraction of top 50,000 websites
Sample size as fraction of top 20,000 websites
20.00 10.00 5.00 2.00 1.00 0.50 0.40 0.30 0.20 0.10 0.00 -0.10 -0.20 -0.30 -0.40 -0.50 -1.00 -2.00 -5.00 -10.00 -20.00
Deviation from baseline (%)
Deviation from baseline (%)
20.00 10.00 5.00 2.00 1.00 0.50 0.40 0.30 0.20 0.10 0.00 -0.10 -0.20 -0.30 -0.40 -0.50 -1.00 -2.00 -5.00 -10.00 -20.00
Top N Systematic (median, 50% interval) Random (median, 50% interval) 0.08 0.16 0.24 0.32 0.40 0.48 0.56 0.64 0.72 0.80 0.88 0.96 (b) Top 20K baseline
Top N Systematic (median, 50% interval) Random (median, 50% interval) 0.08 0.16 0.24 0.32 0.40 0.48 0.56 0.64 0.72 0.80 0.88 0.96
Sample size as fraction of top 100,000 websites
(c) Top 50K baseline
(d) Top 100K baseline
Figure 2: Deviation from the true impact across increasing sampling ratios for Top 𝐾 baselines (5-window smoothing; Y axis on log scale).
Deviation from baseline (%)
15.00 5.00 1.50 0.40 0.10 0.08 0.06 0.04 0.02 0.00 -0.02 -0.04 -0.06 -0.08 -0.10 -0.40 -1.50 -5.00 -15.00
Stability Run-wise Majority Sampling ZCR Peak RMS Over- Under- Over- Under- Tie
Top 10,000
Top N = 5k Random Stratified Bucket
Top 5,000
Top-only Top N plus Random (mean) Top N plus Stratified (mean) Top N plus Bucket (mean)
0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 0.40 0.45 0.50
0 2.025 0.279 2 2.132 0.278 0 2.012 0.277
26.7 27.5 26.9
73.3 72.5 73.1
0 1 0
Top N = 10k 16 0.858 0.142 Random Stratified 18 1.117 0.147 Bucket 19 0.790 0.147
57.9 58.1 57.4
42.1 41.9 42.6
78.1 78.1 76.0
100 96.9 100
0 2.1 0
10.4 11.5 14.6 7.3 17.7 6.3
Table 5: Stability and approximation analysis for hybrid Top 𝑁 sampling strategies.
Sample size as a fraction of full website population
Figure 3: Deviation from the true impact for hybrid sampling strategies across increasing sampling ratios (5-window smoothing; Y axis on log scale).
that cutoff, the hybrid curves begin to resemble the corresponding probability samplers. This correction is real but costly. For Top 10K hybrids, the probabilistic tail can eventually pull the estimate close to the true impact, as reflected in the low RMS values in Table 5. For Top 5K hybrids, the inherited bias is much larger and remains visible much longer.
In other words, hybrids do improve as the tail grows, but the early sample budget is spent offsetting the deterministic prefix rather than improving over a pure probability sample. The three tail samplers behave similarly at the aggregate level. Random, Stratified, and Bucket produce near-identical with less than 0.3% difference of peak errors for a fixed prefix, and their RMS values differ only marginally by 0.005%. However, same sampling strategy differs drastically across different Top N prefixes by up to 100% in terms of peak deviation and RMS. The dominant factor is therefore the prefix size, not the choice of tail sampler.
A (Random) B (Random)
0.5
0.5 A
1.0
Error (Random) A (Stratified) B (Stratified) Error (Stratified) A (Bucket) B (Bucket) Error (Bucket)
Error hybrid
1.5
CCS ’26, November 15–19, 2026, The Hague, Netherlands.
1.6 1.2 0.8
# Domains
Included Issues
0–1,000 2,000–3,000 4,000–5,000 40,000–50,000 70,000–80,000 110,000–120,000
Client-side XSS (CXSS) Content-Type & MIME Handling, CORS TLS misconfiguration Clickjacking (missing or misconfigured) HSTS (missing or misconfigured) Content Injection, Cookie Security
Table 6: Buckets of the robustness experiment, grouped by the number of vulnerable domains.
0.4 0.0
2.5
5.0
7.5
Rest sample size relative to Top-N (nrest / 5000)
10.0
Figure 4: Shapley values of the Top 𝑁 block (𝜙𝐴 ) and the restsampling component (𝜙 𝐵 ) as the rest-sample grows from 1k to 50k domains (0.2× to 10× the Top 𝑁 size).
Takeaway. Hybrid samplers do not escape the central weakness of Top 𝑁 . Their performance is largely determined by how much bias the deterministic prefix injects and how much probabilistic tail sampling is required to undo it. 6.2.2 Contribution Analysis via Shapley Values. The hybrid curves show that the tail corrects the prefix. We now ask which component contributes error and which contributes recovery. To answer this, we use Shapley values [70] to decompose the hybrid estimator into two components: the Top 𝑁 block (𝐴) and the rest-sampler (𝐵). We fix the prefix at Top 5K and measure its standalone error. For each rest-sampling mode (Random, Bucket, Stratified), we then draw rest-samples from 1,000 to 50,000 domains, repeating each draw 50 times. For every rest-sample size, we compute the error of the Top 𝑁 block alone (𝐸𝐴 ), the error of the rest-sample alone (𝐸𝐵 ), and the error of the combined hybrid (𝐸𝐴𝐵 ). A detailed derivation is provided in Appendix E. Figure 4 makes the division of labor explicit. Across all three hybrid strategies, the Top 𝑁 block has strictly negative Shapley values throughout. Its contribution is therefore not merely imperfect; it is systematically harmful. The deterministic prefix starts the estimator in a biased state and continues to reduce accuracy even as the rest-sample grows. The rest-sampling component behaves differently. At very small sizes, it also contributes negatively because its own variance is high. Once the tail becomes large enough, however, its Shapley value turns positive and remains so. This is the point at which the probabilistic component begins to reduce total hybrid error rather than add to it. Stratified and Bucket sampling cross this threshold slightly earlier than Random, consistent with their lower variance. Takeaway. The Shapley decomposition confirms the interpretation suggested by the hybrid curves: hybrid accuracy comes from the probabilistic tail, not from the deterministic prefix. The Top 𝑁 block contributes negative value throughout; the rest-sampler is the only component that eventually improves the estimate.
Fraction of runs within tolerance
Shapley values A, B
B
0.0
Hybrid error (Absolute Relative Error %)
You Get What You Sample: Evaluating Sampling Strategies for Web Security Measurements
1.00 0.75 0.50 Bucket 0 1K Bucket 2K 3K Bucket 4K-5K Bucket 40K-50K Bucket 70K-80K Bucket 110K-120K
0.25 0.00
0.0%
10.0%
20.0%
30.0%
40.0%
Sample size as a fraction of full website population
50.0%
Figure 5: Fraction of probability-sampling runs within a 5% relative-error tolerance of the true impact across vulnerability buckets of different sizes (11-window smoothing).
6.3
Sensitivity Across Vulnerability Buckets
Our analyses so far aggregate all security issues into a single vulnerable population. This reveals overall sampling behavior, but it does not show whether the same pattern holds for issues with very different prevalence. We therefore group vulnerabilities by the number of affected domains and re-evaluate probability sampling within each group. We restrict this analysis to Random, Systematic, and Stratified sampling. The point here is to test how estimator quality changes as prevalence changes. Since Top 𝑁 is a deterministic prefix selector rather than an unbiased estimator of the full baseline, it is not informative for this comparison. Buckets. Table 6 groups the issues in our dataset by the number of affected domains, from rare classes with only a few hundred positives to common classes affecting well over 100k domains. For each bucket, we compute how often the 150 estimates at a given sampling ratio (50 runs for each of the three probability samplers) fall within a 5% relative-error tolerance of the true impact. Takeaway. Figure 5 shows a consistent prevalence-driven pattern. For common issues, probability sampling reaches the true impact quickly and with little uncertainty. For rarer issues, convergence is slower and larger samples are required, but the curves still improve smoothly as the sample grows. The main implication
Zhang et al.
1.0 0.5
Common RMS=0.37 Rand Tranco RMS=0.33
0.2 0.1 0.0 -0.1 -0.2 -0.5 -1.0 -2.5 -5.0
Common |Peak|=4.04 Rand Tranco |Peak|=2.93
20,000
40,000
60,000
Sample size
80,000
Random Tranco (50% interval) Random Common (50% interval)
4.00
Random Tranco (50% interval) Random Common (50% interval)
5.0 2.5
2.00
Deviation from baseline (%)
Deviation from baseline (%)
CCS ’26, November 15–19, 2026, The Hague, Netherlands.
0.00 -2.00 -4.00
Common |Peak|=4.04
-13.50 -15.00 -16.50 -18.00 -19.50
Rand Tranco |Peak|=19.64
20,000
100,000
(a) Respective Baselines
40,000
60,000
Sample size
80,000
100,000
(b) Common Crawl baseline
Figure 6: Deviation of Sample Result w.r.t. Common Crawl Baseline and Self Baseline. Fig. a) shows the relative deviation with respect to the datasets’ own ground truth. Fig. b) shows the deviation with respect to Common Crawl ground truth. Stability Run-wise Majority ZCR Peak RMS Over- Under- Over- Under- Tie
Dataset Tranco Common Crawl
41 2.93 0.33 49.9% 56 4.04 0.37 49.7%
50.1% 50.3%
38% 48%
48% 14% 39% 13%
Table 7: Merged results for the prevalence analysis for Figure 6a.
Prevalence ratio (Tranco / CC)
0.76 0.74 0.72 0.70 0.68 0.66 0.64 0
Ratio Fluctuations Estimated Ratio (Tranco / CC) Ground Truth Ratio (0.69x)
20,000
40,000
60,000
Sample size
80,000
100,000
Figure 7: Prevalence estimation ratio between Common Crawl and Tranco across sample sizes
is that probability sampling remains well-behaved across vulnerability classes; what changes is not the qualitative behavior of the estimator, but the amount of data needed for it to stabilize.
7
Prevalence
Prevalence measurements aim to estimate how frequently a security issue occurs on the Web at large. This objective differs from impact estimation, where the focus is on the fraction of vulnerable domains within a selected popularity-ranked population. For prevalence, the relevant question is not how many domains in a chosen list
are vulnerable, but what fraction of domains in the broader web population is vulnerable. We therefore use Common Crawl as the reference population and define the true prevalence as the fraction of vulnerable domains in Common Crawl.
7.1
Setup and Dataset Scope
Because Common Crawl is not popularity-ranked, rank-dependent strategies such as Top 𝑁 are not applicable in this setting. For the same reason, Systematic and Stratified sampling are not meaningful here, as they require a meaningful ordering or rank structure. We therefore restrict the prevalence evaluation to Random sampling. We exclude client-side XSS and TLS misconfigurations from this analysis. Common Crawl provides a static archival snapshot of the Web, and evaluating these two classes would require revisiting hundreds of millions of pages with dynamic instrumentation or TLS handshakes, which is infeasible at this scale. Still, the remaining classes already reveal substantial differences between the two datasets.
7.2
Differences Between Tranco and Common Crawl
Before evaluating sampling behavior, we first compare the fraction of vulnerable domains in Tranco and Common Crawl. Table 2 shows that the two populations differ markedly. Common Crawl has substantially higher prevalence for Content Type issues (16.43% vs. 0.44%), Clickjacking (18.48% vs. 9.69%), Cookie security (31.28% vs. 23.27%), and HSTS (27.02% vs. 14.55%). Tranco exceeds Common Crawl only for CORS (0.58% vs. 0.28%) and Injection (23.99% vs. 14.69%). These differences matter because they show that popular domains are not representative of the broader web for prevalence estimation. The mismatch is not a uniform offset: for some issue classes the prevalence gap is large, and for others it even changes direction. As a result, sampling from Tranco may distort both the absolute prevalence estimate and the relative profile of which issues appear most widespread.
You Get What You Sample: Evaluating Sampling Strategies for Web Security Measurements
7.3
Random Sampling for Prevalence Estimation
Figure 6 evaluates Random sampling under the Common Crawl prevalence baseline. In this figure, we evaluate the performance of random sampling from both Tranco (blue regions) and Common Crawl (green regions). Similarly to previous figures, we plot the deviation from the baseline with median and the central 50% interval. The two subplots utilize different baselines. In Figure 6a, we plot the deviation for both datasets with their datasets’ own baseline to reflect the stability of the sampling strategy itself. In Figure 6b, we plot the deviation for both dataset with the baseline from Common Crawl to reflect the potential bias introduced by the dataset choice. As shown in Figure 6a, random sampling based on both datasets performs tightly centered around zero deviation with rapid convergence when sampling size increases, which indicates that the random sampling strategy remains stable and unbiased across datasets. Table 7 further shows detailed data for Figure 6a with a balanced runwise over- and under-estimation ratio, which proves the stability of random sampling across datasets. However, despite per-dataset deviation consistency, Figure 6b demonstrates drastic deviation for random sampling towards -16.5%, with a peak absolute deviation of 19.6%. This gap does not close as sample size increases, and there is definitely no zero crossing. To give a clear understanding of the ratio of underestimation, Figure 7 visualizes the ratio of the median prevalence estimates from random sampling on Common Crawl to that from Tranco across sample sizes. When sample size increases, the ratio rapidly converges to 0.70×. As such, compared with Common Crawl as a more comprehensive baseline, if a researcher utilizes Tranco as dataset for the prevalence of security issues, the result may be subject to 30% of underestimation, even if the researcher uses random sampling strategy with appropriately large sample size. All these results on the internal consistency of random sampling and drastic underestimation for mismatched dataset selection highlight the importance of dataset selection. If the dataset is not appropriate, merely optimizing sampling strategy and sizes cannot mitigate the inherited deviation of prevalence estimates.
8
Practical Guidance for Sampling
Designing a web security measurement requires three decisions: which dataset to sample from, which sampling strategy to apply, and how large a sample to collect. The dataset determines the population that a measurement can describe: a popularity-ranked list such as Tranco is natural for impact-oriented questions, whereas a broad archive such as Common Crawl better matches prevalence-oriented questions, and the strategy determines how faithfully a sample reflects that population. Our measurements show that both choices have a first-order effect on the reported results: a deterministic Top 𝑁 prefix can deviate from the rate of a larger population by several percentage points; a probability sample drawn from the wrong population remains far from the target even as the sample grows; and the sample size needed for a reliable estimate depends strongly on the prevalence of the issue. In the following, we first introduce an adaptive sampling strategy that addresses the common situation in which the distribution of the studied issue is unknown, evaluate
CCS ’26, November 15–19, 2026, The Hague, Netherlands.
it in a case study, and then summarize the resulting guidance for impact- and prevalence-oriented measurements.
8.1
Adaptive Probability Sampling
Researchers often begin a study without knowing how common a vulnerability is or where affected domains are located in the list. A fixed sampling size is difficult to choose when the prevalence of the issue is unknown. If the size fraction is too small, the sample may not contain enough cases to reach the desired precision; if it is large enough for rare issues, it wastes resources on common ones. Our sensitivity analysis shows that the fraction required to achieve a given precision differs by orders of magnitude across prevalence levels, even though probability sampling itself behaves consistently. The sampling size should therefore be determined by the data rather than fixed in advance. We therefore propose a new practical sampling strategy, adaptive probability sampling, which begins with a 1–2% probabilitysampling pilot and expands adaptively; most non-sparse issues stabilize by 5–10% sampling. In practice, the researcher can first draw a 2% pilot sample and double the cumulative sample through 4%, 8%, 16%, etc., using nested simple random sampling without replacement. At every later stage, the rule stops when the current estimate is positive and both the relative change from the preceding stage and the relative half-width of a high-confidence interval are sufficiently small; otherwise it doubles the sample, up to the complete dataset. The complete procedure and its confidence-interval details are given in Appendix G.
8.2
Case Study
We evaluate the strategy on three target sample populations that cover the two measurement objectives: for impact, we measure two baselines, the Tranco Top 100K and Tranco Top 500K; and for the prevalence, we measure the strategy on our Common Crawl dataset. For each population, we select one high-, one medium-, and one low-rate outcome—Cookie security, Clickjacking, and CORS. We set the tolerance for both stopping conditions to 5% and use a 95% confidence level for the interval. Since the adaptive strategy is stochastic, we repeat each run 1,000 times for each population and outcome to check the overall results. We then compare the adaptive sampling results to the ground truth of the complete population and evaluate if the result falls within the 5% relative tolerance as accuracy. In Table 8 we show the results of our case study. For each run, the strategy stops after sampling a certain fraction of the population; we summarize these stopping fractions across the 1,000 runs with three statistics: the median, i.e., the fraction at which half of the runs stopped (the typical cost); the mean, i.e., the average stopping fraction (the expected cost); and the maximum, i.e., the largest stopping fraction among all runs (the worst-case cost). The accuracy reports the fraction of runs whose stopped estimate is within 5% relative error of the ground truth. The results show that for common and moderately common issues, the median stopping fraction is only 4%–16% of the population, yet at least 98% of the stopped estimates are within 5% relative error of the full-population rate; in contrast, fixed 2% and 4% samples in Top 100K would have produced correct estimates
CCS ’26, November 15–19, 2026, The Hague, Netherlands.
Outcome
Median stop
Tranco Top 100K Cookie security Clickjacking CORS Tranco Top 500K Cookie security Clickjacking CORS Common Crawl Cookie security Clickjacking CORS
Zhang et al.
Mean stop
Max stop
Accuracy
8% 16% 100%
8.14% 16.37% 100.00%
16% 32% 100%
99.0% 98.3% 100.0%
4% 4% 64%
4.00% 4.10% 65.01%
4% 8% 100%
99.9% 99.2% 99.9%
4% 4% 4%
4.00% 4.00% 4.02%
4% 4% 8%
100.0% 100.0% 99.0%
Table 8: Adaptive sampling case study results over 1,000 retrospective trials.
to impact measurement, comprehensiveness matters more than popularity, and researchers should therefore choose a large dataset such as Common Crawl rather than a smaller, ranked dataset such as Tranco. Similarly, researchers can draw a probability sample from it using the adaptive sampling strategy we propose. 3) Measuring both quantities. Since our prevalence measurements in Section 7 show substantially different rates between Tranco and Common Crawl, we recommend two explicitly separated analyses: an impact-oriented measurement over a popularity-ranked list such as Tranco, and a prevalence-oriented analysis over a broader archive such as Common Crawl. As sampling strategy and research question heavily affect the conclusion, we believe that it is better to separate the dataset for specific impact or prevalence objective and synthesize the conclusion.
9 Discussion 9.1 Lessons Learned for the sparse CORS outcome in only 10.4% and 21.9% of the runs, respectively. For rare issues, the strategy does not stop prematurely: in Top 100K it continues to the complete population, and in Top 500K it uses 64% of the population, thereby making the cost of precision explicit instead of reporting an estimate with unsupported precision. The stopping decision reflects the information in the data rather than the sampling fraction alone: although CORS is also rarer in Common Crawl, a 4% sample of the dataset can also produce satisfactory results that is because the Common Crawl is a massive dataset, 4% already contains a substantial number of positive cases. The strategy therefore spends measurement budget only where the data require it, and it remains operational without prior knowledge of the issue distribution: the researcher only needs to declare the target population and the precision target.
8.3
Recommendation
The recommended practice consists of two elements, i.e., the choice of dataset and the choice of sampling strategy, and heavily relies on the research goal. In the following, we summarize our recommendations for the sampling strategy of web security measurements. 1) Measuring impact. For impact-oriented measurements, the research question determines the choice of sampling strategy. A popularity-based domain list such as Tranco is the natural dataset, because it reflects the most popular websites. If the researcher is only interested in the most popular sites, e.g., the state of the top 1,000 domains, then Top 𝑁 is the obvious choice, and the results should be explicitly reported as valid only for that prefix. In other cases, the researcher may be interested in a much larger range of domains (e.g., the top 500K domains), but cannot measure the entire population or merely wants to conduct a small-scale study. Here, the research goal is to use minimal samples to reflect the overall population of interest, so the researcher can draw a probability sample and use the adaptive sampling strategy introduced in Section 8.1 to properly reflect the distribution of the large population of interest, and if resources allow only a small sample of websites, a plain probability sample remains safer than a Top 𝑁 prefix for this case. 2) Measuring prevalence. For prevalence, the goal is to approximate the rate of a security issue across a large population. Contrary
We now distill the broader lessons from our study, focusing on the main methodological takeaways that generalize across sampling strategies, datasets, and measurement goals. 1) Sampling strategy should be chosen to match the measurement objective. The measurement objective determines both the target dataset and the sampling strategy: impact-oriented questions call for a ranked dataset such as Tranco, while prevalence-oriented questions call for a broad archive such as Common Crawl. Within the declared dataset, a complete measurement remains appropriate when the target is exactly the Top 𝑁 prefix; otherwise, a probability sample drawn from the whole dataset is recommended. We also propose an adaptive sampling strategy for the common case where the prevalence of the issue is unknown in advance. 2) Default strategies (Top 𝑁 ) may not be the best choice. Our literature review shows that Top 𝑁 sampling is by far the dominant strategy in prior web security measurements, and it is the natural choice when the measurement target is exactly the declared 𝑁 prefix. However, when it is used to estimate properties beyond that prefix, it introduces large, persistent bias, and this error does not disappear simply by increasing the sample size. 3) Probability sampling is more robust. Across all of our evaluations, probability-based methods can produce small, stable, and approximately unbiased errors. On the full Tranco Top 500K baseline, their RMS error remains below 0.3 percentage points, compared with 2.62 percentage points for Top 𝑁 . Their estimates remain centered around the baseline, their deviations are sharply bounded, and their accuracy and stability improve predictably as the sample grows. 4) Hybrid strategies do not fix Top 𝑁 prefix, but inherit its bias. Although hybrid methods appear to balance coverage of popular domains with broader representativeness, in practice they spend a substantial part of their sample budget correcting the bias introduced by their deterministic Top 𝑁 prefix. Their performance is therefore determined largely by how harmful that prefix is, rather than by any inherent advantage of the hybrid design. 5) More data cannot fix the wrong population. Sampling correctly from the wrong dataset still yields the wrong answer. In our prevalence analysis, Random sampling on Tranco is statistically wellbehaved with respect to the Tranco population, but it remains
You Get What You Sample: Evaluating Sampling Strategies for Web Security Measurements
systematically far from the Common Crawl baseline: its peak deviation from the Common Crawl baseline reaches 19.6 percentage points. Increasing sample size reduces variance, but it does not eliminate the estimation error for prevalence. 6) Dataset choice must match the measurement goal. For impactoriented studies, ranking lists such as Tranco are appropriate because they reflect the security posture of highly visited domains. For prevalence-oriented studies, datasets such as Common Crawl are more appropriate because they better approximate the Web at large. In our measurements, dataset choice changed observed prevalence by up to about 16 percentage points across reported header classes, e.g., Content-Type issues appear in 16.43% of Common Crawl hosts but only 0.44% of Tranco domains. Choosing the sampling strategy is therefore only part of the methodological decision; choosing the right source population is equally important.
9.2
Limitations
Our study evaluates sampling strategies under realistic conditions and across security issues ranging from rare to common. While this diversity reduces the risk that our findings are tied to a particular class of vulnerabilities or defenses, we cannot claim full generalizability across all security properties, datasets, or domains. Also, our results rely on a large ground-truth baseline, with each domain crawled up to 250 pages. We believe this scale is sufficient for statistically meaningful comparisons, but substantially larger or structurally different populations could still reveal edge cases not captured in our dataset. Likewise, our measurements were performed from a fixed set of servers using standard browser automation. As in prior web measurement work, factors such as network locality, transient instability, CDN variability, bot detection, JavaScript timing, DNS resolution differences, and client-specific TLS behavior may affect what content and vulnerabilities are observed. Our fixed crawling depth also under-represents deep or authenticated content. These factors limit completeness and may add noise, but they are unlikely to change the comparative behavior of the sampling strategies we evaluate. On top of that, our analysis of Common Crawl is limited to static security issues because dynamic security analysis relies on live data, whereas Common Crawl comprises static snapshots of web domains. However, if researchers need to analyze dynamic issues, selecting domains from Common Crawl and then analyzing the selected domains is an existing practice. Prior work has combined these two steps. For example, Squarcina et al. [61] used Common Crawl’s web-graph PageRank scores to select related domains and then performed live, browser-based security analyses on those domains. Future work could explore how sampling may affect dynamic security measurements by selecting domains from Common Crawl, analyzing up-to-date content, and collecting data at a smaller scale. Finally, our research mainly focuses on security issues, but our findings can also contribute to privacy impact. For the privacy issues, for example, HSTS misconfiguration enables protocol downgrade attacks that expose user browsing history to network eavesdroppers, and prior privacy work by Davitt et al. [23] studied the privacy implications of HSTS headers in Tor Browser. Similarly, TLS misconfigurations carry privacy risks beyond their security
CCS ’26, November 15–19, 2026, The Hague, Netherlands.
impact: Foppe et al. [30] have shown that certain TLS deployment flaws can leak identifiable information and enable user tracking. These examples illustrate that sampling bias may carry implications for privacy-oriented studies as well. Future work could extend the sampling analysis to privacy metrics and examine whether the biases we observe in security measurements generalizes to privacy problems.
10
Related Work
Sample representativeness in web and security studies. The representativeness of samples has been investigated by recent works. Zhang et al. [72], in their study on evaluating the accessibility of large websites, pointed out that sampling methods used in prior work may produce bias distributions of policy violations across websites. To address this issue, a new method called URLSamp is proposed, which clusters pages based on URL patterns for analysis, thereby effectively identifying accessibility issues present in webpages that originate from the same template. Tan et al. [65] build an adaptive model leveraging historical update patterns and page popularity to infer which pages are most likely to change when given a sampled webpage and its change status. Their results show that it is most likely to find more updated webpages in the current or upper directories of the changed webpages. Beyond web related studies, Redmiles et al. [59] find that security and privacy researchers often rely on data collected from Amazon Mechanical Turk (MTurk) to evaluate security tools, understand users’ privacy preferences and measure their online behavior. To check MTurk’s performance, a comparison analysis was conducted between a probabilistic telephone sample and the MTurk sample with a census-representative web-panel. Decker et al. [24] instead consider domains participating in vulnerability disclosure programs as the measurement population, and find that such sites exhibit better security practices and fewer vulnerabilities than popular domains. While the focus of that work is not on sampling strategies, both works demonstrate the impact of the dataset choice on the measurement. Web measurement experiment setups. The impact of configurations in web measurement has also been investigated. Jueckstock et al. [35] compare how different measurement tools and setups affect the results of the study. Similarly, Demir et al. [26] investigate how different browser configuration profiles affect measurement results. Some works focus on the impact of crawlers in results. Ahmad et al. [4] analyzed differences between crawlers and how crawler selection impacts understanding of the web ecosystem. Aleksei et al. [62] empirically evaluate web crawling algorithms used in measurement studies. Many studies also analyze the domain ranking lists. Pochat et al. [41] build the Tranco ranking list for research purposes, which is designed to be resistant against manipulation. Xie et al. [71] investigate the issues in popular domain ranking and build a new voting-based domain ranking list. Reproducibility of measurement studies. Approaches to improve reproducibility have been proposed. Paxson et al. [52] propose strategies for sound internet measurement based on their practical experience. Demir et al. [25] investigate the reproducibility and replicability of web measurement studies. These studies reveal that
CCS ’26, November 15–19, 2026, The Hague, Netherlands.
many studies fail to yield reproducible results, and even slightly different setups could result in different conclusions.
11
Conclusion
In this paper, we conducted the systematic analysis of sampling strategies in web security measurement research to analyze how choices of sampling strategy affect the study. We first performed a comprehensive literature review resulting in 107 papers adopting a variety of eight different sampling strategies. Then, we performed large-scale analysis to generate the ground truth of security issue distribution among two large-scale datasets. Our systematic comparison reveals that Top N sampling, despite its popularity among related works and rationality in certain research with specific target datasets, introduces significant and persistent bias when used to estimate properties beyond its declared prefix, whereas probability-based strategies remain accurate and stable. Hybrid sampling, meanwhile, yields no advantage over probability sampling. We further reveal that the choice of dataset must be aligned with the research goal: a misaligned dataset may introduce bias of up to 19.6%. Through a dedicated experiment with case studies, we provide suggestions for performing representative web measurement studies for researchers’ reference, advocating adaptive probability-based sampling.
References [1] 2018. What is max-age property in HSTS security header. https://stackoverflow. com/questions/48497736/. Accessed: 2025. [2] 2025. October 2025 Crawl Archive Now Available. https://commoncrawl.org/blog/ october-2025-crawl-archive-now-available. [3] 2026. Common Crawl web crawl data. https://commoncrawl.org/. [4] Syed Suleman Ahmad, Muhammad Daniyal Dar, Muhammad Fareed Zaffar, Narseo Vallina-Rodriguez, and Rishab Nithyanand. 2020. Apophanies or epiphanies? How crawlers impact our understanding of the web. In Proceedings of The Web Conference 2020. 271–280. [5] Devdatta Akhawe, Johanna Amann, Matthias Vallentin, and Robin Sommer. 2013. Here’s my cert, so trust me, maybe? Understanding TLS errors on the web. In Proceedings of the 22nd international conference on World Wide Web. 59–70. [6] Suood Al Roomi and Frank Li. 2023. A { Large-Scale } Measurement of Website Login Policies. In 32nd USENIX Security Symposium (USENIX Security 23). 2061– 2078. [7] Alexa. 2025. Alexa Top 1 Million. http://s3.amazonaws.com/alexa-static/top1m.csv.zip. Accessed: 2022. [8] Sajjad Arshad, Seyed Ali Mirheidari, Tobias Lauinger, Bruno Crispo, Engin Kirda, and William Robertson. 2018. Large-Scale Analysis of Style Injection by Relative Path Overwrite. In Proceedings of the 2018 World Wide Web Conference (Lyon, France) (WWW ’18). International World Wide Web Conferences Steering Committee, Republic and Canton of Geneva, CHE, 237–246. doi:10.1145/3178876.3186090 [9] Thomas Barber. 2023. Project Foxhound: Dynamic Taint Tracking for JavaScript. https://github.com/SAP/project-foxhound. Accessed: 2025-11-11. [10] Anton Barua, Hossain Shahriar, and Mohammad Zulkernine. 2011. Server side detection of content sniffing attacks. In 2011 IEEE 22nd International Symposium on Software Reliability Engineering. IEEE, 20–29. [11] Andrea E Berndt. 2020. Sampling methods. Journal of human lactation 36, 2 (2020), 224–226. [12] Karthikeyan Bhargavan, Cédric Fournet, Markulf Kohlweiss, Alfredo Pironti, and Pierre-Yves Strub. 2013. Implementing TLS with verified cryptographic security. In 2013 IEEE Symposium on Security and Privacy. IEEE, 445–459. [13] CAIDA. 2026. IPv4 Routed /24 DNS Names Dataset. https://www-old.caida.org/ data/active/ipv4_dnsnames_dataset.xml. Accessed: 2026. [14] Stefano Calzavara, Riccardo Focardi, Matus Nemec, Alvise Rabitti, and Marco Squarcina. 2019. Postcards from the post-http world: Amplification of https vulnerabilities in the web ecosystem. In 2019 IEEE Symposium on Security and Privacy (SP). IEEE, 281–298. [15] Stefano Calzavara, Alvise Rabitti, and Michele Bugliesi. 2016. Content security problems? evaluating the effectiveness of content security policy in the wild. In Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security. 1365–1375.
Zhang et al.
[16] Censys. 2026. Censys Certificates and Certificate Transparency Log Data. https: //censys.com/. Accessed: 2026. [17] Chrome Developers. 2025. Chrome UX Report. https://developer.chrome.com/ docs/crux/. Accessed: 2025. [18] Taejoong Chung, Yabing Liu, David Choffnes, Dave Levin, Bruce MacDowell Maggs, Alan Mislove, and Christo Wilson. 2016. Measuring and applying invalid SSL certificates: The silent majority. In Proceedings of the 2016 Internet Measurement Conference. 527–541. [19] Cisco Umbrella. 2026. Umbrella Popularity List. https://umbrella.cisco.com/blog/ cisco-umbrella-1-million. Accessed: 2026. [20] Cloudflare. 2025. 1.1.1.1 (DNS Resolver). https://developers.cloudflare.com/1.1.1. 1/. Accessed: 2025. [21] Cloudflare. 2026. Cloudflare Radar. https://radar.cloudflare.com/. Accessed: 2026. [22] Curlie. 2026. Curlie: The Curated Directory of the Web. https://curlie.org/. Accessed: 2026. [23] Killian Davitt, Dan Ristea, Duncan Russell, and Steven J Murdoch. 2024. Costrictor: collaborative http strict transport security in tor browser. Proceedings on Privacy Enhancing Technologies 2024, 1 (2024), 343–356. [24] Philip Decker and Florian Hantke. 2026. VDPCollect: Vulnerability Disclosure Programs as a Complement to Web Security Measurements. In Proceedings of the ACM Asia Conference on Computer and Communications Security. 1246–1260. [25] Nurullah Demir, Matteo Große-Kampmann, Tobias Urban, Christian Wressnegger, Thorsten Holz, and Norbert Pohlmann. 2022. Reproducibility and replicability of web measurement studies. In Proceedings of the ACM Web Conference 2022. 533–544. [26] Nurullah Demir, Jan Hörnemann, Matteo Große-Kampmann, Tobias Urban, Norbert Pohlmann, Thorsten Holz, and Christian Wressnegger. 2023. On the Similarity of Web Measurements Under Different Experimental Setups. In Proceedings of the 2023 ACM on Internet Measurement Conference. 356–369. [27] Nurullah Demir, Jan Hörnemann, Matteo Große-Kampmann, Tobias Urban, Norbert Pohlmann, Thorsten Holz, and Christian Wressnegger. 2023. On the Similarity of Web Measurements Under Different Experimental Setups. In Proceedings of the 2023 ACM on Internet Measurement Conference (Montreal QC, Canada) (IMC ’23). Association for Computing Machinery, New York, NY, USA, 356–369. doi:10.1145/3618257.3624795 [28] Ivan Dolnák and Ján Litvik. 2017. Introduction to HTTP security headers and implementation of HTTP strict transport security (HSTS) header for HTTPS enforcing. In 2017 15th International Conference on Emerging eLearning Technologies and Applications (ICETA). IEEE, 1–4. [29] Ilker Etikan and Kabiru Bala. 2017. Sampling and sampling methods. Biometrics & Biostatistics International Journal 5, 6 (2017), 00149. [30] Lucas Foppe, Jeremy Martin, Travis Mayberry, Erik C Rye, and Lamont Brown. 2018. Exploiting tls client authentication for widespread user tracking. Proceedings on Privacy Enhancing Technologies (2018). [31] Google. 2025. Google Public DNS. https://developers.google.com/speed/publicdns. Accessed: 2025. [32] Florian Hantke, Stefano Calzavara, Moritz Wilhelm, Alvise Rabitti, and Ben Stock. 2023. You call this archaeology? evaluating web archives for reproducible web security measurements. In Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security. 3168–3182. [33] Lin-Shung Huang, Alex Moshchuk, Helen J Wang, Stuart Schecter, and Collin Jackson. 2012. Clickjacking: Attacks and defenses. In 21st USENIX Security Symposium (USENIX Security 12). 413–428. [34] Ronaldo Iachan. 1982. Systematic sampling: a critical review. International Statistical Review/Revue Internationale de Statistique (1982), 293–303. [35] Jordan Jueckstock, Shaown Sarker, Peter Snyder, Aidan Beggs, Panagiotis Papadopoulos, Matteo Varvello, Benjamin Livshits, and Alexandros Kapravelos. 2021. Towards realistic and reproducibleweb crawl measurements. In Proceedings of the Web Conference 2021. 80–91. [36] Soheil Khodayari, Thomas Barber, and Giancarlo Pellegrino. 2024. The great request robbery: An empirical study of client-side request hijacking vulnerabilities on the web. In 2024 IEEE Symposium on Security and Privacy (SP). IEEE, 166–184. [37] Soheil Khodayari and Giancarlo Pellegrino. 2021. { JAW } : Studying client-side { CSRF } with hybrid property graphs and declarative traversals. In 30th usenix security symposium (usenix security 21). 2525–2542. [38] Soheil Khodayari and Giancarlo Pellegrino. 2023. It’s (dom) clobbering time: Attack techniques, prevalence, and defenses. In 2023 IEEE Symposium on Security and Privacy (SP). IEEE, 1041–1058. [39] Leslie Kish. 1965. Survey sampling. new york: John wesley & sons. Am Polit Sci Rev 59, 4 (1965), 1025. [40] Hyunsoo Kwon, Hyunjae Nam, Sangtae Lee, Changhee Hahn, and Junbeom Hur. 2019. (In-) security of cookies in HTTPS: Cookie theft by removing cookie flags. IEEE Transactions on Information Forensics and Security 15 (2019), 1204–1215. [41] Victor Le Pochat, Tom Van Goethem, Samaneh Tajalizadehkhoob, Maciej Korczyński, and Wouter Joosen. 2019. Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation. In Proceedings of the 26th Annual Network and Distributed System Security Symposium (NDSS 2019). doi:10.14722/ndss. 2019.23386
You Get What You Sample: Evaluating Sampling Strategies for Web Security Measurements
[42] Changmin Lee and Sooel Son. 2023. Adcpg: Classifying javascript code property graphs with explanations for ad and tracker blocking. In Proceedings of the 2023 ACM SIGSAC conference on computer and communications security. 3505–3518. [43] Sebastian Lekies, Ben Stock, and Martin Johns. 2013. 25 million flows later: large-scale detection of DOM-based XSS. In Proceedings of the 2013 ACM SIGSAC conference on Computer & communications security. 1193–1204. [44] Majestic. 2026. Majestic Million. https://majestic.com/reports/majestic-million. Accessed: 2026. [45] Salvatore Manfredi, Silvio Ranise, and Giada Sciarretta. 2019. Lost in tls? no more! assisted deployment of secure TLS configurations. In IFIP Annual Conference on Data and Applications Security and Privacy. Springer, 201–220. [46] Tim McLean. 2024. Why HSTS. https://www.chosenplaintext.ca/articles/whyhsts.html. Accessed: 2025. [47] William Melicher, Anupam Das, Mahmood Sharif, Lujo Bauer, and Limin Jia. 2018. Riding out domsday: Towards detecting and preventing dom cross-site scripting. In 2018 Network and Distributed System Security Symposium (NDSS). [48] Shaoor Munir, Sandra Siby, Umar Iqbal, Steven Englehardt, Zubair Shafiq, and Carmela Troncoso. 2023. Cookiegraph: Understanding and detecting first-party tracking cookies. In Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security. 3490–3504. [49] Mankal Narasinha Murthy. 1967. Sampling theory and methods. (1967). [50] Open PageRank. 2026. Open PageRank. https://www.openpagerank.com/. Accessed: 2026. [51] OWASP. 2024. HTTP Strict Transport Security Cheat Sheet. https: //cheatsheetseries.owasp.org/cheatsheets/HTTP_Strict_Transport_Security_ Cheat_Sheet.html. Accessed: 2025. [52] Vern Paxson. 2004. Strategies for sound internet measurement. In Proceedings of the 4th ACM SIGCOMM Conference on Internet Measurement. 263–271. [53] Playwright. 2025. Fast and relieble end-to-end testing for modern web apps. https://playwright.dev/. Accessed: 2025. [54] Tara Poteat and Frank Li. 2021. Who you gonna call? an empirical evaluation of website security. txt deployment. In Proceedings of the 21st ACM Internet Measurement Conference. 526–532. [55] Quantcast. 2022. Quantcast Top Million Sites. https://www.quantcast.com/topsites/. discontinued. [56] Rapid7. 2026. Sonar Forward DNS (FDNS) Dataset. https://opendata.rapid7.com/ sonar.fdns_v2/. Accessed: 2026. [57] Jannis Rautenstrauch, Metodi Mitkov, Thomas Helbrecht, Lorenz Hetterich, and Ben Stock. 2024. To auth or not to auth? a comparative analysis of the pre-and post-login security landscape. In 2024 IEEE Symposium on Security and Privacy (SP). IEEE, 1500–1516. [58] Jannis Rautenstrauch, Giancarlo Pellegrino, and Ben Stock. 2023. The leaky web: Automated discovery of cross-site information leaks in browsers and the web. In 2023 IEEE Symposium on Security and Privacy (SP). IEEE, 2744–2760. [59] Elissa M Redmiles, Sean Kross, and Michelle L Mazurek. 2019. How well do my results generalize? comparing security and privacy survey results from mturk, web, and telephone samples. In 2019 IEEE symposium on security and privacy (SP). IEEE, 1326–1343. [60] Kimberly Ruth, Deepak Kumar, Brandon Wang, Luke Valenta, and Zakir Durumeric. 2022. Toppling top lists: Evaluating the accuracy of popular website lists. In Proceedings of the 22nd ACM Internet Measurement Conference. 374–387. [61] Marco Squarcina, Mauro Tempesta, Lorenzo Veronese, Stefano Calzavara, and Matteo Maffei. 2021. Can I take your subdomain? exploring { Same-Site } attacks in the modern web. In 30th USENIX Security Symposium (USENIX Security 21). 2917–2934. [62] Aleksei Stafeev and Giancarlo Pellegrino. 2024. { SoK } : State of the Krawlers– Evaluating the Effectiveness of Crawling Algorithms for Web Security Measurements. In 33rd USENIX Security Symposium (USENIX Security 24). 719–737. [63] Marius Steffens, Christian Rossow, Martin Johns, and Ben Stock. 2019. Don’t Trust The Locals: Investigating the Prevalence of Persistent Client-Side Cross-Site Scripting in the Wild. (2019). [64] Marius Steffens and Ben Stock. 2020. Pmforce: Systematically analyzing postmessage handlers at scale. In Proceedings of the 2020 ACM SIGSAC Conference on Computer and Communications Security. 493–505. [65] Qingzhao Tan, Ziming Zhuang, Prasenjit Mitra, and C. Lee Giles. 2007. Designing efficient sampling techniques to detect webpage updates. In Proceedings of the 16th International Conference on World Wide Web (Banff, Alberta, Canada) (WWW ’07). Association for Computing Machinery, New York, NY, USA, 1147–1148. doi:10.1145/1242572.1242738 [66] testssl.sh. 2024. testssl.sh. https://github.com/testssl/testssl.sh. Accessed: 2025. [67] Cem Topcuoglu, Kaan Onarlioglu, Bahruz Jabiyev, and Engin Kirda. 2024. Untangle: Multi-layer web server fingerprinting. In Proceedings 2024 Network and Distributed System Security Symposium, Internet Society. [68] Tranco. 2025. Information on the Tranco list with ID 8LZ3V. https://trancolist.eu/list/8LZ3V. Accessed: 2025. [69] Tranco. 2026. Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation. https://tranco-list.eu/. Accessed: 2026.
CCS ’26, November 15–19, 2026, The Hague, Netherlands.
[70] Eyal Winter. 2002. The shapley value. Handbook of game theory with economic applications 3 (2002), 2025–2054. [71] Qinge Xie, Shujun Tang, Xiaofeng Zheng, Qingran Lin, Baojun Liu, Haixin Duan, and Frank Li. 2022. Building an Open, Robust, and Stable { Voting-Based } Domain Top List. In 31st USENIX Security Symposium (USENIX Security 22). 625–642. [72] Meng-ni Zhang, Can Wang, Jia-jun Bu, Zhi Yu, Yu Zhou, and Chun Chen. 2015. A sampling method based on URL clustering for fast web accessibility evaluation. Frontiers of Information Technology & Electronic Engineering 16, 6 (2015), 449–456. [73] Xiyuan Zhao, Xinhao Deng, Qi Li, Yunpeng Liu, Zhuotao Liu, Kun Sun, and Ke Xu. 2024. Towards fine-grained webpage fingerprinting at scale. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security. 423–436.
A
Ethical Considerations
This section provides more details about our ethical considerations of our research paper. Avoiding Overload of Web and Nameserver Infrastructure. Largescale measurements risk generating excessive traffic to hosting providers and authoritative nameservers, particularly when many domains share the same backend infrastructure. To mitigate this, we distributed the 500,000 domains into 50 shuffled buckets, reducing sequential access patterns that could concentrate load on individual servers. Each machine processed at most two buckets concurrently, and our crawler limited each domain to a maximum of 250 pages, with a maximum crawl depth of two and a 30 s timeout. These restrictions ensured that our traffic remained comparable to typical benign crawlers and avoided creating undue strain on shared infrastructures. Avoiding Overload of Public DNS Resolvers. Because DNS-based preprocessing could generate significant load on the public resolvers of Google and Cloudflare, we applied strict controls on query rate. DNS lookups were issued in batches of 100 domains, with at most 10 concurrent queries and a mandatory 2 s pause between batches to ensure a low sustained query rate. We observed no DNS lookup failures attributable to resolver rate limits or throttling, suggesting that our approach did not stress these services. Non-intrusive Security Testing. Security measurements may unintentionally trigger server-side vulnerabilities or interfere with remote systems if they involve active exploitation attempts. Our methodology strictly avoided any server-side payload injection, fuzzing, or exploitation. All vulnerability detection—such as identifying client-side XSS or header misconfigurations—was performed solely through passive analysis of fetched page content, JavaScript execution in our controlled environment, and TLS inspection. Thus, our tests did not alter or probe server-side state beyond standard web browsing behavior. Handling Sensitive Vulnerability Information. Our study identified approximately 100k domains exhibiting at least one security issue, which poses a risk of harm if sensitive data were released publicly. Due to the scale of the dataset, individually notifying all affected site operators is impractical. To prevent misuse, we publicly share only aggregated results that cannot be traced back to specific domains. Access to the full 500,000-domain dataset is granted solely through a formal request process restricted to researchers who provide a clear explanation of research goals, intended data use, and risk mitigation procedures. We have started the vulnerability-disclosure process. We first revisit affected websites to verify that each finding
CCS ’26, November 15–19, 2026, The Hague, Netherlands.
is still reproducible. For confirmed findings, we identify an official security or organizational contact and submit a concise report containing the relevant configuration evidence and remediation guidance. Risks of Publishing Methodological Insights. Beyond concerns related to data collection, one might question whether disseminating our findings on effective sampling strategies introduces risks by enabling misuse of improved measurement methodologies. Our intended audience consists of researchers seeking to design more representative and reliable large-scale security studies. The knowledge we present focuses on methodological rigor, e.g., which sampling strategies yield stable measurements, not on how to exploit systems or identify specific vulnerable services. As such, our contributions do not provide actionable information for harming users, website operators, or infrastructure providers. Instead, they support reproducibility and robustness within the research community. We therefore assess the residual risk of publishing these methodological results as negligible and outweighed by the scientific value of improved measurement practices.
B
Open Science
We provide a GitHub repository3 . This repository includes our tools for crawling and analyzing the security issues of domains in the Tranco ranking list, collecting the Common Crawl dataset, and identifying security header issues. We also release the reviewed paper list, the data and code for our case study, an example tool to help researchers apply our adaptive sampling strategy, and nonidentifying results derived from our measurements. However, due to the concerns presented in Appendix A, we provide restricted access to the complete domain dataset or the sampled domain sets for each strategy. The datasets are available only upon formal request from researchers who must describe their intended use, justify the necessity of accessing sensitive information, and outline appropriate safeguards.
C
Identification of Header Misconfiguration
Content injection protection: Content Security Policy (CSP) has been crucial in preventing attackers from inserting unexpected scripts to trusted websites. In terms of misconfiguration, we specifically check if the CSP configuration in script-src is missing or too permissive, i.e., allowing wildcard ∗, containing unsafe-inline, or unsafe-eval. We identify these attributes as security-sensitive misconfiguration because wildcard allows loading scripts from any sources, and unsafe-inline allows inline scripts which still can execute malicious code if developers not configured nonce or hash for the inline code. Similarly, the attribute unsafe-eval allows the use of eval() and similar methods, which can execute arbitrary code. On top of script-src field, stylesheet scripts may also incur code injection attacks from the attackers [8]. Hence, we also consider whether style-src field contains unsafe-inline or wildcard.
3 https://github.com/viewv/lsweb_ccs26
Zhang et al.
Category Error Type
Count
Network Timeout Network Network Network Timeout Unknown Network Network Unknown
64,447 32,135 6,040 2,993 1,316 404 340 299 271 261
Unknown Host Exceeded 30000ms Connection Refused Network Reset Binding Aborted Network Timeout SSL Unknown Error Redirect Loop Error Abort General Unknown Error
Table 9: Error Types and Their Frequencies
Clickjacking protection: For clickjacking protection, we check whether clickjacking protection is not configured (i.e., neither XFrame-Options nor CSP frame-ancestors is configured) or configured with invalid values (e.g., X-Frame-Options not set to DENY or SAMEORIGIN). Cookie security: As a fundamental building block of CSRF prevention, we check the configuration on cookies by analyzing whether the SameSite parameter or whether secure options are properly configured. For SameSite attribute, we check if it is set to None without Secure attribute, which is an invalid configuration that may lead to cookie being exploited. For secure options, we check whether any of the three attributes SameSite, Secure, and HttpOnly are missing, as they allow Cross-site requests, transmission over insecure HTTP, and cookie access via JavaScript that can be fetched via XSS attack, separately. HTTP Strict Transport Security: For the configuration on HTTP Strict Transport Security (HSTS), we check whether the configuration is missing or max age is configured too short, because missing HSTS allows protocol downgrade and results in cookie hijacking over HTTP, and short max age may not provide sufficient protection duration. While prior work suggests HSTS max-age values from 180 days up to one year [1, 46, 51, 66], we set our threshold at 180 days to focus on configurations that are unambiguously insecure, avoiding the gray area between 180 days and one year to maintain a conservative and reliable baseline. Cross-Origin Resource Sharing: Cross-Origin Resource Sharing (CORS) misconfigurations can lead to sensitive data exposure. We specifically check for the misconfiguration by identifying whether two criteria are met. First, Access-Control-Allow-Origin is set to wildcard, allowing resources to be accessed from any origin. Second, Allow-Credentials is set to true, allowing cross-origin to include credentials such as Cookies, which may lead to sensitive data exposure. Content Type and MIME Handling: As an essential part of the HTTP protocol, Content Type and MIME Handling configuration indicates the media type of resources being sent. In terms of misconfiguration, we care about whether the content-type is missing which leads to potential misinterpretation, or whether X-Content-TypeOptions is set to an invalid value (i.e., not set to nosniff) which allows potential MIME sniffing attacks.
You Get What You Sample: Evaluating Sampling Strategies for Web Security Measurements
CCS ’26, November 15–19, 2026, The Hague, Netherlands.
The combinatorial weight |𝑆 |! (|𝑁 | − |𝑆 | − 1)! |𝑁 |! ensures that each ordering of players joining a coalition receives equal probability.
Figure 8: Distribution of errors, accessibility, and vulnerability of crawled domains in 10k ranking bins.
Two-Component Hybrid Sampling as a Shapley Game. In our setting, a hybrid sampling algorithm consists of two components: • A: the Top N deterministic prefix (e.g., Top 5k), and • B: the rest-sampling mechanism (Random, Bucket, or Stratified). Thus the player set is simply: 𝑁 = {𝐴, 𝐵}. With two players, there are only two possible joining orders, and the Shapley formula reduces to: 𝜙𝐴 =
D
Breakdown of Error Types at Collection Time
During the experiment on Tranco 500k, we identified roughly 20% of the domains in the list that failed to return result during crawling. To quantify their impact, we analyzed all encountered errors and summarized the top 10 most frequent types in Table 9. Among them, network issues dominate these issues, accounting for 6 of the top 10 errors. The most common one is failure to resolve the A record of a domain name, which may occur when the website owners configure their websites incorrectly, resulting in failure to resolve the domain that prevents us from crawling the data from the website. Timeout and connection-refused errors are the next two most frequent ones, reflecting servers that either did not respond within 30 seconds or actively rejected connections. Figure 8 shows counts of valid domains and detected issues across 50 buckets of 10,000 domains ordered by Tranco rank. While error rates (solid red and orange) vary across buckets, the share of valid domains (solid blue) consistently remains between 72% and 96%. Moreover, domains with security issues (solid green) exhibit a nonuniform distribution, confirming that our dataset still meaningfully reflects real-world variation and remains appropriate for evaluating sampling-strategy behavior.
E
Shapley Value Formulation for Hybrid Sampling
This appendix provides additional detail on how we formulate the component analysis of hybrid sampling using Shapley values. The main paper presents only the conceptual intuition and highlevel interpretation; here, we include the formal definitions used to compute the marginal contribution of each component in a hybrid sampler. General Shapley Value Definition. Let 𝑁 be the set of players and 𝑣 (𝑆) the value (or payoff) associated with any coalition 𝑆 ⊆ 𝑁 . The Shapley value of player 𝑖 ∈ 𝑁 is:
𝜙𝑖 =
∑︁ 𝑆 ⊆𝑁 \{𝑖 }
|𝑆 |! (|𝑁 | − |𝑆 | − 1)! [𝑣 (𝑆 ∪ {𝑖}) − 𝑣 (𝑆)] . |𝑁 |!
1 1 𝑣 (𝐴) − 𝑣 (∅) + 𝑣 (𝐴, 𝐵) − 𝑣 (𝐵) , 2 2
1 1 𝑣 (𝐵) − 𝑣 (∅) + 𝑣 (𝐴, 𝐵) − 𝑣 (𝐴) . 2 2 As is common in error-based applications, we set 𝑣 (∅) = 0. 𝜙𝐵 =
Value Function. To quantify estimation accuracy, we define the value of any coalition as the negated expected error: 𝑣 (𝑆) = − E error(𝑆) . Higher values correspond to lower estimation error. Using this definition, the relevant coalition values become: 𝑣 (𝐴) = −𝐸𝐴 ,
𝑣 (𝐵) = −𝐸𝐵 ,
𝑣 (𝐴, 𝐵) = −𝐸𝐴𝐵 ,
where 𝐸𝐴 , 𝐸𝐵 , and 𝐸𝐴𝐵 denote the Top N-only, Rest-only, and Hybrid errors, respectively Substituting these into the two-player Shapley equations yields: 𝜙𝐴 = 12 (−𝐸𝐴 ) + 21 (−𝐸𝐴𝐵 + 𝐸𝐵 ),
𝜙 𝐵 = 21 (−𝐸𝐵 ) + 21 (−𝐸𝐴𝐵 + 𝐸𝐴 ).
A positive Shapley value indicates that the corresponding component reduces estimation error; a negative value indicates that it increases error. Experimental Procedure. For each security issue, we compute Shapley values using the following workflow: (1) Compute the ground-truth prevalence over the full dataset. (2) Fix a Top N cutoff to define component 𝐴. (3) For a range of rest-sample sizes 𝑛 rest : (a) draw rest-samples using Random, Bucket, or Stratified sampling, (b) measure rest-only error 𝐸𝐵 , (c) measure the hybrid error 𝐸𝐴𝐵 by combining Top N with the rest-sample. (4) Compute 𝜙𝐴 and 𝜙 𝐵 using the equations above. Interpretation. A positive value indicates that adding the component reduces estimation error on average, whereas a negative value means that the component increases error. Thus, 𝜙𝐴 > 0 means the Top N block contributes beneficially, while 𝜙𝐴 < 0 indicates that it introduces bias or distortion. Likewise, 𝜙 𝐵 > 0 shows that the rest-sampling mechanism improves estimation, and 𝜙 𝐵 < 0 signals that it worsens accuracy.
CCS ’26, November 15–19, 2026, The Hague, Netherlands.
Strata
If 𝐴 < 𝑛, a remainder of
Rank Range
1 2 3 4 5
[1, 5,000] [5,001, 10,000] [10,001, 50,000] [50,001, 250,000] [250,001, 500,000]
Table 10: Stratified stratum configurations based on website rankings.
F
Zhang et al.
Sampling Strategies Details
This appendix provides additional detail on how we implement the sampling strategies discussed in Section 4. The Top N strategy and Random N strategy are straightforward, so we will focus instead on the Systematic sampling and Stratified sampling strategies. Systematic sampling. Systematic sampling involves selecting samples at fixed intervals 𝑘, starting from a random starting point 𝑠 in an ordered list of 𝑁 data points. As we could not find prior work in web measurement offering a reference for the calculation and use of fixed intervals, we used the fractional sampling interval approach [34] discussed by Kish et al. [39] and Murthy et al. [49]. The fractional sampling interval first selects the element at a random position 𝑠 ∈ [1, 𝑘). Then, it uses a real fixed interval 𝑘 = 𝑁 /𝑛, where 𝑁 is the total sample set and 𝑛 is the desired sample size. The index of the 𝑖-th element is then computed as: 𝑁 idx𝑖 = 𝑠 − 1 + 𝑖 · mod 𝑁 + 1 𝑛 The floor operation ensures the index is an integer, whereas the modulo operation prevents index overflow. In our implementation, we generate a random starting position 𝑠 using a uniform distribution over the first interval. Stratified random sampling. Stratified random sampling first partitions the population into strata (subgroups) according to relevant characteristics. A sample is then drawn from each stratum in proportion to the size of the stratum in the overall population, and the selected elements are combined to form the final sample. In web measurement studies, strata are often constructed based on website rankings, where lower-ranked websites typically form larger strata. In this work, we adopt the stratification configuration proposed by Demir et al. [26]. The strata are defined in Table 10. For a given total sample size 𝑛, we allocate samples to strata using proportional allocation. Let the strata be indexed by 𝑖 ∈ {1, . . . , 5}, and let 𝑁𝑖 denote the population size of stratum 𝑖 (i.e., the number of domains within the corresponding rank interval). Í5 The total population size is 𝑁 = 𝑖=1 𝑁𝑖 . The ideal (fractional) allocation for stratum 𝑖 is 𝑁𝑖 . 𝑁 We then convert the fractional allocations into integer sample counts. First, we take the floor of each value: 𝑡𝑖 = 𝑛 ·
𝑎𝑖 = ⌊𝑡𝑖 ⌋,
𝐴=
5 ∑︁ 𝑖=1
𝑎𝑖 .
𝑟 =𝑛 −𝐴 samples must still be assigned. Following adjustment practice, we distribute these 𝑟 samples to the strata with the largest 𝑁𝑖 , since the largest strata should have more elements. After obtaining the integer allocation vector a = (𝑎 1, . . . , 𝑎 5 ), we randomly draw 𝑎𝑖 samples uniformly from each stratum 𝑖. Finally union all selected elements constitutes one stratified random sample of size 𝑛. In addition, the buckets sampling strategy can be seen as a special case of stratified random sampling, where each stratum contains an equal number of elements, we adopt the same allocation and sampling procedure as described above and we use the buckets numbers 10 proposed by Hantke et al. [32].
G
Adaptive Sampling strategy
Researchers generally do not know the affected-unit rate before conducting a measurement. A fixed sample may therefore be unnecessarily large for a common issue but insufficient for a sparse one. To make the recommendation in Section 8 actionable, this appendix specifies the adaptive probability sampling strategy and evaluates it in a case study.
G.1
The Adaptive Strategy
Let F be the declared population containing 𝑁 units. We draw one random permutation of F and inspect nested prefixes of that permutation. Consequently, sampling is uniform and without replacement, and each new stage retains all units measured at previous stages while drawing only the additional units from the remaining population. For example, we expand the sample by doubling, giving cumulative sampling fractions of 2%, 4%, 8%, 16%, 32%, 64%, and, if necessary, 100%. The 2% pilot is not allowed to stop because no preceding estimate is available for comparison. At each stage, the researcher performs the following steps. Step 1: Estimate the affected-unit rate. Let 𝑛𝑘 be the cumulative number of inspected units at stage 𝑘, and let 𝑥𝑘 be the number of affected units observed among them. The current estimate is 𝑥𝑘 . (1) 𝑛𝑘 If 𝑥𝑘 = 0, the procedure does not stop because the available sample contains insufficient information about the frequency of the issue. Instead, the cumulative sample size is doubled and the next stage is evaluated. 𝑝ˆ𝑘 =
Step 2: Quantify the uncertainty of the current estimate. If at least one affected unit has been observed, we calculate the half-width of the approximate confidence interval for a proportion. For example, with a 95% confidence level, the half-width is given by √︄ 𝑝ˆ𝑘 (1 − 𝑝ˆ𝑘 ) 𝑁 − 𝑛𝑘 ℎ𝑘 = 1.96 . (2) 𝑛𝑘 𝑁 −1 The value 1.96 corresponds to the 95% confidence level. The factor (𝑁 − 𝑛𝑘 )/(𝑁 − 1) is the finite-population correction: it accounts for sampling without replacement, so the uncertainty decreases as more of the population is inspected and becomes zero when the complete population is measured.
You Get What You Sample: Evaluating Sampling Strategies for Web Security Measurements
use 𝜏 = 0.05 in our main evaluation. A study with a more limited measurement budget may select a looser tolerance, such as 𝜏 = 0.10, whereas a study requiring greater precision may select a smaller value. A smaller 𝜏 requires both greater agreement between consecutive estimates and a narrower confidence interval, and therefore generally requires a larger sample; a larger 𝜏 permits earlier stopping but produces a less precise estimate. Researchers should predeclare the tolerance according to the precision required by their research question and report it together with the final sample size and confidence interval.
Algorithm 1: Adaptive probability sampling Input: Declared population F of size 𝑁 ; tolerance 𝜏 = 0.05 1 Draw one random permutation of F 2 𝑛 1 ← ⌈0.02𝑁 ⌉ 3 Inspect the first 𝑛 1 units and compute 𝑝ˆ1 ← 𝑥 1 /𝑛 1 4 for 𝑘 ← 2, 3, . . . do 5 𝑛𝑘 ← min(2𝑛𝑘 −1, 𝑁 ) 6 Inspect only the additional units in positions 𝑛𝑘 −1 + 1, . . . , 𝑛𝑘 7 𝑝ˆ𝑘 ← 𝑥𝑘 /𝑛𝑘 8 if 𝑥𝑘 > 0 then 9 Compute the 95% CI half-width ℎ𝑘 using Equation 2 10 𝐷𝑘 ← |𝑝ˆ𝑘 − 𝑝ˆ𝑘 −1 |/𝑝ˆ𝑘 11 𝑅𝑘 ← ℎ𝑘 /𝑝ˆ𝑘 12 if 𝐷𝑘 ≤ 𝜏 and 𝑅𝑘 ≤ 𝜏 then 13 return (𝑝ˆ𝑘 , 𝑛𝑘 ) 14 end 15 end 16 if 𝑛𝑘 = 𝑁 then 17 return (𝑝ˆ𝑘 , 𝑛𝑘 ) 18 end 19 end
G.2
Step 3: Check stability and precision. Starting from the second stage, we compare the current estimate with the estimate from the preceding stage. We calculate 𝐷𝑘 =
|𝑝ˆ𝑘 − 𝑝ˆ𝑘 −1 | , 𝑝ˆ𝑘
𝑅𝑘 =
ℎ𝑘 . 𝑝ˆ𝑘
Step 4: Stop or expand the sample. We use a user-specified relative tolerance 𝜏, for example we set to 0.05 in our evaluation. The strategy stops at stage 𝑘 only when 𝐷𝑘 ≤ 𝜏,
𝑅𝑘 ≤ 𝜏 .
Adaptive Sampling Strategy: Case Study
We evaluate whether the adaptive strategy stops at an appropriate stage without access to the population-wide rate. We use three outcomes representing high, medium, and low affected-unit rates: Cookie security, Clickjacking, and CORS. We apply the same strategy to three declared populations: the Tranco Top 100K, the Tranco Top 500K, and the Common Crawl hosts. The Top 100K experiment models a resource-constrained study whose question concerns exactly the Top 100K domains. The Common Crawl experiment models a prevalence-oriented study over a much broader host population. Table 8 in Section 8 summarizes the results; the following paragraphs provide the detailed per-population analysis. A researcher would run one adaptive sequence in a real measurement. To evaluate the reliability of that one-run strategy, we repeat the experiment over 1,000 independently seeded random permutations of each known population. The adaptive stopping function receives only the observations available at its current stage. In particular, it does not receive the full-population rate. We use the latter only after stopping to calculate the relative error
(3)
Here, 𝐷𝑘 is the relative change between two consecutive estimates. It checks whether doubling the sample materially changes the result. In contrast, 𝑅𝑘 is the relative half-width of the current 95% confidence interval. It checks whether the current estimate has sufficient statistical precision. These conditions serve different purposes: two consecutive estimates may be similar by chance even when both are imprecise, so the relative-change condition alone is insufficient, while the confidence interval describes uncertainty at the current stage but does not verify that the estimate remained stable after expanding the sample. We therefore require both conditions.
𝑥𝑘 > 0,
CCS ’26, November 15–19, 2026, The Hague, Netherlands.
(4)
In other words, the strategy stops only when doubling the sample changes the estimated rate by at most 𝜏 (5% in our evaluation) and the half-width of the current 95% confidence interval is at most 𝜏 relative to the estimate. If either condition is not satisfied, the cumulative sample size is doubled and the next stage is evaluated. Choosing the tolerance. The tolerance 𝜏 is a user-specified precision parameter rather than a fixed property of the strategy. We
𝐸=
|𝑝ˆstop − 𝑝 F | , 𝑝F
(5)
where 𝑝 F is the affected-unit rate obtained from the complete population. We define retrospective reliability as the fraction of the 1,000 trials for which 𝐸 ≤ 0.05. Top 100K results. For Cookie security, 98.3% of the trials stop at 8%, and the remaining 1.7% stop at 16%. The mean sample is 8.14% of the population, and 99.0% of the stopped estimates are within 5% relative error of the complete-population rate. Clickjacking stops at 16% in 97.7% of trials and at 32% in the remaining 2.3%, producing 98.3% retrospective reliability with a mean sample of 16.37%. The low-frequency CORS outcome behaves differently. Its fullpopulation rate is only 0.759%, and the strategy reaches the complete Top 100K in every trial. Although this setting provides no measurement-cost reduction, it prevents the study from reporting an estimate that does not satisfy the requested precision. For comparison, only 10.4% of fixed 2% samples and 21.9% of fixed 4% samples are within the same 5% relative-error tolerance, and a rule based only on cross-stage change reaches 46.5%. The adaptive result therefore communicates an actionable limitation: accurately estimating this sparse outcome within the small Top 100K population requires measuring substantially more than the initially planned budget.
CCS ’26, November 15–19, 2026, The Hague, Netherlands.
Top 500K results. For the larger Tranco population, Cookie security stops at 4% in every trial, and Clickjacking stops at 4% in 97.6% of trials. Their mean sampling fractions are 4.00% and 4.10%, with respective reliabilities of 99.9% and 99.2%. Sparse CORS requires considerably more data: 97.2% of trials stop at 64%, and the mean sampling fraction is 65.01%. This confirms that a common issue can be estimated economically, whereas a strict relative-precision target for a rare issue can require most of the declared population. Common Crawl results. On the Common Crawl dataset, Cookie security and Clickjacking stop at 4% in every trial. CORS stops at 4% in 99.4% of trials and at 8% in 0.6%. The corresponding reliabilities are 100.0%, 100.0%, and 99.0%. Although the CORS rate is lower in Common Crawl than in either Tranco population, a 4% Common Crawl sample contains many more positive observations in absolute terms because the population is much larger. The stopping behavior therefore depends jointly on the affected-unit rate, population size, and requested precision, rather than on the percentage alone.
Zhang et al.