Conceptio › Archive › arXiv CS
arXiv CSopen access

Understanding the Privacy-Preserving Potential of HTTP/2 Against Webpage Fingerprinting

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

Understanding the Privacy-Preserving Potential of HTTP/2 Against Webpage Fingerprinting

arXiv:2609.05119v1 [cs.CR] 4 Sep 2026

Full version of the paper accepted at ACM CCS 2026; includes extended appendices. Please cite the conference version.

Bogdan Cebere

Prateek Kumar

Sylvain Chatel

[email protected] CISPA Helmholtz Center for Information Security Dortmund, Germany

[email protected] CISPA Helmholtz Center for Information Security Saarbrücken, Germany

[email protected] CISPA Helmholtz Center for Information Security Saarbrücken, Germany

Wouter Lueks

Christian Rossow

[email protected] CISPA Helmholtz Center for Information Security Saarbrücken, Germany

[email protected] CISPA Helmholtz Center for Information Security Dortmund, Germany

Abstract Website fingerprinting (WF) attacks can infer which webpage a user visits from encrypted HTTPS traffic alone, compromising privacy even without decryption. WF defenses commonly shape traffic through noise, padding, delays, or flow splitting — yet they are most often studied from the perspective of encapsulating protocols like Tor or VPN rather than at the application layer (HTTP). In this work, we focus on application-layer defenses enabled by the most widely deployed version of HTTP — HTTP/2. We demonstrate how known defenses can be emulated through HTTP/2 features at the client side (HTTPOS, LLaMA, FRONT, Tamaraw) and the server side (ALPaCA, Tamaraw). We further show that HTTP/2 features — such as proactive resource suggestion, multiplexing, and flow control — offer untapped potential for lightweight yet effective defenses deployable at both endpoints. We evaluate these defenses using a unified blueprint that calibrates defense parameters per dataset, then combines practical attacks, information-theoretic leakage estimates, and overhead measurements. For each defense, this framework identifies the strongest hyperparameter-tuned fingerprinting model and estimates the residual uncertainty induced by the defense using two information-theoretic leakage estimators — all while accounting for the defense’s privacy–overhead trade-offs.

CCS Concepts • Security and privacy → Web protocol security; • Computing methodologies → Machine learning.

Keywords Website Fingerprinting, Machine Learning, HTTP/2

1

Introduction

Website fingerprinting is a traffic analysis technique that enables adversaries to infer users’ online activities — such as political or social behavior — and to facilitate adversarial behavior, including censorship and targeted surveillance. Website fingerprinting remains feasible despite the widespread deployment of HTTP over

TLS (i.e., HTTPS), since TLS encrypts only the application payload, leaving network- and transport-layer headers and traffic characteristics (packet sizes, timings, and ordering) observable to on-path adversaries. Website fingerprinting leverages this residual metadata to reconstruct browsing patterns and compromise user privacy. In response, fingerprinting defenses aim to make the IP/TCP/UDP metadata uninformative. Popular obfuscation techniques rely on the Tor and VPN protocols, which encapsulate traffic and obscure its content from on-path observers. However, most web traffic does not use Tor or VPN, leaving HTTP packet metadata directly exposed to on-path adversaries such as ISPs or network operators. For this reason, we focus on fingerprinting scenarios that can be defended by clients or servers at the application-layer (HTTP). Application-layer fingerprinting defenses generally pursue one of two strategies: (1) uniformity, wherein the defense transforms webpage metadata into fixed packet lengths and rates; or (2) unpredictability, wherein the defense ensures that metadata differs across subsequent page loads. To follow these strategies, fingerprinting defenses commonly rely on four obfuscation functionalities: (1) noise traffic insertion, (2) packet padding, (3) packet delay, and (4) packet splitting [63]. These mechanisms do not hide IP-layer metadata or the destination domain and therefore complement (not replace) Tor or VPN solutions when network-layer anonymity is required. Instead, they target subpage fingerprinting over ordinary HTTPS by perturbing transport-visible features such as packet sizes, timing, and ordering, thereby reducing leakage about the specific content accessed within a known domain. This distinction is relevant when the domain itself is not sensitive, but the visited subpage is, such as a particular product or medical page. The privacy potential of modern HTTP-layer features like multiplexing, reprioritization, and flow control has not been fully explored yet. Prior work has shown that established fingerprinting defenses can be emulated in HTTP/3 [63], but no equivalent study exists for HTTP/2. This gap is practically important: as of 2026, HTTP/2 remains the predominant deployed HTTP version, accounting for approximately 60% of browser traffic, compared with roughly 31% for HTTP/3 [4, 13]. Thus, HTTP/2 is not merely a legacy target, but the protocol on which the majority of today’s Web traffic still depends.

Bogdan Cebere, Prateek Kumar, Sylvain Chatel, Wouter Lueks, and Christian Rossow

2

Moreover, the HTTP/2 defenses are more challenging to design, as TCP’s flow control is typically handled by the OS kernel, whereas QUIC (HTTP/3) implementations commonly run in userspace, allowing applications more direct control over flow-control–driven traffic shaping. Beyond emulating existing defenses, HTTP/2 also opens the door to new strategies — such as leveraging proactive resource hints or stream multiplexing — but this potential likewise remains largely unexplored. Before such defenses can be meaningfully assessed, we must first confront a fundamental challenge: how to measure whether they truly provide privacy. Evaluating defenses is inherently harder than evaluating attacks. While attacks can often succeed with a single model and dataset, defenses must be tested against multiple adversaries to establish credible security guarantees. Yet the common practice of reporting reduced fingerprinting accuracy under defense can be misleading, as degradation may also result from preprocessing errors, poor hyperparameter tuning, or near-misses by attackers. Prior work has proposed complementary techniques, including privacy-overhead calibration curves [7] and model-agnostic estimators of information leakage [10, 31, 67]. Building on these ideas, we propose a unified framework that calibrates each defense per dataset to select a practical privacy–overhead operating point, then stress-tests the selected configuration against hyperparametertuned attackers and leakage estimators. Contributions. This paper explores how HTTP/2 can be harnessed for website fingerprinting defenses. Our contributions are threefold: (1) Evaluation framework. We introduce a benchmarking blueprint for assessing the quality of fingerprinting defenses that unifies practical attacks, information-theoretic leakage estimates, and overhead measurement (Section 4). For each defense, the framework calibrates its operating parameters per dataset, evaluates its privacy–overhead trade-off, identifies the strongest hyperparameter-tuned attack, and measures the residual leakage that remains. (2) Defense emulation. We demonstrate how established defenses — both client-side [7, 11, 23, 38] and server-side [7, 11] — can be reproduced using HTTP/2 primitives, and evaluate their effectiveness on real-world case studies using our framework (Section 5). The results show substantial but dataset-dependent privacy gains: even the strongest attackers cannot narrow the set of plausible webpages to fewer than 26 candidates on four of five datasets with client-side defenses, or fewer than 63 candidates on any dataset with appropriately placed server-side defenses, whether at the 1st -party server, CDN, or across all involved servers. (3) New opportunities. We explore untapped HTTP/2 features for privacy, highlighting novel defense strategies such as clientside multiplexing and server-side proactive resource suggestion (Section 6). These client-side techniques provide more consistent protection across datasets, keeping the attacker’s estimated candidate set above seven webpages in all five cases, while reducing latency overhead by 55–98% and downstream overhead by 20–66%. On the server side, HTTP/2 enables the 1st party server to deploy defenses without controlling all contentdelivery servers, keeping the plausible candidate set above 13 webpages across all datasets while reducing bandwidth overhead by 54–80% on most datasets.

Threat Model and Attacker Goals

We consider a passive on-path adversary with access to the TCP layer, who attempts to determine the visited webpage from an encrypted HTTP/2 traffic trace by analyzing traffic characteristics such as packet lengths, inter-arrival times, and ordering. We assume the adversary can observe individual TCP connections during a page load (one per subdomain involved in the page load), grouping them into a single trace while keeping each connection identifiable. We assume the adversary cannot alter the communication – i.e., they do not actively intercept the communication using Machine-inthe-Middle attacks. This scenario is common (e.g., Internet or VPN providers observing communication) and more stealthy — no need to install custom certificates on the client. However, we also assume a defense-aware adversary who knows which defense mechanism is deployed and can collect labeled defended traces for training. We focus on subpage fingerprinting: identifying which specific subpage of a known domain a user visits — for example, the exact Amazon product page, the particular BBC article, or the specific Reddit thread. In this setting, all contacted subdomains during webpage load are already observable via IP addresses, DNS resolves, and TLS handshakes. In principle, the sequence of subpage-specific DNS lookups and TLS SNI values for third-party services (e.g., CDNs) could additionally leak subpage identity, but such leakage is orthogonally mitigated by DNS-over-HTTPS (DoH) [26] or Encrypted Client Hello (ECH) [52]. Specific to our empirical studies, assuming legacy clients without these features, DNS contact ordering alone yields near-random subpage classification on four of five datasets (F1 ≤ 0.10), as we show empirically in Section 3.2.3. The HTTP layer, therefore, remains the primary actionable source of leakage in our threat model and the only one that application-layer HTTP/2 defenses can address. Note that DoH and ECH can mitigate DNS/SNI leakage, but do not mitigate subpage fingerprinting as studied here.

3

Website Fingerprinting

To motivate the need for HTTP/2 privacy techniques, we first introduce established website fingerprinting techniques (Section 3.1), then apply them on real-world sensitive case studies (Section 3.2).

3.1

Fingerprinting Techniques

In website fingerprinting, an observer attempts to extract sensitive information from encrypted communication by using networklayer information (e.g., IP addresses), transport-layer metadata (e.g., TCP segment sizes, connection counts), and derived traffic statistics (e.g., throughput, inter-packet timing, burst patterns) [8, 9, 19, 24, 25, 30, 46, 58, 64]. This metadata is used to find the most similar page in an already-collected training set using MachineLearning methods. Previous website fingerprinting works explored the potential of Gaussian Distributions [41], linear models [14], random forests [35], clustering methods [24, 59], and neural networks (NN) [5, 16–18, 34, 36, 57, 60, 62, 66, 73]. Most prior works focus on domain fingerprinting (using as labels the hosting domain, e.g., google.com or fb.com) under Tor encapsulation; some also explore the subpage-fingerprinting problem [60, 74, 75]. In contrast to these works on subpage WF, our work focuses on subpage-WF in HTTP/2 (previously discussed only in [74] on multiplexing effect). 2

Understanding the Privacy-Preserving Potential of HTTP/2 Against Webpage Fingerprinting

3.1.1 Fingerprinting Models. For the attacker effectiveness benchmarks, we evaluate each dataset against five fingerprinting models: (1) k-FP [24], which combines 𝑘-nearest neighbor clustering with a Random Forest; (2) Deep Fingerprinting (DF) [62], a CNN-based architecture that operates directly on raw traces; (3) VarCNN [5], a ResNet-style architecture designed to capture multi-scale burst patterns; (4) Holmes [18], which combines spatial–temporal pattern analysis with contrastive learning to identify websites before the page fully loads; (5) Robust-Fingerprinting (RobustFP-CNN) [58], for which we reuse only the CNN classifier architecture, applying 2D convolutions over local packet windows to jointly capture packet-size and timing patterns in defended traces. We do not use RF’s time-aggregated representation, which improves robustness to padding but would discard the informative variable packet lengths present in our HTTP/2 traces; consequently, RobustFP-CNN does not evaluate the full RF approach. Note: All WF models above were originally developed for Tor traffic; in contrast, in our benchmarks, the packet lengths are also informative (compared to constant-size Tor cells), and omitting them would artificially cripple the models. Therefore, the benchmarks are executed on different dataset types and features than those in the related work. Implementation details in Section E.

timing statistics. Each trace is first converted to an array of statistics, and the final dataset is a matrix of shape (n_traces, LIMconns ∗ n_features). In our datasets, the bulk of the leakage is concentrated in at most three connections, so we set LIMconns = 3. The neural network models operate on a raw 3D time-series tensor with two channels: signed packet lengths (positive for upload, negative for download) and inter-arrival times. This differs from the original NN model designs [5, 18, 62], which were built for Tor traffic where cell sizes are fixed and only direction, counts, and inter-arrival times are informative. We therefore include packet lengths as a dedicated channel in the tensor representation, acknowledging that this is not a direct architectural comparison but rather an adaptation to the richer information available at the TCP layer. Within each trace, connection blocks are concatenated in arrival order, ordered by the timestamp of each connection’s first packet. Inter-arrival times are recorded relative to the start of each connection, so the timestamp resets to zero at every connection boundary; this acts as an implicit connection-boundary signal that the model can learn from. Following prior work [5, 51, 53, 62], we treat the maximum trace length 𝐿 as a hyperparameter and set 𝐿 = 5000 packets, which does not truncate any trace in our benchmarks. Each model outputs a probability distribution over the 𝑁 webpages in the training set.

3.1.2 Dataset Creation. Throughout the empirical studies, we evaluate HTTP/2 defenses from both client and server perspectives, requiring control over both endpoints while replaying the same real webpage resources. For each webpage, we first collect and cache browser requests and responses, including headers and body content (Section D.1), then replay them through a custom Python client-server with defenses enabled or disabled, (Section D.2) and capture the resulting PCAP traces (Section D.3). This replay setup is necessary to evaluate server-side defenses under identical content, enable same-resource comparisons, and isolate HTTP/2-layer effects from content and network-path variability. The resulting measurements therefore characterize defense-induced changes rather than live-Web variability: browser/CDN scheduling and resource dependencies are not reproduced, which may affect latency and intra-page variance but does not undermine the controlled comparison targeted by our study. Each PCAP trace is represented as a collection of connections, where each connection is a sequence of packet lengths, inter-arrival timings, and directions. For every webpage and defense scenario, we generate 500 defended PCAP traces, independently resampling all randomized defense choices for each replay. Under five-fold cross-validation, each attacker is therefore trained on approximately 400 defended realizations per webpage. This provides hundreds of directly measured realizations of the defense-induced distribution for every class, so additional synthetic augmentation is not needed to compensate for limited training diversity. Moreover, artificially perturbing packet sizes or timings could introduce traces that deviate from the traffic distribution produced by the evaluated defense. The traces are converted into two distinct evaluation datasets: a 2D representation (for k-FP), and a 3D representation (for DF, VarCNN, Holmes, and RobustFP-CNN models). k-FP operates on handcrafted per-connection statistics: packet counts and unique sizes per direction; top-𝑁 burst size statistics; cumulative incoming/outgoing byte totals interpolated to a fixed-length vector; and inter-arrival

3.1.3 Fingerprinting Benchmarks. The primary objective of the fingerprinting benchmarks is to assess the effectiveness of the attacker, represented by the fingerprinting model. Given 𝐶 website classes, we evaluate each model using cross-validation 𝐹 1𝑐 = Í 2TP𝑐 /(2TP𝑐 + FP𝑐 + FN𝑐 ) and report Macro-F1 = 𝐶 −1 𝐶𝑐=1 𝐹 1𝑐 , where TP𝑐 , FP𝑐 , and FN𝑐 denote the true positives, false positives, and false negatives for website 𝑐, respectively. For each benchmark, we report the mean score and the 95% confidence interval (mean ±CI95% ) across all folds. As a rule of thumb, when using 𝐶 = 100 classes, the macro-F1 value is ∼ 0.01 for random guesses, ∼ 0.1 − 0.5 for weak attackers, and 0.8+ for strong attackers. 3.1.4 Datasets Types. We investigate website fingerprinting against users who browse without additional encapsulation layers such as Tor or VPN. Most website fingerprinting research focuses on domain-level inference — determining which domain a user visits — using datasets such as Tranco or Alexa [12, 40, 41, 46, 48]. To avoid trivial leakage via TLS certificates or IP addresses, these studies typically assume clients use anonymity systems like Tor or VPN, which conceal both endpoints and traffic features. In contrast, subpage fingerprinting seeks to identify which specific subpage of a known domain a user visits (e.g., a product page on Amazon). This setting is more fine-grained, as subpages usually share similar DNS, TLS, and IP traits, leaving most differences at the HTTP layer. These architectural differences also influence how defenses behave. In Tor or VPN, a passive observer sees a single long-lived tunnel that encapsulates all TLS connections. Thus, a defense applied to any one TLS connection modifies the metadata of the entire tunnel — for example, protected traffic from one connection may overlap with the TLS handshake of another. In contrast, with plain HTTPS (without Tor or VPN), each connection is exposed individually — therefore, it must be protected independently. Since our work targets plain HTTPS traffic, we focus on subpage-fingerprinting case studies to evaluate application-layer defenses. 3

Bogdan Cebere, Prateek Kumar, Sylvain Chatel, Wouter Lueks, and Christian Rossow

Table 1: Real-world datasets and their average number of unique servers, requests, the image count, and their sizes.

1.0

k-FP

DF

VarCNN

Holmes

RobustFP-CNN

# Servers Used per Page

# Avg. Req. per Conn.

# Avg. Images per Page

Avg. Image Size (KB)

Amazon BBC Reddit Udemy Wiki

4.14 ± 0.1 8.03 ± 0.1 8.21 ± 0.1 2.03 ± 0.1 3.83 ± 0.1

19.3 ± 1.2 6.97 ± 0.3 6.87 ± 0.2 5.72 ± 0.1 8.14 ± 0.3

29.22 ± 1.5 12.8 ± 16.3 10.66 ± 0.2 8.75 ± 1.3 15.42 ± 0.7

101.8 ± 17.2 197.04 ± 173.4 84.71 ± 20.5 16.38 ± 4.4 32.16 ± 1.4

3.2

Macro-F1

0.9 Dataset

0.8 0.7 0.6

Amazon

BBC

Reddit

Dataset

Udemy

Wikipedia

Figure 1: Fingerprinting performance (F1 score) on the undefended real-world datasets.

Fingerprinting Case Studies

We now discuss five fingerprinting case studies and their privacy risks. These results will lay the foundation for the defensive benchmarks in the following sections.

transfer, such as how upload or download bytes accumulate over time). Manual inspection shows per-site fingerprints resulted from: • Amazon: number, size, and IATs of image requests to third-party CDNs (media-amazon, ssl-images-amazon). • BBC: IAT, burst, and CUMUL across ≥ 2 connections — first-party plus a CDN (static.files.bbci). • Reddit: burst and CUMUL on first-party and a CDN (redditmedia). • Udemy: burst and CUMUL, almost entirely on the first-party connection (www.udemy.com). • Wikipedia: first-party (wikipedia.org) dominates — packet counts, bursts, and CUMUL from image downloads.

3.2.1 Datasets. We collect 100 subpages from five public websites — Amazon, BBC, Reddit, Udemy, and Wikipedia — as balanced datasets with 500 samples per page — using the steps detailed in Section 3.1.2. These five websites are case studies rather than a representative sample of the Internet, chosen to capture variation in sensitive connections, requests per connection, and image counts. This diversity is central to our analysis: both the calibrated defense parameters and the strongest hyperparameter-tuned attackers vary across datasets (Section 5.3), showing that defense effectiveness is structure-dependent and that no single configuration is universally representative. We therefore evaluate multiple websites to capture dataset-dependent defense behavior rather than draw general conclusions from a single dataset. To assess the effectiveness of fingerprinting defenses, we focus on closed-world scenarios, where the attacker has full knowledge of all target pages. This setting is intentionally chosen because it represents the most challenging case to defend, and an upper bound on attacker accuracy. Table 1 summarizes some of the per-webpage statistics of each dataset: the average number of servers contacted per webpage, the average request count per connection, and the average image size observed in each webpage for each dataset. Per page load, each connection is a sequence of requests (e.g., images, script downloads) and can create a fingerprinting side-channel (alone or combined with other connections). Due to larger resource-specific sizes, images are a primary source for fingerprinting, as previous studies also analyzed [8, 68], and thus we report their statistics separately.

3.2.3 Scope of Leakage: HTTP vs. Sequence of Domains (DNS or SNI) . HTTP/2 is not, in itself, sufficient to protect the subpage’s identity. For legacy clients — without DoH or ECH support — the target webpage may also be visible through other channels, such as the sequence of plaintext DNS queries or TLS SNI within the webpage load. To assess whether DNS contact ordering alone enables fingerprinting, we encoded the sequence of contacted domain names during webpage load as features and trained a k-FP classifier on the dataset. The resulting macro-F1 scores are: Amazon 0.10, BBC 0.07, Reddit 0.74, Udemy 0.01, Wikipedia 0.03. DNS contact ordering is therefore largely uninformative for subpage identity across four of five datasets. Reddit is the only outlier, likely due to its large number of unique servers per subpage (8.21 ± 0.1, Table 1), which causes subpage-specific TLS handshakes. BBC, despite a similarly large server count (8.03 ± 0.1), contacts largely the same servers across subpages, leaving DNS uninformative. Since HTTP/2 application-layer defenses cannot address DNS or TLS leakage — that requires orthogonal mechanisms such as DoH or ECH — we treat Reddit’s DNS leakage as a known limitation and focus our defenses on the HTTP packet-level side channel, which dominates in four of five datasets.

3.2.2 Fingerprinting on the Undefended Datasets. Figure 1 summarizes the fingerprinting performance of each model across the five datasets. We observe that, for every dataset, multiple hypertuned fingerprinting architectures achieve a macro-F1 score ≥ 0.9, indicating high webpage-identification accuracy. In particular, strong performance is achieved by both feature-engineered approaches such as k-FP, which rely on global traffic statistics, and deep-learning approaches that learn representations directly from packet sequences. Overall, these results show that an effective defense must mitigate both feature-engineered and deep-learningbased fingerprinting attacks. Intuitively, leakage sources group into: packet counts, timing/IAT (inter-arrival times), bursts (how data is grouped, e.g., initial HTML followed by resources), and CUMUL (the overall growth of the

Takeaways. We presented five datasets that exhibit diverse fingerprinting leakage characteristics: (i) some exhibit at least one dominant leaking subdomain/connection (Amazon, Udemy, Wikipedia); (ii) others involve multiple connections contributing to page identification (Reddit, BBC); (iii) certain datasets contain informative packet counts (Amazon, Wikipedia); (iv) some datasets leak information through processing duration (Amazon, BBC); and (v) all exhibit leakage through CUMUL and burst statistics. Any successful fingerprinting defense must address all of these side channels. 4

Understanding the Privacy-Preserving Potential of HTTP/2 Against Webpage Fingerprinting

4

Benchmarking WF Defenses

website 𝑦𝑖 appears among the 𝑘 highest-scoring predictions:

Defense benchmarks aim to quantify how closely any attacker could infer the correct website. However, relying solely on fingerprinting accuracy (Section 3.1.3) can be misleading: a defense may only defeat specific models or provide privacy at the cost of impractically high overhead. Moreover, implementation errors such as faulty preprocessing or incorrect feature scaling can artificially lower an attacker’s performance, creating a false sense of security. To address these pitfalls, we propose a novel Blueprint for Benchmarks and Quality Assurance (BBQ) of fingerprinting defenses. The framework calibrates each defense and evaluates it using five practical attackers (k-FP, DF, VarCNN, RobustFP-CNN, Holmes), information-theoretic leakage estimates from two security estimators (WeFDE [31] and DeepSE-WF [67]), and the privacy–overhead trade-off. We provide takeaways of our benchmarking blueprint at the end of Section 4. BBQ 0 Defense calibration: Which defense configuration is recommended for this website and the underlying protocol (here: HTTP/2)? Defense parameters inherited from prior protocols may not transfer directly to HTTP/2, and multiple parameters may jointly affect privacy and overhead. We therefore define and evaluate several defense configurations at increasing intensity levels. Cai et al. [7] similarly evaluated privacy–overhead trade-offs across defense parameter settings; however, their defenses were applied offline to packet sequences, whereas we implement each configuration in HTTP/2 and generate defended traces by replaying real webpages. Using a calibration set containing 100 traces per webpage, we characterize each configuration by the maximum Macro-F1 across the five practical attackers in BBQ 1, together with bandwidth and latency overheads measured as in BBQ 4. We refer to this maximum as the calibration Macro-F1. For each dataset, we then select a practical operating point from the observed privacy–overhead curve, favoring configurations that provide substantial reductions in calibration Macro-F1 while avoiding higher-intensity settings whose additional overhead yields limited privacy improvement. The selected defense parameters are then fixed and evaluated under BBQ 1–4 on a separately generated dataset containing 500 traces per webpage. BBQ 1 Practical Attacker Effectiveness: How accurately does an attacker perform against the defense? For each defense, we measure and report the maximum Macro-F1 score (Section 3.1.3) of the fingerprinting models introduced in Section 3.1.1 — the “baseline resilience" of each defense. Each attacker is hyperparameter-tuned independently for the corresponding dataset–defense pair, retaining its default configuration when tuning does not improve validation Macro-F1. Intuitively, the macroF1 value is ∼ 0.1 − 0.6 for strong defenses (against the evaluated attacker), and 0.8+ for a weak defense. BBQ 2 Practical Attacker Predictive Proximity: How close are the attackers to the correct answer? The F1 scores hide if a model “almost” predicted a perfect match. Using the same tuned attacker configurations as in BBQ 1, we therefore also report Top-𝑘 accuracy to measure an attacker’s proximity to the correct answer, i.e., the fraction of traces for which the correct

𝑁

Top-𝑘 Accuracy =

1 ∑︁ 1[ 𝑦𝑖 ∈ TopK(𝑝𝑖 , 𝑘) ] , 𝑁 𝑖=1

where 1[·] is the indicator function and 𝑝𝑖 is the scoring prediction. A score of 0.9 means the attacker correctly identifies the correct page within their top 𝑘 predicted candidates in 90% of the cases. BBQ 3 Theoretical Attacker Uncertainty under Optimal Conditions: How much uncertainty does the defense induce in an ideal attacker’s predictions? Fingerprinting defenses can be tuned to evade known ML techniques, creating a false sense of security — that is, BBQ 1–2 remain tightly coupled to specific fingerprinting models and may fail to capture the full leakage potential present in the traffic traces. To overcome this limitation, we measure the model-agnostic uncertainty — not just classification errors — using information-theoretic methods directly on the datasets. We characterize the defense’s impact on the attackers through the prior and the conditional entropies: 𝐻 (𝑌 ) = −

∑︁

𝑝 (𝑦) log2 𝑝 (𝑦),

  𝐻 (𝑌 | 𝑋 ) = E𝑥 𝐻 𝑃 (𝑌 | 𝑋 =𝑥) .

𝑦

where 𝑋 is the (defense-processed) traffic representation and 𝑌 the page label. 𝐻 (𝑌 ) quantifies the uncertainty about the webpage before observing the traffic, while 𝐻 (𝑌 | 𝑋 ) measures the uncertainty about webpage 𝑌 after observing the features 𝑋 (model agnostic). Using these entropies, the mutual information (MI) between the traffic representation 𝑋 and the label 𝑌 (with 𝐶 unique websites) is: 𝐼 (𝑋 ; 𝑌 ) = 𝐻 (𝑌 ) − 𝐻 (𝑌 | 𝑋 ) ∈ [0, log2 𝐶] bits. Unlike accuracy, 𝐼 (𝑋 ; 𝑌 ) also credits near misses: if a model consistently narrows the candidate set without identifying the exact page, 𝐼 still increases. Thus, 𝐼 measures how much uncertainty an optimal attacker could remove, independently of any specific classifier. Two well-established yet complementary MI estimators for WF are: (a) WeFDE [31] estimates the mutual information using manually selected features and kernel density estimators (KDE). WeFDE maps each trace to a vector of handcrafted features 𝜙 (𝑥) (the same features as for k-FP — IAT, packet, burst statistics, CUMUL), then estimates class-conditional and marginal densities with kernel density estimation (KDE) on low-dimensional feature groups. Plugging those into the MI identity yields:    𝑁 𝑝b 𝜙 (𝑥𝑖 ), 𝑦𝑖 1 ∑︁ b  log2 𝐼 WeFDE ≈ , 𝑁 𝑖=1 𝑝b 𝜙 (𝑥𝑖 ) 𝑝b(𝑦𝑖 ) where 𝑝b(𝜙, 𝑦) and 𝑝b(𝜙) are KDE estimates. Intuitively, the estimator approximates the relationship between the joint probability of observing a feature set 𝜙 together with a webpage label 𝑦 (b 𝑝 (𝜙, 𝑦)) and the product of their marginal probabilities 𝑝b(𝜙) and 𝑝b(𝑦). If the features are independent of the webpage, the joint probability equals the product of the marginal probabilities, yielding log2 (1) = 0. Otherwise, dependence between the features and the webpage yields a nonzero value, indicating higher dependence between the features and the labels. This produces a total MI estimate in bits and per-feature/group leakage, pinpointing what still leaks. (b) DeepSE-WF [67] leverages deep learning-derived latent spaces combined with specialized 𝑘-NN estimators to approximate MI. 5

Bogdan Cebere, Prateek Kumar, Sylvain Chatel, Wouter Lueks, and Christian Rossow

While WeFDE provides an interpretable method for measuring information leakage directly from raw data, its accuracy depends heavily on the quality of the selected input features, and poor choices can create a false sense of security. To better approximate information leakage, the framework trains a deep neural network to project website traces into a continuous latent representation 𝑍 = 𝑓 (𝑋 ). It then estimates 𝐼 (𝑍 ; 𝑌 ) using a 𝑘-NN-based estimator [54], which is specifically designed for mixtures of discrete labels (𝑌 ) and continuous features (𝑍 ):

Takeaways. All experiments in the following section use identical preprocessing and evaluation steps, ensuring that score differences arise solely from the traffic itself. Guided by our blueprint, we can: • Calibrate each defense per dataset to identify a practical privacy– overhead operating point (BBQ 0); • Detect when a dataset is easy to fingerprint (BBQ 1) or is trivial to get close to the real answer (BBQ 2); • Identify cases where current ML models underperform, compared to the actual data leakage (BBQ 3); • Compare manual-feature learning against deep-learning, keeping the stronger attacker in each case (BBQ 1–2: k-FP vs. VarCNN; BBQ 3: WeFDE vs. DeepSE); • Uncover design flaws (or bugs) that inflate overhead, that falsely appear beneficial (BBQ 4).

𝑁 𝑁 1 ∑︁ 1 ∑︁ 𝐼ˆ(𝑍 ; 𝑌 ) = 𝜓 (𝑁 ) − 𝜓 (𝑛𝑖 ) + 𝜓 (𝑘) − 𝜓 (𝑚𝑖 ), 𝑁 𝑖=1 𝑁 𝑖=1

where 𝜓 is the digamma function, 𝑁 is the total number of samples, 𝑛𝑖 is the number of samples sharing the class label of sample 𝑖, 𝑘 is a hyperparameter (set to 5 in this work), and 𝑚𝑖 is the number of points regardless of class lying within distance 𝜀𝑖 of 𝑥𝑖 , where 𝜀𝑖 is the distance from 𝑥𝑖 to its 𝑘-th nearest same-class neighbor. Finally, b 𝐼 DeepSE = max 𝑓 𝐼ˆ(𝑓 (𝑋 ); 𝑌 ). This geometric approach allows dimensionality-dependent parameters to cancel out, making the estimator more robust in high-dimensional latent spaces. However, by the data processing inequality [15], 𝐼 (𝑍 ; 𝑌 ) ≤ 𝐼 (𝑋 ; 𝑌 ), so 𝐼ˆ(𝑍 ; 𝑌 ) is a lower bound on the true leakage — meaning the actual information available to an adversary may be larger, and any privacy guarantee implied by 𝐼ˆ(𝑍 ; 𝑌 ) is therefore optimistic. Interpreting the uncertainty estimators. To make the privacy implications more interpretable, we express conditional entropy as an effective number of plausible webpage labels, 2𝐻 (𝑌 |𝑋 ) . Larger values indicate greater residual uncertainty. Since the true mutual information 𝐼 (𝑋 ; 𝑌 ) is unknown, we use the larger of the two estimated leakage values as a proxy and derive the estimator-based uncertainty measure K ∗ :

5

We now discuss how to model existing fingerprinting defenses using HTTP/2. We first review established defensive techniques (Section 5.1). We then highlight HTTP/2 features and how they can enable defenses (Section 5.2). We finally tie these threads of knowledge by emulating HTTP/2 application-layer defenses for clients (Section 5.3) and for servers (Section 5.4).

5.1

5.1.1 Fingerprinting Defenses Categories. Defenses can be broadly grouped into four categories: (1) Noise traffic techniques add dummy packets to obscure bandwidth patterns, often effective but costly in bandwidth [1, 23, 27, 29, 49, 63]; (2) Padding techniques (either through traffic morphing or adversarial perturbations) modify packet sizes to mimic other traffic patterns, usually with lower bandwidth overhead than pure noise injection, but often leaving timing patterns intact [3, 44, 50, 56, 71]. (3) Traffic splitting techniques, which alter the number of requests and responses [23, 29, 38, 63]; and (4) Traffic-shaping techniques enforce fixed-size packets at fixed intervals, achieving strong uniformity guarantees but at the cost of higher latency and bandwidth [6, 7, 20, 27, 37, 55].

b

Here, max{b 𝐼 WeFDE, b 𝐼 DeepSE } is used as a proxy for the unknown 𝐼 (𝑋 ; 𝑌 ); consequently, K ∗ is an estimator-derived quantity rather than a bound on the true 2𝐻 (𝑌 |𝑋 ) . If the estimators underestimate the true leakage, the true effective candidate-set size may be smaller than the reported K ∗ . The mutual-information estimates and the derived K ∗ values are reported as point estimates, as the employed estimators do not provide confidence intervals. BBQ 4 Defense Operational Cost: What is the added latency and traffic volume by the defense? For a webpage 𝑦 ∈ Y under defense 𝑑, let Up𝑑 (𝑦) denote the total upload bytes (client → server), Down𝑑 (𝑦) the total download bytes (server → client), and 𝑇𝑑 (𝑦) the duration until the last real content packet. Baseline values without defense are Up𝑏 (𝑦), Down𝑏 (𝑦), and 𝑇𝑏 (𝑦). To quantify overheads, we compute the relative per-page increase ratio: Δ𝑀 (𝑦) =

𝑀𝑑 (𝑦) − 𝑀𝑏 (𝑦) , 𝑀𝑏 (𝑦)

Fingerprinting Defenses

WF defenses aim to disrupt site-specific bandwidth and latency patterns, using either uniformity (making all page loads appear identical) or unpredictability (making each page visit look random).

K ∗ = 2𝐻 (𝑌 ) −max{ 𝐼WeFDE , 𝐼DeepSE } . b

Modelling WF Defenses in HTTP/2

5.1.2 Fingerprinting Defenses: Tor vs. Plain HTTPS. Most fingerprinting defense research focuses on Tor traffic [6, 7, 20, 23, 29, 37, 55, 69], where all data is multiplexed into fixed-size cells within a single tunnel. Plain HTTPS defenses instead complement Tor/VPN solutions: they do not hide the destination domain, but aim to reduce subpage leakage. This setting introduces additional challenges compared to the Tor/VPN defenses: packet sizes vary, and a page load spans multiple parallel TCP connections that can be fingerprinted independently. Some defenses are designed specifically for HTTPS [11], while others originally built for Tor can, in principle, be adapted [6, 23]. However, (1) traffic-shaping defenses must also address packet-size leakage and (2) unlike Tor, noise added to one connection does not propagate to the encapsulating layer — every leaking connection must be defended individually. Defenses can be deployed at different points in the communication path: (1) Client-side (in the browser, OS, or client application),

𝑀 ∈ {Up, Down,𝑇 }.

Similar to prior work [63], (1) we report the median (Q2) and the first and third quantiles (Q1–Q3) of Δ𝑀 across all pages, pooled across all case studies; and (2) the latency overhead excludes trailing defensive traffic after the last real content packet, as this does not affect the user’s perceived page load time. 6

Understanding the Privacy-Preserving Potential of HTTP/2 Against Webpage Fingerprinting

Table 2: Comparison between the HTTP versions. Feature

HTTP/1.1

HTTP/2

HTTP/3

Request Resource Range Requests Transport Flow Control Concurrent Streams Stream Prioritization Header Compression Server-Push 103 Early Hints Built-in Padding Connection Probing

✔ ✔ TCP ✕ ✕ ✕ ✕ ✕ ✔ ✕ ✕

✔ ✔ TCP ✔ ✔ ✔ ✔ (✔) ✔ Limited ≤ 255 B PING (8 B)

✔ ✔ UDP ✔ ✔ ✔ ✔ ✕ ✔ Unlimited PING (no payload)

be in-flight on the network at once, HTTP/2 flow control operates per-stream: the receiver advertises an initial window size for each stream and can increment it via WINDOW_UPDATE frames, allowing it to pace or pause individual streams independently without affecting others on the same connection. This per-stream granularity has direct fingerprinting implications: by varying window sizes or deliberately delaying WINDOW_UPDATE frames, the receiver can alter the sizes and timing of downloaded DATA frames at the HTTP layer, independently of what TCP would otherwise permit. 5.2.5 HTTP/2 Resource Suggestions. HTTP/2 servers can proactively suggest or even deliver resources to the clients. One option is Server-Push frames, which let a server proactively send responses linked to a client’s earlier request. For instance, after an HTML request, the server can push JavaScript, CSS, or images without waiting for additional requests. However, support for Server-Push has declined in recent years, with Chrome deprecating its support in 2022. Our evaluation of Server Push-based defenses serves primarily to establish theoretical bounds on server-side defense efficacy. A more widely adopted alternative is the HTTP ‘103 Early Hints’ status code [45], which lets servers indicate critical resources for the client to fetch proactively. Unlike Server-Push, this approach gives the client full control over which resources to request, at the cost of an extra round trip.

where traffic can be perturbed before it leaves the user’s device [3, 6, 11, 23, 29, 38, 63]; or (2) Server-side (at the web server or CDN), where content delivery patterns, resource sizes, or response delays can be modified [1, 11]. In the following sections, we focus on one-sided (client-only or server-only) HTTP/2 application-layer defenses powered by unpredictability.

5.2

HTTP/2 Privacy Potential

HTTP/2 represents roughly 60% of browser traffic [4] and introduces several features beyond HTTP/1.1 that can be relevant to fingerprinting defenses. Table 2 outlines the key similarities and differences across HTTP versions. As defined in RFC 9113 [65], the basic unit of the HTTP/2 communication is a frame. A stream is a bidirectional sequence of frames between the client and the server, mapped to a single client request. An HTTP/2 connection corresponds to a single TCP connection and comprises multiple streams, i.e., each carrying one or more requests and responses exchanged with the same server/subdomain.

5.2.6 HTTP/2 Built-in Padding. HTTP/2 offers limited support for padding for DATA, HEADERS, and PUSH_PROMISE frames, with up-to 255 bytes of extra noise. In contrast, QUIC provides dedicated PADDING frames [28], and HTTP/3 DATA frames can include arbitrary-length padding, making padding strategies more flexible than HTTP/2 ’s 255-byte limit per frame. 5.2.7 HTTP/2 Connection Probing. HTTP/2 uses PING frames for liveness checks and RTT measurements. Each PING frame must contain an 8-byte payload, which the remote party will echo.

5.2.1 HTTP/1.1 Inheritance. HTTP/2 inherits HTTP/1.1’s header semantics — the same fields for method, path, status, and content negotiation. It also inherits Range requests (RFC 7233 [21]), which allow clients to fetch specific byte ranges of a resource, enabling application-layer traffic splitting.

5.2.8 HTTP/2 Opportunities for Fingerprinting Defenses. Unlike QUIC, which offers fine-grained stream-level flow control and lets applications dictate when and how packets are sent, HTTP/2 relies on TCP, where the OS stack manages transmission. As a result, defenses that depend on strict fixed-rate packet flows can only be approximated in HTTP/2, since it lacks QUIC’s frame-level control. While HTTP/2 offers features that could enhance privacy, their defensive potential has not been systematically quantified. Prior studies generally found that website fingerprinting remains effective against HTTP/2 traffic [22, 32, 33, 39, 42, 42, 43, 70], but did not actively leverage HTTP/2-specific mechanisms to strengthen communication privacy. To that end, in addition to emulating established defenses with HTTP/2, we are also interested in exploring the privacy impact of the aforementioned HTTP/2 features.

5.2.2 HTTP/2 Multiplexing. On the transport side, HTTP/2 replaces HTTP/1.1 pipelining with stream multiplexing: multiple requests and responses are interleaved concurrently over a single persistent TCP connection without ordering constraints, eliminating head-of-line blocking. Beyond performance gains [2], multiplexing blurs the boundaries between individual encrypted requests and responses, with direct implications for connection privacy. 5.2.3 HTTP/2 Stream Prioritization. In HTTP/2, clients and servers can change the priority of streams depending on the content type (e.g., HTML, CSS, JS), the content location in the pages (above or below the visible fold), or from developer hints [47]. Stream reprioritization can also lead to unpredictable burst patterns, thus improving communication privacy.

5.3

HTTP/2 Client-Side Defenses

We now show how HTTP/2 clients can increase privacy in the five case studies from Section 3.2. To that end, we emulate and evaluate four defensive strategies using HTTP/2 features: (1) HTTPOS [38], a traffic splitting technique; (2) LLaMA [11], a traffic morphing defense; (3) FRONT [23], a noise insertion strategy; and (4) Tamaraw [7], a traffic-shaping defense emulated from the client-side. The

5.2.4 HTTP/2 Flow Control. HTTP/2 introduces its own flow control at the application layer, independent of TCP’s transport-layer flow control. While TCP flow control governs how much data can 7

Bogdan Cebere, Prateek Kumar, Sylvain Chatel, Wouter Lueks, and Christian Rossow send

send batch

ENTRYPOINT New Connection C Wrecv ← 512–16384 B

Pending requests? yes ⇒ continue no ⇒ exit

empty

U{𝑆 min, . . . , 𝑆 max }, across six candidate configurations spanning W = 16384–512 B and range-split bounds from [3, 5] to [20, 40]. Compared to the original implementation, we additionally vary the initial flow-control window. The six candidate configurations use W ∈ {16384, 8192, 4096, 2048, 1024, 512} B with corresponding range-split bounds [𝑆 min, 𝑆 max ] ∈ {[3, 5], [4, 7], [5, 8], [5, 10], [10, 20], [20, 40]}. Figure 3 shows that HTTPOS provides limited protection across the calibration sweep. On Amazon, increasing the defense intensity reduces the calibration Macro-F1 to 0.699 for W = 1024 B and N ∼ U{10, . . . , 20}. On BBC and Reddit, the calibration Macro-F1 remains above 0.97 under the same configuration, and on Udemy and Wikipedia it remains above 0.97 even with W = 512 B and N ∼ U{20, . . . , 40}. These results indicate that more aggressive range splitting alone does not substantially suppress fingerprinting leakage on most datasets. Following calibration, we therefore retain W = 1024 B with N ∼ U{10, . . . , 20} for Amazon, BBC, and Reddit, and W = 512 B with N ∼ U{20, . . . , 40} for Udemy and Wikipedia.

EXIT

non-empty

Split into byte-ranges N ∼ U{3, . . . , 40} overlapping chunks C ← {R1 , . . . , RN }

yes

Is binary content? content-type image/, video/, . . .

Unchanged C←R

no

Figure 2: HTTP/2 emulation of the HTTPOS Defense [38].

following describes how each defense is adapted to HTTP/2 and includes a state machine diagram showing its page-load behavior.

Macro-F1 (strongest attack)

Macro-F1 (strongest attack)

5.3.1 HTTPOS. HTTPOS [38] obscures traffic patterns by splitting requests using the Range header to fetch specific or overlapping parts of binary resources (e.g., images), and by constraining the receive window so that responses arrive in smaller units. The original defense manipulates the TCP advertised window; we approximate this behavior using HTTP/2 flow control, which constrains DATA delivery while leaving the frame-to-segment mapping to the kernel. Request splitting changes the number of transmitted units, while overlapping ranges change how many bytes are transferred. Yet, with this approach, the Range header is mainly applicable to binary data, and servers may ignore it altogether — e.g., by returning the full image even if only a byte range is requested. In our experiments, we assume the server cooperates. Figure 2 shows the HTTP/2 version of the defense. For the HTTPOS defense calibration (BBQ 0), we jointly vary the initial HTTP/2 flow-control window W and the lower and upper limits [𝑆 min, 𝑆 max ] of the range-split count N , with N ∼

5.3.2 LLaMA. LLaMA [11] is an application-layer defense that obfuscates traffic by randomizing request order, using HTTP pipelining, introducing random delays, and injecting dummy traffic. In HTTP/2, request pipelining is natively supported through stream multiplexing. Each stream can be delayed independently without stalling the entire connection. LLaMA can leverage these features by grouping requests into random batches (up to 5 in our implementation) and reordering them before sending. Every request is delayed by T ∼ U (0, 𝜏); the original draws 𝜏 as half the median page load time, which assumes concurrent dispatch, so we scale it by the mean number of requests per page to obtain the same latency budget under sequential replay. To further increase ambiguity, LLaMA probabilistically issues dummy HTTP requests after each real request and received response, injecting noise into the traffic pattern. Figure 4 illustrates the HTTP/2 version of the defense. For the LLaMA defense calibration (BBQ 0), we vary the dummyrequest probability over 𝑝𝑑 ∈ {0, 0.15, 0.25, 0.30, 0.50, 0.80} while keeping the dataset-specific delay bound 𝜏 fixed. Figure 5 shows that LLaMA generally benefits from increasing the dummy-request probability, but the magnitude of the improvement is strongly datasetdependent. On Amazon, the calibration Macro-F1 decreases to 0.296 at 𝑝𝑑 = 0.8, while on BBC, Reddit, and Udemy the same configuration reaches 0.931, 0.691, and 0.756, respectively. On Wikipedia, 𝑝𝑑 = 0.5 already reaches 0.784, while increasing 𝑝𝑑 to 0.8 improves

Calibration plot for the HTTPOS defense Amazon BBC Reddit Udemy Wikipedia

1.0 0.8 0.6 0.4 0.2 0.0

0x

1x

2x

3x

Upload Overhead compared to undefended traces

4x

1.0 0.8 0.6 0.4

ENTRYPOINT: New Connection C

0.2 0.0

non-empty

0x

0.5x

1x

2x

3x

4x

5x

loop

Latency Overhead compared to undefended traces

Figure 3: Calibration for the HTTPOS defense (BBQ 0). Strongest-attacker Macro-F1 per dataset against bandwidth (top) and latency (bottom) overhead. Horizontal lines mark the weak, moderate, and strong attacker thresholds; the dotted line is random guessing.

Shuffle queue re-permute pending requests

Pending Requests? yes ⇒ continue no ⇒ exit

Pick next - pop head of queue - append it to dummies

Send batch dispatch all queued (real + dummies)

empty

EXIT

Delay T ∼ U(0, 50ms)

Noise? p = 0.3 enqueue 1 random dummy

Figure 4: HTTP/2 emulation of the LLaMA Defense [11]. 8

Macro-F1 (strongest attack)

Calibration plot for the LLaMA defense Amazon BBC Reddit Udemy Wikipedia

1.0 0.8 0.6 0.4 0.2 0.0

0x

1x

2x

3x

Download Overhead compared to undefended traces

4x

1.0 0.8 0.6 0.4 0.2 0.0

0x

0.5x

1x

2x

3x

Latency Overhead compared to undefended traces

Calibration plot for the FRONT defense Amazon BBC Reddit Udemy Wikipedia

1.0 0.8 0.6 0.4 0.2 0.0

0x 1x 2x 3x 4x 5x 6x

8x

10x

15x

20x

Download Overhead compared to undefended traces

Macro-F1 (strongest attack)

Macro-F1 (strongest attack)

Macro-F1 (strongest attack)

Understanding the Privacy-Preserving Potential of HTTP/2 Against Webpage Fingerprinting

4x

1.0 0.8 0.6 0.4 0.2 0.0

0x

0.5x

1x

2x

3x

Latency Overhead compared to undefended traces

4x

Figure 5: Calibration for the LLaMA defense (BBQ 0). Strongest-attacker Macro-F1 per dataset against bandwidth (top) and latency (bottom) overhead. Horizontal lines mark the weak, moderate, and strong attacker thresholds; the dotted line is random guessing.

Figure 7: Calibration for the FRONT defense (BBQ 0). Strongest-attacker Macro-F1 per dataset against bandwidth (top) and latency (bottom) overhead. Horizontal lines mark the weak, moderate, and strong attacker thresholds; the dotted line is random guessing.

calibration Macro-F1 only to 0.758 while increasing both bandwidth and latency overhead. The calibration shows that LLaMA generally benefits from increasing the dummy-request probability, but the magnitude of the improvement is strongly dataset-dependent. Consequently, we select 𝑝𝑑 = 0.8 for Amazon, BBC, Reddit, and Udemy, and 𝑝𝑑 = 0.5 for Wikipedia.

QUIC, and the same noise-scheduling principles apply to HTTP/2. Figure 6 summarizes our HTTP/2 emulation and its parameters. For the FRONT defense calibration (BBQ 0), we vary the maximum per-connection dummy-stream budget 𝑁 max ∈ {20, 50, 110, 200, 350, 500}. Figure 7 shows a clear privacy–overhead knee on all five datasets. On Amazon, the calibration Macro-F1 decreases from 0.358 at 𝑁 max = 110 to 0.180 at 𝑁 max = 200, whereas increasing the budget to 350 provides almost no further improvement (0.176) while substantially increasing bandwidth overhead. Similarly, Reddit improves from 0.518 at 𝑁 max = 200 to 0.343 at 𝑁 max = 350, while increasing the budget to 500 reaches only 0.334; Wikipedia reaches 0.065 at 𝑁 max = 350 and 0.064 at 𝑁 max = 500. On BBC, 𝑁 max = 350 achieves the lowest observed Macro-F1 (0.593), with no benefit from increasing the budget to 500. On Udemy, increasing 𝑁 max from 200 to 350 improves the calibration Macro-F1 from 0.392 to 0.320, but increases downstream and upstream overhead from 8.97 and 7.16 to 21.55 and 16.21, respectively; we therefore retain the lower-overhead configuration. The calibration shows that FRONT exhibits a clear privacy– overhead knee on all five datasets, with the selected operating points yielding a calibration Macro-F1 of at most 0.593. We therefore select 𝑁 max = 200 for Amazon and Udemy, and 𝑁 max = 350 for BBC, Reddit, and Wikipedia.

5.3.3 FRONT. The FRONT defense [23] injects 𝑁 dummy packets at timestamps drawn from a Rayleigh(W) distribution, where the scale W ∼ U (0, 1) is randomized per connection. In our HTTP/2 adaptation, for each connection 𝑐, we sample the number of scheduled dummy transmissions as 𝑁𝑐 ∼ U{1, . . . , 𝑁 max }, independently sample W𝑐 ∼ U (0, 1) seconds, and draw 𝑁𝑐 transmission times from Rayleigh(W𝑐 ). Thus, both the dummy volume and timing pattern vary across connections. QCSD [63] adapted FRONT to ENTRYPOINT: New Connection c Wc ∼ U(0, 1) s Nc ∼ U{1, . . . , Nmax } Tc,i ∼ Rayleigh(W  c ), i = 1,. . . , Nc

Pending Requests on c? yes ⇒ continue no ⇒ EXIT

c Tc ← sort {Tc,i }N i=1

non-empty

spawn noise thread

Tc not empty

Noise Traffic Thread Tc empty? ⇒ exit Tc,i ← min(Tc ) remove Tc,i from Tc wait until Tc,i send noise request on c

Tc empty

5.3.4 CL-Tamaraw. The Tamaraw defense [7] obfuscates traffic by enforcing fixed-size packets at a constant rate. While feasible in Tor entry nodes — where traffic can be modified bidirectionally — this makes Tamaraw difficult to adapt to HTTP/2. QCSD [63] showed that similar behavior can be reproduced in QUIC using flow control and multiplexing. HTTP/2 frame sizes are under application

Noise Thread Done

Figure 6: HTTP/2 emulation of the FRONT defense [23]. 9

Bogdan Cebere, Prateek Kumar, Sylvain Chatel, Wouter Lueks, and Christian Rossow

ENTRYPOINT New Connection C W = 4096 B; TU ms; TD ms

Pending requests? yes ⇒ handle next no ⇒ EXIT

holding the flow-control window at W = 4096 B and the receivedelay bound T𝐷 at 10 ms. The receive-delay threshold is fixed to 4096 B. Figure 9 shows that CL-Tamaraw exhibits strong dataset dependence and early saturation on several datasets. On Amazon, the calibration Macro-F1 reaches 0.082 at T𝑈 = 20 ms and improves only to 0.070 at 10 ms, so we select 20 ms. On BBC, T𝑈 = 40 ms already achieves the lowest observed Macro-F1 (0.120); shorter intervals increase overhead without improving protection, so 40 ms is selected. Reddit continues to benefit from stronger shaping throughout the evaluated range, reaching 0.327 at T𝑈 = 5 ms, which is therefore selected. Udemy presents a substantially less favorable trade-off: reducing T𝑈 from 20 ms to 5 ms reduces the calibration Macro-F1 from 0.644 to 0.403, but increases downstream overhead from 8.23 to 26.41 and upstream overhead from 10.47 to 28.29. We therefore retain 20 ms as the practical operating point. Finally, Wikipedia reaches 0.175 at T𝑈 = 10 ms, while 5 ms provides only a modest additional reduction to 0.133; we select 10 ms. Using a calibration target of strongest-attacker Macro-F1 closest below 0.5, while avoiding unnecessarily costly configurations, we select T𝑈 = 20 ms for Amazon and Udemy, 40 ms for BBC, 5 ms for Reddit, and 10 ms for Wikipedia.

yes

Noise Thread wait TU ms C ← noise req. stop after 2 s

loop

Flow-Control Thread wait TD ms C ← WIN UPD(+W)

loop

Figure 8: HTTP/2 emulation of the CL-Tamaraw Defense [7].

Macro-F1 (strongest attack)

control, but the mapping of frames to TCP segments is governed by the kernel, so a fixed-rate fixed-size packet stream cannot be enforced from userspace; instead, we approximate constant-rate behavior through flow-control window updates — i.e., the client can request a small initial flow-control window, thereby limiting the size of DATA frames the server can send. After receiving a fixed amount of data, the client can intentionally delay sending window update frames — pausing the server’s transmission — to mimic fixed-rate response defenses. Due to the TCP flow control in the kernel, this does not achieve a perfectly constant rate; yet, it still alters the observable timing patterns of the response. In addition, the defense issues noise requests at a fixed rate, further disrupting burst and cumulative statistics. Figure 8 outlines our client-side emulation of Tamaraw using HTTP/2, noting that it is not a full replication of the original Tor defense. For the CL-Tamaraw defense calibration (BBQ 0), we vary only the dummy-request interval T𝑈 ∈ {320, 80, 40, 20, 10, 5} ms, while

Calibration plot for the CL-Tamaraw defense Amazon BBC Reddit Udemy Wikipedia

1.0 0.8 0.6 0.4 0.2 0.0

0x 1x 2x 3x 4x 5x 6x

8x 10x

15x

20x

Download Overhead compared to undefended traces

Macro-F1 (strongest attack)

5.3.5 Fingerprinting on Client-Defended Datasets. Table 3 summarizes the Macro-F1, Top-5 and K ∗ scores for each dataset under calibrated client-side defenses, using the strongest attacker for each dataset–defense pair, with attackers hyperparameter-tuned independently for each pair (BBQ 1 – 3). CL-Tamaraw provides the strongest resilience across all five datasets. We first look at the Macro-F1 score. FRONT provides the second strongest protection overall, with its best results on Amazon (0.54) and Wikipedia (0.43), but remains much less effective on BBC, Reddit, and Udemy (0.89, 0.88, and 0.89, respectively). CLTamaraw achieves the lowest Macro-F1 on every dataset, reaching 0.24 on Amazon and BBC, 0.38 on Reddit, and 0.42 on Wikipedia, but remains weak on Udemy (0.88). In contrast, HTTPOS is largely ineffective across all five datasets, with Macro-F1 ranging from 0.86 to 0.99. LLaMA similarly fails to provide meaningful protection, with Macro-F1 remaining above 0.80 for all datasets. From the attacker’s perspective, RobustFP-CNN and Holmes shine against heavier defenses (FRONT, CL-Tamaraw, LLaMA), while k-FP dominates against lighter defenses (HTTPOS). One explanation is that defended traces are longer (noise packets extend the sequence), giving the CNN models more context windows to find stable patterns that survive noise injection. The Top-5 results further distinguish defenses that merely disrupt the attacker’s top-ranked prediction from those that induce broader ranking uncertainty. With CL-Tamaraw, Top-5 accuracy falls to under 0.55 on Amazon, BBC, and Reddit, respectively, indicating that the true webpage is absent from even the attacker’s five highest-ranked candidates in roughly half of the traces. Wikipedia shows a similar, though weaker, effect (0.62). In contrast, Udemy remains highly identifiable under every defense: even with CLTamaraw, Top-5 accuracy is 0.99. FRONT exhibits meaningful Top-5 robustness primarily on Wikipedia (0.70), while its Top-5 accuracy remains at least 0.87 on the other datasets.

25x

1.0 0.8 0.6 0.4 0.2 0.0

0x

0.5x

1x

2x

3x

4x

Latency Overhead compared to undefended traces

5x

Figure 9: Calibration for the CL-Tamaraw defense (BBQ 0). Strongest-attacker Macro-F1 per dataset against bandwidth (top) and latency (bottom) overhead. Horizontal lines mark the weak, moderate, and strong attacker thresholds; the dotted line is random guessing. 10

Understanding the Privacy-Preserving Potential of HTTP/2 Against Webpage Fingerprinting

Table 3: Baseline resilience of the HTTP/2 Client Defenses against the best-performing evaluated attacker (BBQ 1 and 2), together with the estimator-derived anonymity-set proxy K ∗ (BBQ 3). Macro-F1 and Top-5 are means; their 95% CIs are below 0.02 and omitted.

substantial cost, while FRONT offers a lower-latency alternative on Amazon and Wikipedia. More broadly, the dataset-specific defense settings and strongest attackers vary across workloads, underscoring the need for multi-dataset evaluation even when defenses and attackers are individually tuned.

Dataset Metric HTTPOS LLaMA FRONT CL-TAM

Takeaways. We demonstrated practical emulations of popular client-side defenses using HTTP/2 primitives, highlighting both their strengths and limitations. Client-side defenses uniformly cover all connections within a page load, but cannot directly control download leakage and may incur substantial bandwidth and latency overhead. Their effectiveness is also strongly dataset-dependent, requiring defense parameters to be calibrated separately for each dataset rather than using a single operating point. Overall, CLTamaraw provides the strongest privacy, reaching K ∗ = 38.43 on Amazon but dropping to K ∗ = 2.48 on Udemy, at median overheads of ΔDown = 3.7 and Δ𝑇 = 1.0; FRONT provides a lower-latency alternative on Amazon and Wikipedia, with a median Δ𝑇 = 0.5 but similarly substantial bandwidth overhead.

Amazon

Macro-F1 Top-5 K∗

0.86 0.96 2.84

0.82 0.96 3.62

0.54 0.87 16.80

0.24 0.51 38.43

BBC

Macro-F1 Top-5 K∗

0.99 1.00 1.10

0.99 1.00 1.13

0.89 0.95 2.28

0.24 0.54 26.64

Reddit

Macro-F1 Top-5 K∗

0.98 1.00 1.22

0.97 0.99 1.31

0.88 0.98 2.57

0.38 0.47 34.08

Udemy

Macro-F1 Top-5 K∗

0.99 1.00 1.10

0.96 1.00 1.22

0.89 1.00 1.82

0.88 0.99 2.48

Wiki

Macro-F1 Top-5 K∗

0.99 1.00 1.16

0.98 1.00 1.24

0.43 0.70 27.66

0.42 0.62 27.89

5.4

Table 3 further reports model-agnostic anonymity-set sizes K ∗ under client defenses (BBQ 3), capturing attacker uncertainty beyond Top-𝑘 metrics (higher K ∗ indicates greater uncertainty). The estimates show that CL-Tamaraw induces genuine ambiguity rather than merely perturbing the top-ranked guesses: on BBC, despite Top-5= 0.54, K ∗ = 26.64, leaving a substantially broader set of plausible webpages. Across BBQ 1–3, CL-Tamaraw is the most robust defense, with FRONT competitive on Amazon and Wikipedia. Finally, Table 4 summarizes the privacy–cost trade-offs of the client defenses (BBQ 4). HTTPOS adds little downstream traffic (ΔDown = 0.04) but increases upload and latency (ΔUp = 0.81, Δ𝑇 = 1.04), while providing weak privacy. LLaMA has the lowest latency overhead (Δ𝑇 = 0.35) but roughly doubles traffic in both directions and likewise provides limited protection. FRONT and CL-Tamaraw incur the largest bandwidth costs: FRONT reaches ΔDown = 5.76, while CL-Tamaraw reaches 3.73. However, FRONT shows a smaller impact on the perceived latency. Overall (BBQ 1–4), CL-Tamaraw provides the strongest client-side privacy but at

5.4.1 ALPaCA. The ALPaCA defense [11] supports two modes: a mimicking mode, which reshapes a site toward a larger reference site, and a noise mode, which perturbs resource sizes and injects additional resources. We evaluate the noise mode because it does not require selecting a reference website and maps directly to randomized padding and noise injection in our HTTP/2 setting. ALPaCA alters the webpage’s actual content length (images, HTML, CSS, etc.) so that traffic patterns become indistinguishable and unpredictable to a passive adversary. The original design works by (1) padding each object (e.g., an image, CSS file, HTML) with random data and by (2) inserting noise traffic (fake resources) in the main HTML. Figure 10 illustrates the HTTP/2 adaptation of the defense. Here, object-size perturbation is emulated by injecting dummy bytes into HEADERS frames. Unlike the original ALPaCA, which embeds fake resource references in HTML, the HTTP/2 adaptation proactively

Table 4: Client-side defenses overhead: median (Q1–Q3) of the relative increase ratio Δ𝑀. Baseline averages shown for reference (BBQ 4). Defense

ΔUp

ΔDown

Δ𝑇

HTTPOS LLaMA FRONT CL-TAM

0.8 (0.6 − 1.4) 1.9 (1.7 − 2.6) 7.4 (4.4 − 11.1) 6.7 (2.8 − 11.6)

0.0 (0.0 − 0.1) 2.1 (1.7 − 2.4) 5.8 (3.1 − 8.9) 3.7 (1.9 − 8.4)

1.0 (0.3 − 2.1) 0.4 (0.2 − 0.5) 0.5 (0.2 − 1.1) 1.0 (0.5 − 3.0)

5.37 KB

2638.1 KB

2.97 s

Baseline avg.

HTTP/2 Server-Side Defenses

In this section, we benchmark established defenses from the perspective of the HTTP/2 server. The server-side defenses follow the same principles as the client-side defenses: they must be able to pad, split, delay, or add additional noise. We adopt two known fingerprinting defenses, which can be emulated at the application layer (1) ALPaCA [11], a web object morphing strategy; and (2) Tamaraw [7], emulated from the server perspective. At first glance, server-side defenses seem attractive since they could, in principle, protect all users. In practice, however, modern webpages are rarely monolithic. As shown in Table 1 and prior work [61], most page loads span multiple servers. Unlike clients, which observe and can defend all connections uniformly, servers act independently — and many may not deploy any protection. In our benchmarks, we therefore assume no shared proxy or centralized cloud deployment: each subdomain involved in the webpage load applies defenses individually, while clients follow proactive resource suggestions. The HTTP/2 adaptations deliver noise via Server Push frames as a proof-of-concept; while some browsers have dropped Server Push support [72], the mechanism can be replaced by 103 Early Hints — as demonstrated in Section 6.

11

Bogdan Cebere, Prateek Kumar, Sylvain Chatel, Wouter Lueks, and Christian Rossow

$

ENTRYPOINT New Connection C

Insert Noise? ∼ Bern(ppush )

PAD ∼ U{PADmin , . . . , PADmax } B Synthetic Noise Paths /noiseM , M ∈ [4·PAD, 100·PAD] = M

Mi ∼ U[M], i = 1..N PUSH PROMISE /noiseMi

Request received? yes ⇒ R, SZ, H no (conn. closed) ⇒ EXIT

Add padding? ∼ {0, 1} HSZ = (⌊SZ/PAD⌋ + 1) · PAD − SZ append HSZ random bytes to H

also incurs substantial downstream and latency overhead, the additional privacy gain remains large on every dataset. The calibration shows a sharp transition: padding drawn from [1024, 8000] B, with up to 10 noise objects and 𝑝 push = 0.5, reduces calibration MacroF1 below 0.15 on every dataset, whereas the immediately weaker configuration remains between 0.19 and 0.71. We therefore select this defense configuration for all five datasets.

$

N ∼ U{1, . . . , Nmax }

$

$

Send X Bytes HEADERS DATA

$

Figure 10: HTTP/2 ALPaCA Server-Side Defense [11]. $

0.4 0.2 0.0

0x 1x 2x 3x 4x 5x 6x

8x

10x

15x

Calibration plot for the SRV-Tamaraw defense 0.8 0.6 0.4 0.2 0.0

Macro-F1 (strongest attack)

Macro-F1 (strongest attack)

0.8 0.6 0.4 0.2 0.5x

1x

2x

3x

4x

Latency Overhead compared to undefended traces

0x 1x 2x 3x 4x 5x 6x

8x

10x

15x

Download Overhead compared to undefended traces

1.0

0x

Amazon BBC Reddit Udemy Wikipedia

1.0

Download Overhead compared to undefended traces

0.0

Pad response HSZ = (⌊SZ/P⌋ + 1) · P − SZ append HSZ random bytes to H

5.4.2 SRV-Tamaraw (SRV-TAM). We emulate Tamaraw from the server side using three mechanisms: (1) noise injection via Server Push, (2) response padding to a fixed-size multiple, and (3) paced transmission. Unlike ALPaCA’s randomized per-connection padding, SRV-Tamaraw uses a fixed padding size to enforce more uniform traffic shaping. The server tracks transmitted bytes with a counter Δ and pauses for 𝜏 whenever Δ reaches the threshold 𝐵, before resetting the counter. Figure 12 illustrates the HTTP/2 adaptation.

Amazon BBC Reddit Udemy Wikipedia

0.6

Send X Bytes ∆ += min(X , W, WCl ) if ∆ ≥ ∆LIM → Wait τ → ∆ = 0

Figure 12: HTTP/2 emulation of the SRV-Tamaraw Defense [7].

Calibration plot for the ALPaCA defense 0.8

$

N ∼ U{1, . . . , Nmax } PUSH PROMISE /noiseP

Request received? yes ⇒ R, SZ, H, client window WCl no (conn. closed) ⇒ EXIT

Macro-F1 (strongest attack)

Macro-F1 (strongest attack)

delivers noisy resources via Server Push frames, without modifying HTML content. For the ALPaCA defense calibration (BBQ 0), we jointly vary ALPaCA’s per-connection padding range, maximum number of pushed noise objects, and push probability 𝑝 push . Calibration is performed with ALPaCA enabled on all connections involved in the page load. The three parameters scale together because each strengthens the same padding-and-noise mechanism, and noise-object sizes are defined relative to the padding granularity. The four candidate configurations use padding ranges [256, 1024], [512, 4000], [768, 6000], and [1024, 8000] B, with corresponding maximum numbers of noise objects 2, 4, 6, and 10, and 𝑝 push ∈ {0.2, 0.3, 0.4, 0.5}, respectively. Figure 11 shows that ALPaCA requires relatively aggressive padding before providing substantial protection. Increasing the padding range from [768, 6000] B to [1024, 8000] B, the maximum number of noise objects from 6 to 10, and 𝑝 push from 0.4 to 0.5 reduces the calibration Macro-F1 from 0.312 to 0.147 on Amazon, 0.401 to 0.068 on BBC, 0.229 to 0.027 on Reddit, 0.187 to 0.035 on Udemy, and 0.708 to 0.033 on Wikipedia. Although this increase

1.0

Insert Noise? ∼ Bern(ppush )

ENTRYPOINT New Connection C W, τ, ∆ = 0, B P (padding constant) Synthetic Noise /noiseP

5x

Figure 11: Calibration for the ALPaCA defense (BBQ 0). Strongest-attacker Macro-F1 per dataset against bandwidth (top) and latency (bottom) overhead. Horizontal lines mark the weak, moderate, and strong attacker thresholds; the dotted line is random guessing.

1.0 0.8 0.6 0.4 0.2 0.0

0x 0.5x 1x

2x

3x

4x

5x

7x

Latency Overhead compared to undefended traces

Figure 13: Calibration for the SRV-Tamaraw defense (BBQ 0). Strongest-attacker Macro-F1 per dataset against bandwidth (top) and latency (bottom) overhead. Horizontal lines mark the weak, moderate, and strong attacker thresholds; the dotted line is random guessing. 12

Understanding the Privacy-Preserving Potential of HTTP/2 Against Webpage Fingerprinting

Table 5: Baseline resilience of the HTTP/2 Server Defenses against the best-performing evaluated attacker (BBQ 1 and 2), together with the estimator-derived anonymity-set proxy K ∗ (BBQ 3). The best scores among single-server deployments are underlined, while the best-performing deployments overall are highlighted in bold. Macro-F1 and Top-5 are means; their 95% CIs are below 0.02 and omitted.

For the SRV-Tamaraw defense calibration (BBQ 0), we jointly vary the padding constant 𝑃, output flow-control window 𝑊 , send delay 𝜏, pacing threshold 𝐵, maximum number of pushed noise objects 𝑁 max , and push probability 𝑝 push . These parameters scale together to increase shaping intensity through larger padding, tighter flow control, longer pauses, and more noise. The four candidate configurations use (𝑃,𝑊 , 𝜏, 𝐵, 𝑁 max, 𝑝 push ) values (1024, 16384, 0.2, 16384, 2, 0.2), (4096, 8192, 0.5, 8192, 4, 0.3), (6144, 4096, 0.8, 6144, 6, 0.4), and (8092, 2048, 1.0, 4096, 10, 0.5), respectively, where 𝑃, 𝑊 , and 𝐵 are in bytes and 𝜏 is in milliseconds. Figure 13 shows a clear privacy–overhead knee at 𝑃 = 6144 B, 𝑊 = 4096 B, 𝜏 = 0.8 ms, 𝐵 = 6144 B, 𝑁 max = 6, and 𝑝 push = 0.4. On Amazon, this configuration reaches a calibration Macro-F1 of 0.063. On BBC and Reddit, it reaches 0.035 and 0.031, respectively, and the stronger configuration does not improve protection. On Udemy, the stronger configuration reduces Macro-F1 only from 0.031 to 0.015, while latency overhead increases from 2.11 to 8.19. Similarly, on Wikipedia, it reduces Macro-F1 from 0.055 to 0.020 while downstream overhead increases from 5.84 to 10.51 and latency overhead from 2.56 to 4.52. The calibration shows a clear knee at 𝑃 = 6144 B, 𝑊 = 4096 B, 𝜏 = 0.8 ms, 𝐵 = 6144 B, 𝑁 max = 6, and 𝑝 push = 0.4, reducing calibration Macro-F1 to 0.031–0.063 across all five datasets. Stronger shaping provides limited additional protection at substantially higher overhead, so we select this configuration for all datasets.

ALPaCA Data Metric st 1 Party CDN

SRV-TAM All 1st Party CDN

F1 Amz. Top-5 K∗

1.00 1.00 1.01

0.76 0.94 5.21

0.21 0.59 62.15

1.00 1.00 1.02

0.26 0.51 24.47

0.08 0.30 77.10

F1 Top-5 K∗

0.99 1.00 1.06

0.87 0.99 2.32

0.20 0.54 23.56

0.94 0.96 1.62

0.90 0.98 1.70

0.15 0.43 63.08

F1 Reddit Top-5 K∗

0.33 0.45 39.31

0.97 1.00 1.37

0.04 0.18 83.42

0.76 0.92 5.23

0.98 1.00 1.23

0.06 0.22 82.19

F1 Udemy Top-5 K∗

0.64 0.91 8.14

0.49 0.85 9.46

0.08 0.27 57.93

0.41 0.54 14.55

0.83 0.99 3.43

0.03 0.12 97.51

F1 Wiki Top-5 K∗

0.09 0.17 76.26

0.97 1.00 1.36

0.06 0.20 65.72

0.69 0.93 7.83

0.97 0.99 1.34

0.08 0.25 74.92

BBC

5.4.3 Fingerprinting on Server-Defended Datasets. For each website, we benchmark the defenses applied to either the main page connection (1st Party), the leakiest third-party server (CDN), or all servers involved (All). Table 5 reports the privacy of the calibrated server-side defenses against the strongest hyperparameter-tuned attacker for each dataset–defense pair (BBQ 1 and BBQ 2). Placement remains the dominant factor. On Amazon, defending the CDN connection is already highly effective for SRV-TAM (F1 = 0.26, K ∗ = 24.47), whereas defending only the first-party server provides essentially no protection. BBC shows the opposite extreme: neither singleserver placement is sufficient, but deployment across all servers reduces F1 to 0.20 for ALPaCA and 0.15 for SRV-TAM. Reddit, Udemy, and Wikipedia exhibit stronger first-party leakage: ALPaCA reaches F1 = 0.33 and 0.09 on Reddit and Wikipedia, respectively, while SRV-TAM reaches 0.41 on Udemy. The anonymity-set estimates (BBQ 3) reinforce this placement effect. Targeting a single dominant connection can already provide substantial ambiguity—for example, K ∗ = 24.47 for SRV-TAM on Amazon’s CDN, 39.31 for ALPaCA on Reddit’s first-party server, and 76.26 for ALPaCA on Wikipedia’s first-party server. When leakage is distributed across connections, however, full deployment is necessary: on BBC, K ∗ rises from at most 2.32 under any singleserver placement to 23.56 with ALPaCA-All and 63.08 with SRVTAM-All. Full deployment is also particularly effective on Udemy, where SRV-TAM reaches F1 = 0.03 and K ∗ = 97.51. Overall, the results show that server-side defenses can provide strong privacy, but their effectiveness depends critically on placing the defense at the connections carrying the dominant fingerprinting signal rather than simply instrumenting the first-party server. Table 6 reports the overhead of the calibrated server-side defenses (BBQ 4). Both defenses incur zero upload overhead because

All

the client sends no additional requests. Because each defense is calibrated per dataset toward its strongest practical privacy operating point, the reported overhead also reflects the cost of the selected protection level rather than a matched-cost comparison. In these conditions, SRV-TAM is nevertheless consistently cheaper than ALPaCA. Thus, ALPaCA’s higher overhead partly reflects the stronger configuration selected during calibration, while SRV-TAM provides the more favorable privacy–overhead trade-off overall. 5.4.4 Takeaways. We demonstrated practical emulations of established server-side defenses using HTTP/2 features, with defense parameters calibrated per dataset and attackers’ hyperparameters Table 6: Server-side defenses overhead ratio: median (Q1– Q3) of the relative increase ratio Δ𝑀 pooled across all five case studies, for single-server (top) and all-servers (bottom) deployments. (BBQ 4). ΔUp = 0 throughout, as the client sends no additional requests. Defense

Δ𝑇

ALPaCA

Single Srv. All Srv.

4.9 (2.9 − 7.8) 11.7 (7.1 − 15.8)

0.7 (0.4 − 1.7) 2.4 (1.4 − 4.2)

SRV-TAM

Single Srv. All Srv.

3.6 (2.1 − 6.3) 6.8 (4.4 − 9.6)

0.5 (0.3 − 1.5) 1.7 (1.0 − 2.6)

2638.10 KB

2.97 s

Baseline avg. 13

ΔDown

Bogdan Cebere, Prateek Kumar, Sylvain Chatel, Wouter Lueks, and Christian Rossow

tuned independently for each dataset–defense pair. The results remain strongly dataset- and placement-dependent: the most effective single-server deployment varies across case studies, while protecting all participating servers generally provides the strongest privacy at higher cost. Overall, SRV-TAM provides the stronger privacy– overhead trade-off for full deployment, while targeted single-server defenses can be highly effective when one connection dominates the leakage (e.g., Amazon, Reddit, Udemy or Wikipedia).

ENTRYPOINT New Connection C Sample Initial Window W

Flow Control Thread Must Update Window? If > B Bytes received Wait ∼ U(0, τ ) C ← window(+W) Resample B

PING Thread $

Insert PING? ∼ Bern(pping ) Send PING. Wait ACK Wait ∼ U(0, 0.01)

Shuffle and Batch N ∼ U{1, . . . , 5} reqs. Batch and Shuffle max N reqs. → R

Pending requests in C? yes ⇒ continue no ⇒ EXIT

Noise Guard $

Insert Noise? ∼ Bern(pguard ) Sample Nd noise streams [NA , R, NB ] → R

Figure 14: A privacy-conscious HTTP/2 client (H2PC) flow.

6

The Untapped Potential of HTTP/2 for Fingerprinting Defenses

HTTP/2 features against primary leakage sources, yielding more efficient and scalable defenses. A notable idea comes from the HTTPOS [38] defense, which focuses on binary resources — often the leakiest. Its strategy can be generalized: large resources can be guarded by surrounding them with noise streams, allowing HTTP/2 to multiplex guarding streams alongside legitimate traffic, with the frame scheduler interleaving DATA frames from multiple streams in a pattern that depends on window sizes, priorities, and implementation-specific scheduling. Unlike HTTPOS, this approach is not limited to binary content and does not require server cooperation. And unlike LLaMA, FRONT, or CL-Tamaraw, it targets the sensitive streams directly rather than relying on opportunistic overlaps between noise and real traffic. Figure 14 depicts the client-side defense workflow. For each connection C, the client (i) randomizes the initial receive window W using the flow-control parameters and (ii) launches an independent PING thread whose frames, when serialized into TCP segments, probabilistically interleave with DATA-frame segments, perturbing observable burst-direction patterns. Pending requests are then grouped into randomized batches and reordered (as in LLaMA), while each request may be accompanied by sampled guarding noise, altering CUMUL and packet-level statistics. After receiving 𝐵 bytes, the client introduces additional randomness by delaying the receivewindow update by a value sampled from U (0, 𝜏) and resampling 𝐵, further obscuring timing patterns. Resampling prevents window updates from occurring at deterministic byte intervals, which could otherwise become a learnable signature. For defense calibration (BBQ 0), we calibrate five parameters: the guarding-stream limit 𝑁𝑑 , PING probability 𝑝 ping and count range, receive-delay bound 𝜏, and receive-delay threshold 𝐵. The two weakest configurations use no guarding streams, isolating the effect of HTTP/2-native randomization before guarding noise is introduced: the first uses randomized flow control alone, while the second adds PING padding, both with negligible measured bandwidth overhead. The six candidate configurations progressively increase the guardingstream and PING activity while tightening the receive-delay parameters. The first uses 𝑁𝑑 = 0, 𝑝 ping = 0, PING count [1, 1], 𝜏 = 50 𝜇s, and 𝐵 = 20000 B; the second uses 𝑁𝑑 = 0 𝑝 ping = 0.25, PING count [1, 2], 𝜏 = 50 𝜇s, and 𝐵 = 20000 B; and the third uses 𝑁𝑑 = 1, 𝑝 ping = 0.35, PING count [1, 2], 𝜏 = 80 𝜇s, and 𝐵 = 15000 B. The remaining configurations use (𝑁𝑑 , 𝑝 ping ) = (1, 0.50), (2, 0.70), and (3, 1.00), with PING-count ranges [1, 3], [1, 5], and [2, 8], receivedelay bounds of 100, 200, and 500 𝜇s, and thresholds of 10000, 5000, and 2500 B, respectively. Figure 15 shows a comparatively consistent reduction in attacker performance as H2PC intensity increases.

We showed that HTTP/2 users can emulate established WF defenses, achieving privacy improvements across all case studies and deployment perspectives (client and server). Yet, these guarantees often come at considerable overhead or depend on carefully targeting the right servers — limitations that stem from the fact that these defenses were never designed with HTTP/2 ’s architecture in mind.

6.1

Opportunities with HTTP/2

Can we do better? In this section, we outline various novel strategies for leveraging HTTP/2 features to raise the baseline privacy of page loads and to overcome key limitations of existing defense designs. 6.1.1 Client Side. Most defenses assume noise is essential, overlooking the privacy benefits inherent in application-layer behavior, such as multiplexing and flow control. We first detail HTTP/2’s features, which can improve privacy from the client perspective, then discuss how to insert noise more efficiently. Client Opportunity 1: Randomizing client behavior by using HTTP/2 features can improve the baseline privacy. A light privacy-conscious HTTP/2 client (H2PC) can leverage the following features: • Multiplexing & Prioritization. The client can shuffle and batch a subset of the pending requests, proactively altering the burst patterns of the connection. • Flow Control. While the Tamaraw defense uses flow control for uniform traffic shaping, we can also turn it into an unpredictabilitydriven defense. Concretely, the client can randomize flow control window sizes per connection, enforcing different burst patterns on each page reload. • Connection Probing. PING frames carry only 8 bytes, so they minimally affect bandwidth. As long as the mechanism is not abused, inserting a random number of PINGs disrupts burst patterns when combined with prior techniques. Client Opportunity 2: HTTP/2 primitives allow clients to create ‘guarding’ noise streams. A key lesson from existing defenses is that more noise generally yields stronger privacy. For example, if we vary the noise quantity ceiling in the FRONT defense for the Udemy dataset between 10 → 500, we get a monotonic F1 score variation between [0.48 − 0.76]. Yet, the critical question of “how much noise is enough” remains largely unaddressed. To move beyond brute-force noise, we leverage 14

Macro-F1 (strongest attack)

Macro-F1 (strongest attack)

Understanding the Privacy-Preserving Potential of HTTP/2 Against Webpage Fingerprinting

Calibration plot for the H2PC defense

Table 7: Client-Side Defense Landscape with HTTP/2. H2PC vs. best performing defense (BBQ 1, 2, 3, 4). Macro-F1 and Top-5 are means; their 95% CIs are below 0.02 and omitted.

Amazon BBC Reddit Udemy Wikipedia

1.0 0.8 0.6 0.4 0.2 0.0

0x

1x

2x

3x

Download Overhead compared to undefended traces

4x

1.0

Dataset

Defense

F1

Top-5

K∗

∆Up

∆Down

∆T

Amz.

CL-TAM H2PC

0.24 0.32

0.51 0.60

38.4 20.1

3.3 2.1

3.2 2.0

1.4 0.4

BBC

CL-TAM H2PC

0.24 0.51

0.54 0.77

26.6 10.8

2.5 2.3

4.1 3.3

0.3 0.01

Reddit

CL-TAM H2PC

0.38 0.68

0.47 0.84

34.1 8.11

16.3 1.1

3.0 1.3

0.6 0.01

Udemy

CL-TAM H2PC

0.88 0.70

0.99 0.87

2.48 7.31

10.5 2.8

8.2 2.8

4.7 0.2

Wiki

CL-TAM H2PC

0.42 0.70

0.62 0.87

27.9 7.32

6.6 1.4

2.2 1.6

1.1 0.5

0.8 0.6 0.4 0.2 0.0

0x

0.5x

1x

2x

3x

Latency Overhead compared to undefended traces

4x

Server Opportunity 1: Randomizing HTTP/2’s built-in feature behavior can improve web-browsing privacy.

Figure 15: Calibration for the H2PC defense (BBQ 0). Strongest-attacker Macro-F1 per dataset against bandwidth (top) and latency (bottom) overhead. Horizontal lines mark the weak, moderate, and strong attacker thresholds; the dotted line is random guessing.

A privacy-conscious HTTP/2 server (H2PS) can use: • Multiplexing & Prioritization. The server can buffer multiple requests or delay streams, leading to unpredictable burst patterns. • Flow Control. HTTP/2 servers can artificially constrain their transmission rate by maintaining an internal flow-control limit smaller than the receiver’s advertised window, effectively under-utilizing the available window capacity. • Connection Probing. Similar to clients, servers can send PING frames to disrupt burst patterns.

On Amazon, the calibration Macro-F1 decreases from 0.167 with 𝑁𝑑 = 1, 𝑝 ping = 0.5, 𝜏 = 100 𝜇s, and 𝐵 = 10000 B to 0.086 with 𝑁𝑑 = 2, 𝑝 ping = 0.7, 𝜏 = 200 𝜇s, and 𝐵 = 5000 B; the strongest configuration further reduces it to 0.046, but with additional bandwidth and latency overhead. Wikipedia shows a similar knee, decreasing from 0.290 to 0.130 and then 0.105 across the same configurations. In contrast, BBC, Reddit, and Udemy continue to obtain substantial privacy gains at the strongest configuration, reaching Macro-F1 values of 0.232, 0.444, and 0.109, respectively. Calibration reveals two operating regimes: Amazon and Wikipedia favor an intermediate privacy–overhead point, whereas BBC, Reddit, and Udemy justify the strongest configuration. For Amazon and Wikipedia, we select 𝑁𝑑 = 2, 𝑝 ping = 0.7, PING count [1, 5], 𝜏 = 200 𝜇s, and 𝐵 = 5000 B, reaching calibration Macro-F1 ≤ 0.13. For BBC, Reddit, and Udemy, we select the strongest configuration, 𝑁𝑑 = 3, 𝑝 ping = 1.0, PING count [2, 8], 𝜏 = 500 𝜇s, and 𝐵 = 2500 B, reaching calibration Macro-F1 ≤ 0.44. Table 7 compares H2PC against CL-Tamaraw, the strongest clientside defense overall, across privacy and overhead (BBQ 1–4). H2PC keeps attacker Macro-F1 at or below 0.70 on all five datasets while substantially reducing latency and downstream overhead. On Amazon, BBC, Reddit, and Wikipedia, this yields a lower-cost privacy operating point: H2PC maintains K ∗ ≥ 7.32 while reducing Δ𝑇 and ΔDown relative to CL-Tamaraw. The trade-off is especially favorable on Udemy, where H2PC improves K ∗ from 2.48 to 7.31, while reducing Δ𝑇 from 4.7 to 0.2 and ΔDown from 8.2 to 2.8. Overall, H2PC provides competitive privacy at substantially lower cost.

Server Opportunity 2: HTTP/2 primitives allow 1st Party servers to defend the entire webpage. When a HTTP/2-conformant client prefetches suggested resources, servers can suggest noise traffic using mechanisms such as ServerPush or “103 Early Hints” (as with ALPaCA and SRV-Tamaraw). Notably, “103 Early Hints” [45] work across servers: a server can reference resources from any domain, which a client would fetch. This adds a new dimension to server-side defenses, enabling first parties to protect the entire webpage.

ENTRYPOINT New Connection C Wout ∼ U{212 , . . . , 214 } B Local Hints Noise /noiseM , M ∼ sizes on C 3rd-party Hints Noise S: Links in webpage

On Request R Client Window WCl $

Batch? ∼ {0, 1} Delay response

6.1.2 Server Side. We conduct a similar analysis on the privacypreserving impact of HTTP/2 features from the server perspective.

PING Thread $

Insert PING? ∼ {0, 1} Send PING. Wait ACK Wait ∼ U(0, 0.01) s

Send X Bytes X = min(WCl , Wout , |data|)

Local 103 Early Hints

3rd-Party 103 Early Hints

$

Noise? ∼ {0, 1} k ∼ U{1, . . . , kmax } for i = 1, . . . , k Sample URLi ∈ S Hint → URLi

Noise? ∼ {0, 1} k ∼ U{1, . . . , kmax } for i = 1, . . . , k Mi ∼ sizes on C, jittered Hint → /noiseMi

$

Figure 16: A privacy-conscious HTTP/2 server (H2PS) flow. 15

Macro-F1 (strongest attack)

Macro-F1 (strongest attack)

Bogdan Cebere, Prateek Kumar, Sylvain Chatel, Wouter Lueks, and Christian Rossow

Calibration plot for the H2PS defense

Table 8: Server-Side Defense Landscape with HTTP/2. Comparison of deploying the strongest defense on the leakiest server compared to H2PS 1st Party (BBQ 1-4). Macro-F1 and Top-5 are means; their 95% CIs are below 0.01 and omitted.

Amazon BBC Reddit Udemy Wikipedia

1.0 0.8 0.6 0.4 0.2 0.0

0x

1x

2x

3x

4x

Download Overhead compared to undefended traces

Dataset

Defense

F1

Top-5

K∗

∆Down

∆T

Amz.

TAM (CDN) H2PS

0.26 0.41

0.51 0.73

24.5 29.2

2.2 2.4

1.5 1.1

BBC

ALP (CDN) H2PS

0.87 0.18

0.99 0.39

2.3 42.5

12.1 2.4

0.5 0.0

Reddit

ALP (1st) H2PS

0.33 0.31

0.45 0.49

39.3 44.2

3.4 1.2

0.5 0.8

Udemy

TAM (1st) H2PS

0.41 0.52

0.54 0.79

14.6 13.2

3.7 1.5

1.5 0.1

Wiki

ALP (1st) H2PS

0.09 0.05

0.17 0.15

76.3 94.7

5.0 2.3

1.4 2.5

5x

1.0 0.8 0.6 0.4 0.2 0.0

0x

0.5x

1x

2x

3x

Latency Overhead compared to undefended traces

4x

Figure 17: Calibration for the H2PS defense (BBQ 0). Strongest-attacker Macro-F1 per dataset against bandwidth (top) and latency (bottom) overhead. Horizontal lines mark the weak, moderate, and strong attacker thresholds; the dotted line is random guessing.

Macro-F1 of 0.076 and 0.021, respectively. BBC and Udemy select [6, 18], reaching 0.051 and 0.121, as stronger configurations incur substantially higher overhead for limited additional benefit. Wikipedia requires [24, 72], where calibration Macro-F1 decreases from 0.169 at [12, 36] to 0.016. The calibration shows datasetdependent knees: BBC and Udemy select [6, 18] Early Hints, Amazon and Reddit [12, 36], and Wikipedia [24, 72]; the selected configurations yield calibration Macro-F1 ≤ 0.121 across all five datasets. All selected configurations use one PING, with HPACK randomization enabled for Amazon, Reddit, and Wikipedia. Table 8 summarizes the privacy–overhead trade-off of H2PS against the strongest single-server baseline (BBQ 1–4). For each dataset, the baseline defense (ALPaCA or SRV-Tamaraw) is deployed on the leakiest server, which may be either the 1st -party server or a CDN; H2PS, in contrast, is always deployed only at the 1st party server and requires no third-party cooperation. Despite this constraint, H2PS keeps the attacker’s estimated candidate set above 13 webpages across all five datasets. Its protection is also more consistent across datasets: H2PS reaches K ∗ = 13.2–94.7, whereas the corresponding leakiest-server baselines range from 2.3 to 76.3. The gains are most pronounced on BBC and Wikipedia, while H2PS remains competitive on Reddit. Amazon presents a mixed trade-off: directly defending the leaky CDN yields lower F1 and Top-5, while H2PS achieves a larger candidate set (29.2 vs. 24.5). On Udemy, the leakiest-server baseline retains a small privacy advantage. H2PS also offers a favorable overhead profile. On four of five datasets, it reduces downstream overhead by 54–80% relative to the corresponding leakiest-server baseline; on Amazon, the two are comparable (2.4 vs. 2.2). Latency is lower on Amazon, BBC, and Udemy, while Reddit and Wikipedia trade additional latency for comparable or stronger privacy. Overall, H2PS provides more consistent protection using only the 1st -party server, without requiring cooperation from third-party or CDN servers.

Figure 16 shows the H2PS defense workflow. For each new connection 𝐶 to the 1𝑠𝑡 −Party (e.g., www.bbc.com), the server randomizes its outbound flow-control limit Wout and spawns an independent PING thread to disrupt burst patterns, similar to H2PC. For every incoming request, the server probabilistically issues two types of “103 Early Hints” before returning the actual response: (1) synthetic local hints, served directly by the defended 1st server; and (2) third-party hints, with real resources hosted on other CDN servers observed in the page load (e.g., by extracting them from the HTML). This is a key advantage of “Early Hints” over “Server Push”: as Early Hints send only Link headers rather than actual content, the 1st -party server can hint resources on any origin without CDN cooperation or a shared proxy. Assuming client cooperation, the hint fetches are multiplexed with legitimate HTTP/2 frames, perturbing the network metadata observed by a passive adversary. Requests are also probabilistically batched and delayed, leading to timing unpredictability. For defense calibration (BBQ 0), H2PS varies the number of “103 Early Hints” per defended connection ([1, 1] to [40, 120]), PING padding, and HPACK randomization at the strongest configurations. The eight candidate configurations use Early Hints ranges [1, 1], [1, 2], [1, 5], [1, 10], [6, 18], [12, 36], [24, 72], and [40, 120]. The first configuration uses no PING padding; all remaining configurations use one PING, and HPACK randomization is enabled for the three strongest ranges, starting at [12, 36]. Figure 17 shows that lighter H2PS configurations provide limited protection, with dataset-dependent knees at stronger configurations. Amazon and Reddit select the [12, 36] Early Hints range, reaching calibration 16

Understanding the Privacy-Preserving Potential of HTTP/2 Against Webpage Fingerprinting

Table 9: Summary of the HTTP/2 WF defenses. For server-side baselines, K ∗ ranges use the best single-server placement (1st party or CDN) for each dataset. Defense

7

∗ − K∗ ] [Kmin max

Overhead

Transactions on Dependable and Secure Computing 18, 2 (March 2021), 505–517. doi:10.1109/TDSC.2019.2907240 [4] David Belson and Lucas Pardue. 2023. Examining HTTP/3 usage one year on — blog.cloudflare.com. https://blog.cloudflare.com/http3-usage-one-year-on/ [5] Sanjit Bhat, David Lu, Albert Kwon, et al. 2019. Var-CNN: A Data-Efficient Website Fingerprinting Attack Based on Deep Learning. Proc. Priv. Enhancing Technol. 2019, 4 (2019), 292–310. doi:10.2478/POPETS-2019-0070 [6] Xiang Cai, Rishab Nithyanand, and Rob Johnson. 2014. CS-BuFLO: A Congestion Sensitive Website Fingerprinting Defense. In Proceedings of the 13th Workshop on Privacy in the Electronic Society. Association for Computing Machinery, New York, NY, USA, 121–130. doi:10.1145/2665943.2665949 [7] Xiang Cai, Rishab Nithyanand, Tao Wang, et al. 2014. A Systematic Approach to Developing and Evaluating Website Fingerprinting Defenses. In ACM CCS. 227–238. doi:10.1145/2660267.2660362 [8] Bogdan Cebere and Christian Rossow. 2024. Understanding Web Fingerprinting with a Protocol-Centric Approach. In RAID. 17–34. doi:10.1145/3678890.3678910 [9] Heyning Cheng and Ron Avnur. 1998. Traffic Analysis of SSL Encrypted Web Browsing. Project Paper. University of California, Berkeley. [10] Giovanni Cherubin. 2017. Bayes, not Naïve: Security Bounds on Website Fingerprinting Defenses. Proc. Priv. Enhancing Technol. 2017, 4 (2017), 215–231. doi:10.1515/POPETS-2017-0046 [11] Giovanni Cherubin, Jamie Hayes, and Marc Juarez. 2017. Website Fingerprinting Defenses at the Application Layer. Proc. Priv. Enhancing Technol. 2017, 2 (2017), 186–203. doi:10.1515/POPETS-2017-0023 [12] Giovanni Cherubin, Rob Jansen, and Carmela Troncoso. 2022. Online Website Fingerprinting: Evaluating Website Fingerprinting Attacks on Tor in the Real World. In USENIX Security Symposium. 753–770. https://www.usenix.org/confe rence/usenixsecurity22/presentation/cherubin [13] Cloudflare. 2026. Adoption & Usage Worldwide. https://radar.cloudflare.com/a doption-and-usage?dateRange=52w. Cloudflare Radar, accessed 2026-08-17. [14] Scott E. Coull and Kevin P. Dyer. 2014. Traffic Analysis of Encrypted Messaging Services: Apple iMessage and Beyond. Comput. Commun. Rev. 44, 5 (2014), 5–11. doi:10.1145/2677046.2677048 [15] Thomas M Cover. 1999. Elements of information theory. [16] Tianyu Cui, Gaopeng Gou, Gang Xiong, et al. 2021. SiamHAN: IPv6 Address Correlation Attacks on TLS Encrypted Traffic via Siamese Heterogeneous Graph Attention Network. In USENIX Security Symposium. 4329–4346. https://www.us enix.org/conference/usenixsecurity21/presentation/cui [17] Thilini Dahanayaka, Guillaume Jourjon, and Suranga Seneviratne. 2020. Understanding Traffic Fingerprinting CNNs. In IEEE LCN. 65–76. doi:10.1109/LCN486 67.2020.9314785 [18] Xinhao Deng, Qi Li, and Ke Xu. 2024. Robust and Reliable Early-Stage Website Fingerprinting Attacks via Spatial-Temporal Distribution Analysis. In ACM CCS. 1997–2011. doi:10.1145/3658644.3670272 [19] Xianwen Deng, Ruijie Zhao, Yanhao Wang, Mingwei Zhan, Zhi Xue, and Yijun Wang. 2025. Countmamba: A generalized website fingerprinting attack via coarse-grained representation and fine-grained prediction. In IEEE S&P. IEEE, 1419–1437. [20] Kevin P. Dyer, Scott E. Coull, Thomas Ristenpart, et al. 2012. Peek-a-Boo, I Still See You: Why Efficient Traffic Analysis Countermeasures Fail. In IEEE S&P. 332–346. doi:10.1109/SP.2012.28 [21] Roy T. Fielding, Yves Lafon, and Julian F. Reschke. 2014. Hypertext Transfer Protocol (HTTP/1.1): Range Requests. RFC 7233 (June 2014). doi:10.17487/RFC 7233 [22] Vincent Ghiëtte and Christian Doerr. 2020. Scaling website fingerprinting. In IFIP Networking. 199–207. https://ieeexplore.ieee.org/document/9142795 [23] Jiajun Gong and Tao Wang. 2020. Zero-delay Lightweight Defenses against Website Fingerprinting. In USENIX Security Symposium. 717–734. https://www. usenix.org/conference/usenixsecurity20/presentation/gong [24] Jamie Hayes and George Danezis. 2016. k-fingerprinting: A Robust Scalable Website Fingerprinting Technique. In USENIX Security Symposium. 1187–1203. https://www.usenix.org/conference/usenixsecurity16/technical-sessions/prese ntation/hayes [25] Andrew Hintz. 2002. Fingerprinting websites using traffic analysis. In PETS. Springer, 171–178. [26] Paul Hoffman and Patrick McManus. 2018. DNS Queries over HTTPS (DoH). RFC 8484 (Oct. 2018). doi:10.17487/RFC8484 [27] James K. Holland and Nicholas Hopper. 2022. RegulaTor: A Straightforward Website Fingerprinting Defense. Proc. Priv. Enhancing Technol. 2022, 2 (2022), 344–362. doi:10.2478/POPETS-2022-0049 [28] Jana Iyengar and Martin Thomson. 2021. QUIC: A UDP-Based Multiplexed and Secure Transport. RFC 9000 (May 2021). doi:10.17487/RFC9000 [29] Marc Juarez, Mohsen Imani, Mike Perry, et al. 2015. WTF-PAD: Toward an Efficient Website Fingerprinting Defense for Tor. CoRR abs/1512.00524 (2015). arXiv:1512.00524 http://arxiv.org/abs/1512.00524 [30] Wladimir De la Cadena, Asya Mitseva, Jens Hiller, et al. 2020. TrafficSliver: Fighting Website Fingerprinting Attacks with Traffic Splitting. In ACM CCS. 1971–1985. doi:10.1145/3372297.3423351

Coverage

CL-TAM FRONT H2PC

Client-Side [2.5 − 38.4] High [1.8 − 27.7] High [7.3 − 20.1] Low

Page-wide Page-wide Page-wide

SRV-TAM ALPaCA H2PS 1st P

Server-Side [1.7 − 24.5] High [2.3 − 76.3] High [13.2 − 94.7] Low

Per-server Per-server Page-wide

Conclusion

We show that HTTP/2 supports practical subpage fingerprinting defenses without requiring a multi-hop anonymity network. We demonstrate how established client- and server-side defenses can be emulated using HTTP/2 primitives and identify additional protocol features that improve their privacy–overhead trade-offs and deployment coverage. Further, we highlight overlooked HTTP/2 features that can significantly reduce client-side defense overhead or expand the coverage of server-side defenses. Table 9 summarizes these trade-offs. On the client side, CL-TAM achieves the strongest privacy in several datasets but with high overhead and substantial cross-dataset variability, whereas H2PC provides a higher minimum candidate set (7.31) at lower cost and with page-wide coverage. On the server side, SRV-TAM and ALPaCA provide strong protection when deployed on the appropriate server, but remain placement-dependent and costly. H2PS instead extends protection page-wide from the 1st -party server alone, maintaining a candidate set above 13 webpages at lower overhead and without requiring coordination with third-party servers. Finally, we introduced a blueprint for benchmarking and quality assurance, providing a structured way to assess fingerprinting defenses and their privacy–overhead trade-offs. Our evaluation shows that defense calibration and attacker hyperparameter tuning are integral to this assessment. More broadly, the variation in selected defense parameters and strongest attackers across datasets highlights the benefits of multi-dataset evaluation and suggests that defense configurations should be tailored to the target website, as conclusions drawn from other websites may not transfer reliably.

Acknowledgments We thank the anonymous reviewers for their constructive feedback, which helped improve the paper.

References [1] Ahmed Abusnaina, Rhongho Jang, Aminollah Khormali, et al. 2020. DFD: Adversarial Learning-based Approach to Defend Against Website Fingerprinting. In IEEE INFOCOM. 2459–2468. doi:10.1109/INFOCOM41043.2020.9155465 [2] Akamai Technologies. 2023. Improve UX with HTTP/2 Multiplexed Requests. https://www.akamai.com/blog/perf ormance/improve- ux- with- http2multiplexed-requests. Accessed: 2024-07-31. [3] Khaled Al-Naami, Amir El-Ghamry, Md Shihabul Islam, et al. 2021. BiMorphing: A Bi-Directional Bursting Defense against Website Fingerprinting Attacks. IEEE

17

Bogdan Cebere, Prateek Kumar, Sylvain Chatel, Wouter Lueks, and Christian Rossow

[31] Shuai Li, Huajun Guo, and Nicholas Hopper. 2018. Measuring Information Leakage in Website Fingerprinting Attacks and Defenses. In ACM CCS. 1977– 1992. doi:10.1145/3243734.3243832 [32] Jingyuan Liang, Chansu Yu, Kyoungwon Suh, et al. 2022. Tail Time Defense Against Website Fingerprinting Attacks. IEEE Access 10 (2022), 18516–18525. doi:10.1109/ACCESS.2022.3146236 [33] Weiran Lin, Sanjeev Reddy, and Nikita Borisov. 2019. Measuring the impact of HTTP/2 and server push on web fingerprinting. In MADWeb. [34] Xinjie Lin, Gang Xiong, Gaopeng Gou, Zhen Li, Junzheng Shi, and Jing Yu. 2022. ET-BERT: A Contextualized Datagram Representation with Pre-training Transformers for Encrypted Traffic Classification. In Proceedings of the ACM Web Conference 2022. Association for Computing Machinery, New York, NY, USA, 633–642. doi:10.1145/3485447.3512217 [35] Chang Liu, Zigang Cao, Zhen Li, and Gang Xiong. 2018. LaFFT: Length-Aware FFT Based Fingerprinting for Encrypted Network Traffic Classification. In IEEE ISCC. 1–6. doi:10.1109/ISCC.2018.8538732 [36] Chang Liu, Longtao He, Gang Xiong, et al. 2019. FS-Net: A Flow Sequence Network For Encrypted Traffic Classification. In IEEE INFOCOM. 1171–1179. doi:10.1109/INFOCOM.2019.8737507 [37] David Lu, Sanjit Bhat, Albert Kwon, et al. 2018. DynaFlow: An Efficient Website Fingerprinting Defense Based on Dynamically-Adjusting Flows. In ACM WPES. 109–113. doi:10.1145/3267323.3268960 [38] Xiapu Luo, Peng Zhou, Edmond W. W. Chan, et al. 2011. HTTPOS: Sealing Information Leaks with Browser-side Obfuscation of Encrypted Flows. In NDSS Symposium. https://www.ndss-symposium.org/ndss2011/httpos-sealinginformation-leaks-with-browser-side-obfuscation-of-encrypted-flows [39] Mariano Di Martino, Peter Quax, and Wim Lamotte. 2019. Realistically Fingerprinting Social Media Webpages in HTTPS Traffic. In ARES. 54:1–54:10. doi:10.1145/3339252.3341478 [40] Mariano Di Martino, Pieter Robyns, Peter Quax, et al. 2018. IUPTIS: A Practical, Cache-resistant Fingerprinting Technique for Dynamic Webpages. In WEBIST. 102–112. doi:10.5220/0007226501020112 [41] Brad Miller, Ling Huang, Anthony D. Joseph, et al. 2014. I Know Why You Went to the Clinic: Risks and Realization of HTTPS Traffic Analysis. In PETS (Lecture Notes in Computer Science, Vol. 8555). 143–163. doi:10.1007/978-3-319-08506-7_8 [42] Gargi Mitra, Prasanna Karthik Vairam, Patanjali Slpsk, et al. 2020. Depending on HTTP/2 for Privacy? Good Luck!. In IEEE/IFIP DSN. doi:10.1109/DSN48063.2 020.00044 [43] Ricardo Morla. 2017. Effect of Pipelining and Multiplexing in Estimating HTTP/2.0 Web Object Sizes. CoRR abs/1707.00641 (2017). arXiv:1707.00641 http://arxiv.org/abs/1707.00641 [44] Milad Nasr, Alireza Bahramali, and Amir Houmansadr. 2021. Defeating DNNBased Traffic Analysis Systems in Real-Time With Blind Adversarial Perturbations. In USENIX Security Symposium. 2705–2722. https://www.usenix.org/con ference/usenixsecurity21/presentation/nasr [45] Kazuho Oku. 2017. An HTTP Status Code for Indicating Hints. RFC 8297 (2017), 1–7. doi:10.17487/RFC8297 [46] Andriy Panchenko, Fabian Lanze, Jan Pennekamp, et al. 2016. Website Fingerprinting at Internet Scale. In NDSS Symposium. http://wp.internetsociety.org /ndss/wp-content/uploads/sites/25/2017/09/website-fingerprinting-internetscale.pdf [47] Performance Calendar. 2022. HTTP/3 Prioritization Demystified. https://ca lendar.perfplanet.com/2022/http-3-prioritization-demystif ied/. Accessed: 2024-07-30. [48] Victor Le Pochat, Tom Van Goethem, Samaneh Tajalizadehkhoob, Maciej Korczyński, and Wouter Joosen. 2018. Tranco: A research-oriented top sites ranking hardened against manipulation. arXiv preprint arXiv:1806.01156 (2018). [49] Tobias Pulls. 2020. Towards Effective and Efficient Padding Machines for Tor. CoRR abs/2011.13471 (2020). arXiv:2011.13471 https://arxiv.org/abs/2011.13471 [50] Mohammad Saidur Rahman, Mohsen Imani, Nate Mathews, et al. 2021. Mockingbird: Defending Against Deep-Learning-Based Website Fingerprinting Attacks With Adversarial Traces. IEEE Trans. Inf. Forensics Secur. 16 (2021), 1594–1609. doi:10.1109/TIFS.2020.3039691 [51] Mohammad Saidur Rahman, Payap Sirinam, Nate Mathews, et al. 2020. TikTok: The Utility of Packet Timing in Website Fingerprinting Attacks. Proc. Priv. Enhancing Technol. 2020, 3 (2020), 5–24. doi:10.2478/POPETS-2020-0043 [52] Eric Rescorla, Kazuho Oku, Nick Sullivan, and Christopher A. Wood. 2026. TLS Encrypted Client Hello. RFC 9849 (March 2026). doi:10.17487/RFC9849 [53] Vera Rimmer, Davy Preuveneers, Marc Juarez, et al. 2018. Automated Website Fingerprinting through Deep Learning. In NDSS Symposium. doi:10.14722/ndss. 2018.23105 [54] Brian C Ross. 2014. Mutual information between discrete and continuous data sets. PloS one 9, 2 (2014), e87357. [55] Amir Sabzi, Rut Vora, Swati Goswami, et al. 2024. NetShaper: A Differentially Private Network Side-Channel Mitigation System. In USENIX Security Symposium. https://www.usenix.org/conference/usenixsecurity24/presentation/sabzi [56] Shawn Shan, Arjun Nitin Bhagoji, Haitao Zheng, et al. 2021. A Real-time Defense against Website Fingerprinting Attacks. CoRR abs/2102.04291 (2021).

arXiv:2102.04291 https://arxiv.org/abs/2102.04291 [57] Meng Shen, Zhenbo Gao, Liehuang Zhu, et al. 2021. Efficient Fine-Grained Website Fingerprinting via Encrypted Traffic Analysis with Deep Learning. In IEEE/ACM IWQoS. 1–10. doi:10.1109/IWQOS52092.2021.9521272 [58] Meng Shen, Kexin Ji, Zhenbo Gao, et al. 2023. Subverting Website Fingerprinting Defenses with Robust Traffic Representation. In USENIX Security Symposium. 607– 624. https://www.usenix.org/conference/usenixsecurity23/presentation/shenmeng [59] Meng Shen, Yiting Liu, Siqi Chen, Liehuang Zhu, and Yuchao Zhang. 2019. Webpage Fingerprinting using Only Packet Length Information. In IEEE ICC. 1–6. doi:10.1109/ICC.2019.8761167 [60] Meng Shen, Yiting Liu, Liehuang Zhu, et al. 2021. Fine-Grained Webpage Fingerprinting Using Only Packet Length Information of Encrypted Traffic. IEEE Trans. Inf. Forensics Secur. 16 (2021), 2046–2059. doi:10.1109/TIFS.2020.3046876 [61] Sandra Siby, Ludovic Barman, Christopher A. Wood, et al. 2023. Evaluating practical QUIC website fingerprinting defenses for the masses. Proc. Priv. Enhancing Technol. 2023, 4 (2023), 79–95. doi:10.56553/POPETS-2023-0099 [62] Payap Sirinam, Mohsen Imani, Marc Juarez, et al. 2018. Deep Fingerprinting: Undermining Website Fingerprinting Defenses with Deep Learning. In ACM CCS. 1928–1943. doi:10.1145/3243734.3243768 [63] Jean-Pierre Smith, Luca Dolfi, Prateek Mittal, et al. 2022. QCSD: A QUIC ClientSide Website-Fingerprinting Defence Framework. In USENIX Security Symposium. 771–789. [64] Qixiang Sun, Daniel R Simon, Yi-Min Wang, Wilf Russell, Venkata N Padmanabhan, and Lili Qiu. 2002. Statistical identification of encrypted web browsing traffic. In IEEE S&P. IEEE, 19–30. [65] Martin Thomson and Cory Benfield. 2022. HTTP/2. RFC 9113 (2022), 1–78. doi:10.17487/RFC9113 [66] Thijs Van Ede, Riccardo Bortolameotti, Andrea Continella, et al. 2020. FlowPrint: Semi-Supervised Mobile-App Fingerprinting on Encrypted Network Traffic. In NDSS Symposium. doi:10.14722/ndss.2020.24412 [67] Alexander Veicht, Cédric Renggli, and Diogo Barradas. 2023. DeepSE-WF: Unified Security Estimation for Website Fingerprinting Defenses. Proc. Priv. Enhancing Technol. 2023, 2 (2023), 188–205. doi:10.56553/POPETS-2023-0047 [68] Kailong Wang, Junzhe Zhang, Guangdong Bai, et al. 2021. It’s Not Just the Site, It’s the Contents: Intra-domain Fingerprinting Social Media Websites Through CDN Bursts. In The Web Conference. 2142–2153. doi:10.1145/3442381.3450008 [69] Rong Wang, Zhen Ling, Guangchi Liu, Shaofeng Li, Junzhou Luo, and Xinwen Fu. 2026. Cease at the Ultimate Goodness: Towards Efficient Website Fingerprinting Defense via Iterative Mutual Information Minimization.. In NDSS. [70] Tao Wang. 2020. High Precision Open-World Website Fingerprinting. In IEEE S&P. 152–167. doi:10.1109/SP40000.2020.00015 [71] Tao Wang and Ian Goldberg. 2017. Walkie-Talkie: An Efficient Defense Against Passive Website Fingerprinting Attacks. In USENIX Security Symposium. 1375– 1390. https://www.usenix.org/conference/usenixsecurity17/technicalsessions/presentation/wang-tao [72] Wikipedia contributors. [n. d.]. HTTP/2 Server Push. https://en.wikipedia.org/w iki/HTTP/2_Server_Push. Accessed: 2026-04-23. [73] Renjie Xie, Jiahao Cao, Enhuan Dong, et al. 2023. Rosetta: Enabling Robust TLS Encrypted Traffic Classification in Diverse Network Environments with TCP-Aware Traffic Augmentation. In USENIX Security Symposium. 625–642. https://www.usenix.org/conference/usenixsecurity23/presentation/xie [74] Ziqing Zhang, Cuicui Kang, Gang Xiong, et al. 2019. Deep Forest with LRRS Feature for Fine-grained Website Fingerprinting with Encrypted SSL/TLS. In ACM CIKM. 851–860. doi:10.1145/3357384.3357993 [75] Xiyuan Zhao, Xinhao Deng, Qi Li, Yunpeng Liu, Zhuotao Liu, Kun Sun, and Ke Xu. 2024. Towards fine-grained webpage fingerprinting at scale. In ACM CCS. 423–436.

A

Open Science

Code availability. We release the code for defense calibration, the WF defenses auditing tool (wfaudit), real-world data collection and replay, the HTTP/2 client- and server-side defenses, and the dataset creation and benchmarking pipeline at https://github.com/bcebe re/Understanding-the-Privacy-Preserving-Potential-of-HTTP2Against-Webpage-Fingerprinting.

B

Ethical Considerations

Data Collection. All data used in this work was collected from publicly accessible websites using automated browsing without bypassing authentication, paywalls, or access controls. No personal data, user accounts, or sensitive identifiers were collected; only 18

Understanding the Privacy-Preserving Potential of HTTP/2 Against Webpage Fingerprinting

D.1

client–server traffic from scripted, non-authenticated sessions was recorded. For each case study, we saved the request order and the downloaded resources to replay the content locally under various client or server defenses. Stakeholders and Impacts. This work involves three primary stakeholder groups, as follows: • End Users. Common end users, including privacy-sensitive populations (such as journalists, activists, and users of privacyenhancing technologies), are the primary stakeholders. Our experiments did not involve real users or user-generated traffic. While publication of fingerprinting techniques may increase the risk of traffic analysis by adversaries, the defensive insights provided by this work aim to strengthen user privacy in the long term. • HTTP/2 Clients or Servers Developers. Developers of HTTP/2 libraries may be impacted by our findings. This work identifies potential privacy weaknesses and benefits when employing various HTTP/2 features. • Researchers and Practitioners. The research community benefits from reproducible measurements and benchmarks. To support responsible reuse, we release our code for data collection and the benchmarking framework (Section A). The authors bear responsibility for accurate threat modeling and harm mitigation. Research Impact. This work has both positive and negative impacts, affecting the stakeholders. • Positive Impacts. First, we emulate concrete defenses for strengthening HTTP/2 clients and servers. Second, our open-source implementation enables practitioners to audit, reproduce, and extend our findings (Section A). • Negative Impacts. As with prior fingerprinting research, the techniques discussed could be misused for surveillance, targeted advertising, or aggressive monitoring. These risks reflect the dual-use nature of traffic analysis research. Mitigations. To mitigate potential harms, we implement and evaluate defenses alongside attacks. Specifically, we implement defensive mechanisms at both the client and the server levels, using various HTTP/2 features.

C

D.2

HTTP/2 Client - Server Simulation

The client-server code used to replay the browser traces is available in the code repository, with a proof-of-concept standalone library ‘h2deflib‘. The HTTP/2 clients and server (and their defenses) are implemented using the h2 Python library, version 4.1.0. The h2 library is a pure-Python, fully compliant implementation (RFC 9113 [65]) of the HTTP/2 protocol, that provides the low-level building blocks necessary to implement HTTP/2 clients and servers without requiring a specific I/O framework.

D.3

PCAP Traces Capture

To create the datasets for each benchmark, we replay the previously captured browser traces using the client–server setup described in Section D.2. The client and server run in separate Docker containers connected through a Docker virtual network; traffic is therefore exchanged between the two container network namespaces rather than over the host loopback interface. Concretely, the client_runner.py script automates the capture of network traffic during replay. At startup, it loads the pre-captured browser data (Section D.1), including the requests and response bodies associated with each webpage. For each (webpage, repeat) pair, the script starts a Scapy AsyncSniffer on the client-side network interface and replays the webpage’s HTTP/2 requests against the server container. Requests are grouped by connection (i.e., by domain), and the configured client-side and server-side defenses are applied during replay. Once the replay completes, the sniffer is stopped, and the captured packets are stored in a PCAP file for subsequent processing by the PCAP parsing pipeline.

Generative AI Usage

Large Language Models were used for limited editorial assistance, related-work search support, and code review/debugging. All generated text was reviewed and edited by the authors, all references were independently verified, and all research code was written, verified, and validated by the authors. No LLM was used to generate original research ideas or results forming the scientific contributions of this paper. The authors take full responsibility for the accuracy, originality, and integrity of the work.

D

Baseline Browser Crawlers

The browser traces are collected using a headless Chromium browser automated via Playwright. For each URL, a new browser context is created with caching disabled, and the page is loaded until the domcontentloaded event fires. Two traces are recorded in parallel: a client-side trace logging each request’s URL and headers, and a server-side trace logging each response’s URL, status code, content type, response headers, and round-trip duration (measured as wallclock time between request dispatch and response receipt). Both client and server browser traces are serialized to JSON and saved to disk. The raw response body of each resource is saved to disk and referenced by path in the server trace JSON. These traces (request order and response content) are replayed in a Python HTTP/2 client - server environment, with various defenses enabled. The code responsible for collecting the browser traces is available in “datasets/browser_crawlers” in the code repository.

D.4

Evaluation Datasets Creation

The PCAP traces are parsed and converted into two distinct evaluation datasets: a 2D representation and a 3D representation. For 2D datasets (used by k-FP and WeFDE), each trace is loaded as a two-column CSV of timestamps and signed packet sizes, from which a flat 1D feature vector is extracted per trace covering packet counts, inter-packet timing, burst statistics, and CUMUL features. Stacking these vectors across all traces yields a 2D feature matrix of shape (n_traces, LIMconns ∗ n_features).

Data Collection Methodology

Given that we benchmark defenses at both the server and the client levels, we cannot use the real website deployments for experiments. Instead, we collect browser traces for each webpage and replay them using the HTTP/2 client and server described in Section D.2. 19

Bogdan Cebere, Prateek Kumar, Sylvain Chatel, Wouter Lueks, and Christian Rossow

10−5, 10−4, 10−3, 10−2 }; classifier dropout over [0.3, 0.7]; and batch size over {64, 128, 200, 256}.

For 3D datasets (used by DF, VarCNN, RobustFP, Holmes and DeepSE-WF), each trace is converted into two channels: a signed timing channel (sign × timestamp) and a signed size channel (sign × |size|/2000), each padded or truncated to feature_length (set adaptively as the median non-zero trace length plus 50, capped at the specified maximum 5, 000). The two channels are stacked along axis 0 to give a per-trace tensor of shape (2, feature_length). Collecting all traces yields a 3D tensor of shape (n_samples, 2, feature_length), which is then standardized per channel across samples and time positions using the global mean and standard deviation. The code responsible for parsing the traces and creating the datasets is available in the code repository, in “wfaudit/src/wfaudit/parser.py”.

E

E.1.3 VarCNN [5]. VarCNN is a data-efficient website-fingerprinting attack that uses a ResNet-based CNN, which leads to better performance with fewer training traces (compared to DF). The original implementation is available at github.com/sanjit-bhat/Var-CNN. The VarCNN classifier is built on a ResNet-18 backbone with four stages of [2, 2, 2, 2] residual blocks, each consisting of two Conv1d layers of kernel size 3 with batch normalization (𝜀 = 10−5 ) and ReLU activations. The input embedding uses a Conv1d layer of kernel size 7 and stride 2, followed by batch normalization, ReLU, and max pooling (kernel size 3, stride 2). Channel depths double across stages: 64 → 128 → 256 → 512. Global average pooling is applied after the final stage, followed by a two-layer embedding head that projects 512 → 1024 → embedding_size (= 512) with ReLU and dropout (𝑝 = 0.1), and a classification head with ReLU, dropout (𝑝 = 0.1), and a linear output layer. For hyperparameter search, we tune the learning rate, weight decay, dropout, and batch size. The learning rate is searched logarithmically over [10−4, 5 × 10−3 ]; weight decay over {0, 10−6, 10−5, 10−4, 10−3, 10−2 }; dropout over [0.1, 0.5]; and batch size over {64, 128, 200, 256}.

Security Estimators Details

We employ two categories of security estimators: (1) fingerprinting classifiers that measure defense effectiveness through classification performance, and (2) information-theoretic estimators that quantify residual information leakage independently of specific attack strategies.

E.1

The Machine-Learning Estimators

E.1.1 K-Fingerprinting [24]. k-FP extracts hand-crafted features from labeled network traces and trains a random decision forest classifier. Each trace is then represented by the vector of leaf-node identifiers. This compact fingerprint is then matched against known fingerprints using a nearest-neighbor approach to identify the visited site. For hyperparameter search, we tune the number of trees in the random forest and the number of nearest neighbors used for classification. We search 𝑛 trees ∈ {50, 100, . . . , 500} and 𝑘 ∈ {2, . . . , 15}. These parameters control, respectively, the complexity of the learned fingerprint representation and the resolution of the neighbor-based voting stage.

E.1.4 Holmes [18]. Holmes is a method focusing on early-stage website fingerprinting. The original implementation is available at github.com/Xinhao-Deng/Website-Fingerprinting-Library. The Holmes classifier uses a four-stage convolutional encoder (conv_num_layers = 4), where each stage consists of a residual ConvBlock1d with two Conv1d layers of kernel size 3, same padding, batch normalization, and ReLU activations, plus a 1×1 projection shortcut when channel dimensions change. Between stages, max pooling (kernel size 3) and dropout (𝑝 = 0.3) are applied. Channel depths follow a doubling schedule capped at the embedding size: 128 → 128 → 128 → 128 (since emb_size = 128). Global average pooling is applied after the final stage, followed by a classification head with dropout (𝑝 = 0.3) and a linear output layer. For hyperparameter search, we tune the learning rate, weight decay, dropout, and batch size using the same ranges as VarCNN: learning rate in [10−4, 5×10−3 ] on a logarithmic scale, weight decay in {0, 10−6, 10−5, 10−4, 10−3, 10−2 }, dropout in [0.1, 0.5], and batch size in {64, 128, 200, 256}.

E.1.2 Deep-FP [62]. DF uses a deep convolutional neural network to learn features directly from raw traces, without relying on handcrafted features (unlike K-FP). The original implementation is available at github.com/deep-fingerprinting/df. The DF neural network architecture consists of four convolutional blocks, each with two Conv1d layers of kernel size 5 and same padding, ELU activations (𝛼 = 1.0), and a dropout rate of 0.1. The channel depths double across blocks: 32 → 64 → 128 → 256. A global average pooling layer is followed by a two-layer embedding head that projects to embedding_size = 512 dimensions via a ReLU activation and dropout (𝑝 = 0.1), and a classification head with dropout (𝑝 = 0.5). For hyperparameter search, we tune the learning rate, weight decay, classifier dropout, and batch size. The learning rate is searched logarithmically over [10−4, 5 × 10−3 ]; weight decay over {0, 10−6,

E.1.5 RobustFP-CNN [58]. Robust Fingerprinting combines a dedicated traffic representation with a CNN classifier. In our evaluation, we use its CNN architecture with our packet representation rather than reproducing the complete Robust Fingerprinting preprocessing pipeline; we therefore refer to this classifier as RobustFP-CNN. The original Robust Fingerprinting implementation is available at github.com/robust-fingerprinting/RF. The RobustFP-CNN classifier processes input through two sequential stages. The first is a 2D convolutional frontend with two pairs of Conv2d layers (kernel size (3, 6), same padding), channel depths 1 → 32 → 64, max pooling and dropout (𝑝 = 0.1) between pairs. The second is a 1D convolutional backend with channel configuration [128, 128, M, 256, 256, M, 512, 𝑛 classes ], where M denotes max pooling with dropout (𝑝 = 0.3), and each layer uses kernel size 3. Global average pooling produces the final logits.

We first describe the fingerprinting classifiers used in this study, together with their hyperparameter search spaces. To account for distribution shifts introduced by the defenses, we tune classifier hyperparameters independently for each dataset and defense configuration.

20

Understanding the Privacy-Preserving Potential of HTTP/2 Against Webpage Fingerprinting

For hyperparameter search, we tune the learning rate, weight decay, convolutional dropout, and batch size. The learning rate is searched logarithmically over [10−4, 5×10−3 ]; weight decay over {0, 10−6, 10−5, 10−4, 10−3, 10−2 }; convolutional dropout over [0.1, 0.5]; and batch size over {64, 128, 200, 256}. All CNN-based methods are modified to also process the packetlength information, which is informative in our datasets (and constant per Tor cell in the original implementations).

E.2

Information Leakage Estimators

E.2.1 WeFDE [31]. estimates the mutual information using manually selected features. The original implementation is available at github.com/s0irrlor7m/InfoLeakWebsiteFingerprint, and a Python version is available at github.com/notem/reWeFDE. The WeFDE information leakage estimator uses the following hyperparameters: Bandwidth selection in the KDE uses the Hall plug-in method, with a rule-of-thumb fallback; Individual feature leakages are estimated using n_samples = 5,000 Monte Carlo samples, while the final joint cluster leakage uses n_samples = 50,000. The top topn = 20 features by individual leakage are selected for joint analysis, after pruning redundant features whose normalized mutual information exceeds nmi_threshold = 0.7. E.2.2 DeepSE-WF [67]. estimates the mutual information and the Bayes error by using specialized kNN-based estimators on learned latent feature spaces. The original implementation is available at github.com/veichta/DeepSE-WF. DeepSE-WF estimates mutual information by training an embedding model and applying 𝑘-NN estimators. The embedding backbone is DF with embedding_size = 512, dropout = 0.1, trained with batch_size = 200 and input sequences of length = 5,000. The dataset is split into train, validation, and two held-out test sets via stratified 𝑘-fold cross-validation (k_fold = 5); in each fold, the embedding model is trained on the training split and the two test splits are used to compute pairwise 𝑘-NN distance matrices. MI is then estimated in both directions (test1→test2 and test2→test1) and averaged. The 𝑘-NN estimator uses squared_l2 distance, with mi_k = 5 neighbours for mutual information. All these models are available in the “wfaudit” folder in the repository.

21

Bogdan Cebere, Prateek Kumar, Sylvain Chatel, Wouter Lueks, and Christian Rossow

Table 11: Hyperparameter tuning on the BBC Dataset. Abbreviations: batch size → bs; dropout → do; learning rate → lr; weight decay → wd.

Table 14: Hyperparameter tuning on the Wikipedia Dataset. Abbreviations: batch size → bs; dropout → do; learning rate → lr; weight decay → wd.

Defense

Best Model

Selected parameters

Defense

Best Model

Selected parameters

Client-side HTTPOS LLaMA FRONT CL-Tamaraw H2PC

k-FP RobustFP-CNN RobustFP-CNN Holmes RobustFP-CNN

trees 350, k 9 bs 64, do .448, lr 9.2e-4, wd 0 bs 64, do .457, lr 8.6e-4, wd 0 bs 256, do .147, lr 4.6e-3, wd 1.0e-6 bs 64, do .457, lr 8.6e-4, wd 0

Client-side HTTPOS LLaMA FRONT CL-Tamaraw H2PC

RobustFP-CNN RobustFP-CNN Holmes Holmes RobustFP-CNN

bs 64, do .448, lr 9.2e-4, wd 0 bs 64, do .457, lr 8.6e-4, wd 0 bs 256, do .147, lr 4.6e-3, wd 1.0e-6 bs 256, do .147, lr 4.6e-3, wd 1.0e-6 bs 64, do .448, lr 9.2e-4, wd 0

Server-side SRV-ALPaCA (1st) SRV-ALPaCA (3rd) SRV-ALPaCA (all) SRV-Tamaraw (1st) SRV-Tamaraw (3rd) SRV-Tamaraw (all) H2PS (1st)

RobustFP-CNN RobustFP-CNN RobustFP-CNN Holmes Holmes Holmes Holmes

bs 64, do .457, lr 8.6e-4, wd 0 bs 64, do .448, lr 9.2e-4, wd 0 bs 64, do .448, lr 9.2e-4, wd 0 bs 256, do .147, lr 4.6e-3, wd 1.0e-6 bs 128, do .124, lr 1.1e-3, wd 1.0e-6 bs 256, do .147, lr 4.6e-3, wd 1.0e-6 bs 256, do .147, lr 4.6e-3, wd 1.0e-6

Server-side SRV-ALPaCA (1st) k-FP trees 50, k 3 SRV-ALPaCA (3rd) RobustFP-CNN bs 64, do .457, lr 8.6e-4, wd 0 SRV-ALPaCA (all) RobustFP-CNN defaults kept SRV-Tamaraw (1st) RobustFP-CNN bs 64, do .457, lr 8.6e-4, wd 0 SRV-Tamaraw (3rd) RobustFP-CNN bs 64, do .448, lr 9.2e-4, wd 0 SRV-Tamaraw (all) RobustFP-CNN bs 64, do .448, lr 9.2e-4, wd 0

Table 12: Hyperparameter tuning on the Reddit Dataset. Abbreviations: batch size → bs; dropout → do; learning rate → lr; weight decay → wd.

Table 10: Hyperparameter tuning on the Amazon Dataset. Abbreviations: batch size → bs; dropout → do; learning rate → lr; weight decay → wd.

Defense

Defense

Best Model

Selected parameters

RobustFP-CNN bs 64, do .457, lr 8.6e-4, wd 0 Holmes bs 128, do .124, lr 1.1e-3, wd 1.0e-6 RobustFP-CNN bs 64, do .457, lr 8.6e-4, wd 0 Holmes bs 256, do .147, lr 4.6e-3, wd 1.0e-6 RobustFP-CNN bs 64, do .457, lr 8.6e-4, wd 0

Client-side HTTPOS LLaMA FRONT CL-Tamaraw H2PC

RobustFP-CNN Holmes Holmes Holmes Holmes

bs 64, do .448, lr 9.2e-4, wd 0 bs 256, do .147, lr 4.6e-3, wd 1.0e-6 bs 256, do .147, lr 4.6e-3, wd 1.0e-6 bs 256, do .147, lr 4.6e-3, wd 1.0e-6 bs 256, do .147, lr 4.6e-3, wd 1.0e-6

RobustFP-CNN Holmes RobustFP-CNN Holmes Holmes Holmes

Server-side SRV-ALPaCA (1st) RobustFP-CNN SRV-ALPaCA (3rd) Holmes SRV-ALPaCA (all) Holmes SRV-Tamaraw (1st) RobustFP-CNN SRV-Tamaraw (3rd) Holmes SRV-Tamaraw (all) Holmes

bs 64, do .457, lr 8.6e-4, wd 0 bs 256, do .147, lr 4.6e-3, wd 1.0e-6 bs 256, do .147, lr 4.6e-3, wd 1.0e-6 bs 64, do .165, lr 3.4e-4, wd 1.0e-4 bs 256, do .147, lr 4.6e-3, wd 1.0e-6 bs 128, do .124, lr 1.1e-3, wd 1.0e-6

Client-side HTTPOS LLaMA FRONT CL-Tamaraw H2PC Server-side SRV-ALPaCA (1st) SRV-ALPaCA (3rd) SRV-ALPaCA (all) SRV-Tamaraw (1st) SRV-Tamaraw (3rd) SRV-Tamaraw (all)

Best Model

Selected parameters

bs 64, do .448, lr 9.2e-4, wd 0 bs 128, do .124, lr 1.1e-3, wd 1e-6 defaults kept bs 128, do .124, lr 1.1e-3, wd 1e-6 bs 128, do .124, lr 1.1e-3, wd 1e-6 bs 128, do .124, lr 1.1e-3, wd 1e-6

Table 13: Hyperparameter tuning on the Udemy Dataset. Abbreviations: batch size → bs; dropout → do; learning rate → lr; weight decay → wd. Defense

Best Model

Selected parameters

Client-side HTTPOS FRONT LLaMA CL-Tamaraw H2PC

Holmes Holmes Holmes Holmes RobustFP-CNN

bs 128, do .124, lr 1.1e-3, wd 1.0e-6 bs 128, do .124, lr 1.1e-3, wd 1.0e-6 bs 256, do .147, lr 4.6e-3, wd 1.0e-6 bs 256, do .147, lr 4.6e-3, wd 1.0e-6 bs 64, do .448, lr 9.2e-4, wd 0

Server-side SRV-ALPaCA (1st) SRV-ALPaCA (3rd) SRV-ALPaCA (all) SRV-Tamaraw (1st) SRV-Tamaraw (3rd) SRV-Tamaraw (all)

RobustFP-CNN defaults kept RobustFP-CNN bs 64, do .448, lr 9.2e-4, wd 0 Holmes bs 256, do .148, lr 4.6e-3, wd 1e-6 Holmes bs 128, do .124, lr 1.1e-3, wd 1e-6 Holmes bs 128, do .124, lr 1.1e-3, wd 1e-6 k-FP 𝑛 trees 350, 𝑘 9

F

Security Estimators Hyperparameter Tuning

We implement hyperparameter optimization using Optuna. We tune each attacker independently for each dataset–defense pair using a separate Optuna study, with Macro-F1 on a stratified heldout validation split as the optimization objective. To reduce tuning cost, each search uses up to 150 traces per webpage, of which 20% are reserved for validation. Before optimization, we evaluate the attacker’s default configuration on the same tuning split. We retain the optimized parameters only when the best Optuna trial achieves a higher Macro-F1 than the default configuration; otherwise, we preserve the default parameters. We then retrain the selected configuration and evaluate it on the full dataset using the same benchmarking procedure as the remaining experiments. The following tables (Table 10, Table 11, Table 12, Table 13, Table 14) report, for each dataset–defense pair, the strongest attacker 22

Understanding the Privacy-Preserving Potential of HTTP/2 Against Webpage Fingerprinting

after tuning and the hyperparameters selected for that attacker. We

abbreviate batch size as bs, dropout as do, learning rate as lr, and weight decay as wd.

23

Record · ID 660743 · SHA-256 6be7bf2e11301461
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.