FLOWATOM: ATOM-BASED EVIDENCE AGGREGATION FOR MULTI-LABEL WEBSITE FINGERPRINTING Chongru Fan1,2 , Wentao Huang1 , Wei Wang2 , Zhenquan Ding2,∗ , Jinqiao Shi1,∗ , Wei Cai2 , Zhiyu Hao2
arXiv:2609.29330v1 [cs.LG] 24 Sep 2026
1
School of Cyberspace Security, Beijing University of Posts and Telecommunications, Beijing, China 2 Zhongguancun Laboratory, Beijing, China Website A
ABSTRACT Identifying the set of monitored websites in mixed encrypted traffic is challenging because an individual flow often provides only partial evidence of website identity. To address this challenge, we propose FlowAtom, which constructs shared prototypes, called Atoms, from flow representations without website labels. Specifically, FlowAtom pretrains a flow encoder on external unlabeled traffic and aggregates Atom responses across flows within each observation window into a fixed-dimensional, permutation-invariant representation for monitored website-set prediction. Across Direct HTTPS, Trojan, and VMess, FlowAtom achieves micro-F1 scores of 97.82%, 94.43%, and 93.92% in closedworld evaluation, respectively, and consistently outperforms the evaluated baselines in open-world evaluation on windows containing monitored visits. The code is available at https://github.com/aimafan123/FlowAtom. Index Terms— Website fingerprinting, Multi-instance multi-label learning, Encrypted traffic analysis 1. INTRODUCTION Website fingerprinting (WF) attacks infer the websites a user visits from side-channel features of encrypted traffic [1, 2, 3, 4]. Traditional WF studies typically use the complete traffic trace of a single website visit as one sample and formulate website identification as single-label classification [5, 6, 7]. In realistic browsing, however, multi-tab activity can generate overlapping visits to multiple websites, while background services and unmonitored websites contribute additional traffic [8, 9, 10]. These multi-website scenarios motivate website-set identification: inferring the set of websites represented in mixed traffic [11, 9, 12]. Packet-sequence-based multi-label WF methods, such as [9, 11], take predefined traffic traces as input and predict website labels from packet sequences. For approaches that require predefined visit traces, processing continuous ∗ Corresponding authors: Zhenquan Ding ([email protected]) and Jinqiao Shi ([email protected]). This work was supported by the National Major Science and Technology Project for Cyberspace Security under Grant 2025ZD1501502.
𝑓1𝐴 Shared Atom 1
Observation window
Aggregate evidence and decide
Atom x
Website A
𝑓3𝐴
Atom 3
A1
Website B
Atom 1
𝑓2𝐴
𝑓1𝐵 𝑓2𝐵
𝑓3𝐵
Atom x
A2
A3
Website B A1
A4
A5
Atom 2
Fig. 1. Different websites can generate similar flows; aggregating complementary evidence across flows in an observation window supports prediction of the monitored website set.
traffic depends on packet-level trace segmentation to locate those traces; visit boundaries can be difficult to recover accurately [13]. In the settings considered here, individual connections are distinguishable, allowing an attacker to separate flows using observable connection identifiers [14]. This groups packets by connection, without determining which website visit generated each flow. We therefore take a flowlevel perspective: the model receives an unordered flow set for each observation window and predicts the monitored website set without packet-level trace segmentation by visit or flow-to-website assignment at inference. Our evaluation constructs these windows offline from pre-segmented visit traces. As illustrated in Fig. 1, an individual flow often provides only partial evidence of website identity [14, 3]. A website visit typically generates multiple flows, while shared content delivery networks (CDNs), third-party services, and resources can produce similar flow characteristics across different websites [15, 16, 14]. Consequently, a flow may contain only part of the traffic generated by a website visit and may not identify the website on its own. Discriminative evidence may therefore be distributed across multiple flows within an observation window. We represent shared patterns in the flow representation space using prototypes called Atoms, and aggregate their responses to predict the monitored website set for an observation window.
Self-supervised Pretraining
Atom Construction
Augmentation Flow Pool Unlabeled Flows
Eq (query)
Proj. Head
Momentum Update
Aug. 1
Ek (key)
Atoms (clusters)
Atom IDs Atom 1
Contrastive Learning
Proj. Head
Embeddings
𝑓1
𝑓1 Ek (key)
𝑓2 𝑓3
...
k-Means
𝑓2 𝑓3
Atom 2 Atom 3
...
...
...
Aug. 2
𝑓1
XGBoost
q1 0.02, 0.03, 0.93, ...
𝑓1
𝑓2
Ek (key)
𝑓3
𝑓2 𝑓3
XGBoost
...
...
Window T (Unordered flow set)
Embeddings
q2 0.88, 0.01, 0.04, ...
q3 0.02, 0.84, 0.03, ...
Max Pooling
...
Window feature
MLP
Φ(𝑇) 0.88, 0.84, 0.93, ...
p(site1)
0.91
p(site2)
0.07
p(site3)
0.95
... Website labels
Atom responses
Window-level website set prediction Fig. 2. Overview of FlowAtom: flow representation pretraining, Atom construction, and window-level multi-label prediction. Building on this intuition, we propose FlowAtom, which constructs shared Atoms by clustering flow representations without using website labels and represents each flow by a vector of Atom response scores. For each observation window, FlowAtom applies max pooling to each Atom’s response scores across flows, producing a fixed-dimensional, permutation-invariant window representation for predicting the monitored website set [17]. Our contributions are as follows. First, we formulate multi-label WF from a flow-level perspective, using an unordered flow set as input without requiring packet-level trace segmentation by visit or flow-to-website assignment at inference. We construct an evaluation benchmark from offline mixtures of pre-segmented visit traces under Direct HTTPS and non-multiplexed encrypted proxies. Second, we propose FlowAtom, which constructs shared Atoms without websitelabel supervision and aggregates their response scores across flows into a fixed-dimensional, permutation-invariant window representation for multi-label prediction. Third, we evaluate FlowAtom under Direct HTTPS, Trojan [18], and VMess [19], obtaining closed-world micro-F1 scores of 97.82%, 94.43%, and 93.92% and open-world micro-F1 scores of 92.37%, 92.64%, and 89.85%, respectively. 2. THREAT MODEL We consider Direct HTTPS and non-multiplexed encrypted proxy settings in which the attacker can distinguish individual connections. The attacker observes traffic on the client egress link in the Direct HTTPS setting and on the client–proxy link in the encrypted proxy settings. This passive attacker uses five-tuples or equivalent connection identifiers to separate
flows and observes packet payload lengths and directions, but cannot decrypt or modify the traffic. Connection identifiers are used only for flow separation; the model input excludes IP addresses, DNS information, and Server Name Indication (SNI). For each observation window, the model receives an unordered flow set, without website-visit boundaries or flowto-website assignments. The attacker must still separate flows by connection; the model does not require packet-level trace segmentation by visit. 3. METHOD As shown in Fig. 2, FlowAtom comprises three stages: flow representation pretraining, Atom construction, and windowlevel multi-label prediction. During pretraining, FlowAtom removes zero-payload packets and represents each flow as a sequence of signed packet payload lengths, preserving packet order and using the sign to encode direction. Each sequence is truncated or padded to L = 300 packet positions. We use these sequences from large-scale external unlabeled traffic for MoCo-style contrastive pretraining [20]. For each input flow, we randomly select two distinct augmentation operations from [21] to generate two views for contrastive learning of flow representations. The encoder E maps each input flow fi to a representation zi = E(fi ). After pretraining, we discard the projection head and freeze the encoder E for Atom construction and all subsequent window representation computations. For each target traffic scenario, FlowAtom constructs an Atom space using only flows from that scenario’s training split, without using their website labels. We first extract training-flow representations using the frozen encoder
zi = E(fi ). We then apply k-means clustering [22] to these training-flow representations and discard clusters with fewer members than a minimum cluster-size threshold. The centers of the retained clusters form a set of shared prototypes, called Atoms: A = {a1 , a2 , . . . , aA }, where A is the number of retained clusters. We train an XGBoost mapper M [23] using the k-means cluster assignments as pseudo-labels. The mapper outputs an Atom response vector qi = M (zi ), where qi,a is the response score of flow fi for Atom a. These Atoms are learned without website labels, allowing flows from different website visits to be represented in a shared Atom space. The k-means algorithm is used only for offline Atom construction. For subsequent window representation computation and inference, both E and M remain frozen and are applied sequentially to compute Atom responses. To construct training and evaluation windows offline, we select one visit trace from each of one or more websites within the same data split and combine their valid flows into an unordered flow set T . In this multi-instance multi-label formulation [24], flows are instances, each window is a bag of flows, and the monitored websites present in the window form its label set. For each flow in T , we compute its Atom response vector using the frozen encoder E followed by the frozen mapper M . We then apply max pooling over flows for each Atom [25]: Φa (T ) = maxfi ∈T qi,a . This yields an A-dimensional window representation that is invariant to flow order, where A is the number of retained Atoms. We estimate the standardization parameters for window representations using only downstream training windows. We then train a multilayer perceptron (MLP) decoder with two hidden layers using window-level website labels and binary crossentropy with positive-class weighting. During decoder training, E and M remain frozen; the validation set is used for model and decision-threshold selection. At inference time, we apply flow encoding, Atom mapping, max pooling, standardization, and multi-label decoding in sequence to the input flow set to predict the monitored websites present in the window.
and selection. We compare FlowAtom with ARES [11], BAPM [10], TMWF [9], CAWF [14], and Flow-DF, a flow-based adaptation of DF [5]. ARES uses Transformer-based classifiers to identify websites from local patterns within traffic segments. BAPM combines convolution and attention mechanisms for website identification, whereas TMWF uses a Transformer. Flow-DF encodes each flow’s signed payloadlength sequence to predict website scores, then aggregates these scores across flows to obtain window-level predictions. CAWF predicts flow classes from per-flow statistical features and identifies websites using contextual relationships within the resulting flow-class sequences. All methods use the same data splits, constructed observation windows, and evaluation metrics. The baseline implementations are based on the corresponding published methods. For each traffic scenario, we split single-website visit traces into disjoint training, validation, and test sets. Within each split, we construct observation windows offline by combining the flows from complete, pre-segmented visit traces. Each window contains visits to m ∈ {1, 2, 3, 4, 5} distinct monitored websites, where m counts monitored websites only. FlowAtom receives only the mixed unordered flow set; website-visit boundaries, flow-to-website assignments, and the true number of monitored websites m are not provided. In closed-world testing, m = 1 denotes the single-website condition, and m ∈ {2, 3, 4, 5} denotes the multi-website condition. Within each traffic scenario, we evaluate the same trained model under both conditions. For open-world testing, we add the flows from b ∈ {1, 2, 3, 4, 5} unmonitored website-visit traces to each window containing monitored visits and evaluate the resulting performance in identifying the monitored website set. We use micro-averaged F1 (micro-F1) as the primary evaluation metric. For each evaluation task, we perform five runs with different random seeds and report the mean micro-F1 across the five runs. 4.2. Closed-World Experiments
4. EVALUATION 4.1. Evaluation Setup The experimental data comprise unlabeled traffic for pretraining, monitored website traffic, and unmonitored website traffic used as background traffic. For pretraining, we sample 1,000,000 unlabeled flows from JP-MAWI [26]. For monitored traffic, we randomly select 100 websites from the Tranco Top 10K [27] and collect website-visit traces under Direct HTTPS, Trojan, and VMess. For open-world evaluation, we exclude the 100 monitored websites from the same list and collect one visit trace for each remaining website, which we treat as unmonitored. Unmonitored website traffic is used only for testing and is excluded from model training
In the closed-world setting, we evaluate multi-label WF under single-website (m = 1) and multi-website (m ∈ {2, 3, 4, 5}) conditions. Table 1 shows that FlowAtom achieves the highest micro-F1 among the evaluated methods in all 15 closedworld settings, outperforming the strongest baseline in each setting by 4.64–16.70 percentage points. 4.3. Open-World Experiments Fig. 3 reports results for 30 combinations of three traffic scenarios, two monitored-website conditions (m = 1 and pooled m ∈ {2, 3, 4, 5}), and five values of b. All evaluated windows contain monitored visits, and FlowAtom achieves the highest micro-F1 among the evaluated methods in every combination.
Table 1. Closed-world micro-F1 (%) in three traffic scenarios, with m ∈ {1, 2, 3, 4, 5} distinct monitored websites per observation window. Bold and underlined values indicate the best and second-best results in each column, respectively. Direct HTTPS
Trojan
VMess
Method
1
2
3
4
5
1
2
3
4
5
1
2
3
4
5
ARES BAPM TMWF Flow-DF CAWF FlowAtom
67.91 57.57 68.85 75.92 84.72 99.35
57.14 34.21 35.72 80.20 86.59 98.76
40.49 25.30 29.39 78.82 86.06 96.71
37.41 21.05 25.45 78.66 81.00 97.70
35.77 15.66 20.22 81.72 84.19 96.78
88.20 64.17 79.42 82.29 82.41 95.57
73.37 34.33 42.38 79.62 84.75 96.26
62.42 33.61 36.84 77.05 81.87 95.00
57.07 32.39 36.55 81.30 80.80 92.37
53.73 23.73 30.00 82.07 80.52 93.08
89.89 65.26 82.83 82.56 80.85 94.53
74.74 37.11 46.91 82.49 79.85 93.83
63.73 31.71 39.67 82.21 77.39 94.58
61.79 30.95 39.98 84.23 81.79 93.66
53.82 24.12 32.53 80.10 73.62 92.86
FlowAtom
ARES
BAPM
Direct HTTPS
TMWF
Flow-DF
Trojan
CAWF
Direct HTTPS
1.0
VMess
0.8 Micro-F1
m=1
Trojan
1.0
VMess
0.5
0.6
0.4
1
2
4
8
16
All
1
2
4
8
16
All
1
2
4
8
16
All
Flow budget, k
Micro-F1
0.0
Fig. 5. Effect of the number of flows retained per visit trace, k, on micro-F1 under Direct HTTPS, Trojan, and VMess.
m = 2–5
1.0
0.5
0.0
1
2
3
4
5
1
2
3
4
5
1
2
3
4
5
Background traces, b
Fig. 3. Open-world micro-F1 as the number of unmonitored visit traces b increases. Results are reported separately for windows with one monitored website (m = 1) and pooled windows with multiple monitored websites (m ∈ {2, 3, 4, 5}). Direct HTTPS
Trojan
Fig. 4 shows that micro-F1 generally increases with K in all three traffic scenarios and then plateaus with small fluctuations at larger values. Fig. 5 shows that micro-F1 increases as more flows are retained per visit trace. When only one flow is retained per visit trace, micro-F1 is 32.16%, 37.31%, and 36.24% under Direct HTTPS, Trojan, and VMess, respectively. These results are consistent with website-identifying information being distributed across multiple flows and suggest that retaining more flows can improve window-level prediction.
VMess
Micro-F1 (%)
100 95
5. CONCLUSION
90 85 80 75 100
200
400
600
800
1000 1200 1500 1800 2000 3000 5000 K
Fig. 4. Effect of the parameter K on micro-F1 (%) under Direct HTTPS, Trojan, and VMess. 4.4. Ablation Experiments We separately vary the parameter K and the number of flows retained per visit trace under Direct HTTPS, Trojan, and VMess to examine how performance changes with each parameter.
We propose FlowAtom for multi-label WF in encrypted traffic with distinguishable connections. From a flow-level perspective, FlowAtom models each observation window as an unordered flow set and predicts the monitored website set without packet-level trace segmentation by visit or flowto-website assignment at inference. FlowAtom learns flow representations from external unlabeled traffic and constructs shared Atoms from the training flows of each target traffic scenario. Max pooling of each Atom’s responses across flows yields a fixed-dimensional, permutation-invariant window representation for predicting the monitored website set. Across Direct HTTPS, Trojan, and VMess, FlowAtom achieves the highest micro-F1 among the evaluated methods in all 15 closed-world and 30 open-world settings tested.
6. REFERENCES [1] T. Wang, X. Cai, R. Nithyanand, R. Johnson, and I. Goldberg, “Effective Attacks and Provable Defenses for Website Fingerprinting,” in Proc. USENIX Security Symp., 2014, pp. 143–157. [2] J. Hayes and G. Danezis, “k-fingerprinting: A Robust Scalable Website Fingerprinting Technique,” in Proc. USENIX Security Symp., 2016, pp. 1187–1203. [3] C. Li, L. Nie, L. Zhao, and K. Li, “Robust website fingerprinting through resource loading sequence,” World Wide Web, vol. 26, no. 5, pp. 2329–2349, 2023. [4] Yifei Cheng, Yujia Zhu, Baiyang Li, Xin-Hao Deng, Yitong Cai, Yaochen Ren, and Qingyun Liu, “Star: Semantic-traffic alignment and retrieval for zero-shot https website fingerprinting,” IEEE INFOCOM 2026 - IEEE Conference on Computer Communications, pp. 1–10, 2025. [5] P. Sirinam, M. Imani, M. Juarez, and M. Wright, “Deep Fingerprinting: Undermining Website Fingerprinting Defenses with Deep Learning,” in Proc. ACM CCS, 2018, pp. 1928–1943. [6] S. Bhat, D. Lu, A. Kwon, and S. Devadas, “Var-CNN: A Data-Efficient Website Fingerprinting Attack Based on Deep Learning,” Proc. Privacy Enhancing Technol., vol. 2019, no. 4, pp. 292–310, 2019. [7] M. S. Rahman, P. Sirinam, N. Mathews, K. G. Gangadhara, and M. Wright, “Tik-Tok: The Utility of Packet Timing in Website Fingerprinting Attacks,” Proc. Privacy Enhancing Technol., vol. 2020, no. 3, pp. 5–24, 2020. [8] M. Juarez, S. Afroz, G. Acar, C. Diaz, and R. Greenstadt, “A Critical Evaluation of Website Fingerprinting Attacks,” in Proc. ACM CCS, 2014, pp. 263–274. [9] Z. Jin, T. Lu, S. Luo, and J. Shang, “Transformer-based Model for Multi-tab Website Fingerprinting Attack,” in Proc. ACM CCS, 2023, pp. 1050–1064. [10] Z. Guan, G. Xiong, G. Gou, Z. Li, M. Cui, and C. Liu, “BAPM: Block Attention Profiling Model for Multi-tab Website Fingerprinting Attacks on Tor,” in Proc. ACSAC, 2021, pp. 248–259. [11] X. Deng, Q. Yin, Z. Liu, X. Zhao, Q. Li, M. Xu, K. Xu, and J. Wu, “Robust Multi-tab Website Fingerprinting Attacks in the Wild,” in Proc. IEEE Symp. Security Privacy, 2023, pp. 1005–1022. [12] W. Meng, C. Ma, M. Ding, C. Ge, Y. Qian, and T. Xiang, “Beyond Single Tabs: A Transformative Few-Shot Approach to Multi-Tab Website Fingerprinting Attacks,” in Proc. ACM Web Conf. (WWW), 2025, pp. 1068–1077. [13] Tao Wang and Ian Goldberg, “On realistically attacking tor with website fingerprinting,” Proc. Priv. Enhancing Technol., vol. 2016, no. 4, pp. 21–36, 2016. [14] X. Ma, J. Qu, M. Shi, B. An, J. Li, X. Luo, J. Zhang, Z. Li, and X. Guan, “Website Fingerprinting on En-
crypted Proxies: A Flow-Context-Aware Approach and Countermeasures,” IEEE/ACM Trans. Netw., vol. 32, no. 3, pp. 1904–1919, 2024. [15] Trinh Viet Doan, Roland van Rijswijk-Deij, Oliver Hohlfeld, and Vaibhav Bajpai, “An empirical view on consolidation of the web,” ACM Trans. Internet Technol., vol. 22, no. 3, Feb. 2022. [16] X. S. Wang, A. Balasubramanian, A. Krishnamurthy, and D. Wetherall, “Demystifying Page Load Performance with WProf,” in Proc. USENIX NSDI, 2013, pp. 473–485. [17] M. Zaheer, S. Kottur, S. Ravanbakhsh, B. Póczos, R. Salakhutdinov, and A. J. Smola, “Deep Sets,” in Adv. Neural Inf. Process. Syst. (NeurIPS), 2017. [18] trojan-gfw, “Trojan Documentation — trojangfw.github.io,” https://trojan-gfw.github. io/trojan/, 2023, Accessed: 2026-04-28. [19] Project X Community, “Project X Xray-core,” https: //xtls.github.io/, 2020, Accessed: 2026-0428. [20] K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick, “Momentum Contrast for Unsupervised Visual Representation Learning,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2020, pp. 9729–9738. [21] Renjie Xie, Jiahao Cao, Enhuan Dong, Mingwei Xu, Kun Sun, Qi Li, Licheng Shen, and Menghao Zhang, “Rosetta: Enabling robust TLS encrypted traffic classification in diverse network environments with TCPAware traffic augmentation,” in 32nd USENIX Security Symposium (USENIX Security 23), Anaheim, CA, Aug. 2023, pp. 625–642, USENIX Association. [22] S. Lloyd, “Least squares quantization in PCM,” IEEE Trans. Inf. Theory, vol. 28, no. 2, pp. 129–137, 1982. [23] T. Chen and C. Guestrin, “XGBoost: A Scalable Tree Boosting System,” in Proc. ACM SIGKDD Int. Conf. Knowl. Discovery Data Mining (KDD), 2016, pp. 785– 794. [24] Z.-H. Zhou and M.-L. Zhang, “Multi-Instance MultiLabel Learning with Application to Scene Classification,” in Adv. Neural Inf. Process. Syst. (NeurIPS), 2006. [25] C. R. Qi, H. Su, K. Mo, and L. J. Guibas, “PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), 2017, pp. 652–660. [26] K. Cho, K. Mitsuya, and A. Kato, “Traffic Data Repository at the WIDE Project,” in Proc. USENIX Annu. Tech. Conf., FREENIX Track, 2000, pp. 263–270. [27] Victor Le Pochat, Tom van Goethem, Samaneh Tajalizadehkhoob, Maciej Korczynski, and Wouter Joosen, “Tranco: A research-oriented top sites ranking hardened against manipulation,” in Proc. NDSS, 2019.