Conceptio › Archive › arXiv CS
arXiv CSopen access

Beyond Measurement Metrics: A Human-Centered Framework for Semantic Validation of Network Traffic Classification

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
distributed-systemsinternetnetworkingprotocols
networking, internet, protocols, distributed systems

Beyond Measurement Metrics: A Human-Centered Framework for Semantic Validation of Network Traffic Classification

arXiv:2609.17014v1 [cs.NI] 15 Sep 2026

Igor Cherepanov 1*, David Sessler 1 , Alex Ulmer Thorsten May 1 , Jörn Kohlhammer 1,2

1

,

1

2

Fraunhofer IGD, Darmstadt, 64283, Germany. Technische Universität Darmstadt, Darmstadt, 64289, Germany.

*Corresponding author(s). E-mail(s): [email protected]; Abstract Machine learning (ML) has become the dominant approach for network traffic classification, achieving very high predictive performance. However, a model is only valuable if it learns semantically meaningful and trustworthy patterns rather than exploiting spurious correlations. Conventional evaluation practices predominantly assess predictive performance. Consequently, whether the model relies on semantically meaningful patterns remains unknown. To address these challenges, we adapt the knowledge generation framework for network traffic classification. The adapted framework combines data, ML models, explainability, visualization, and expert reasoning to support the iterative exploration, verification, and refinement of model behavior and data preprocessing. The framework is grounded in findings from the literature, benchmark dataset analyses, practical experience with XAI-based traffic classification, and expert feedback, providing practical guidance for semantic model validation. By complementing predictive performance with semantic validation and human expertise, the proposed framework supports the development of network traffic classification models that are not only accurate but also robust and trustworthy. Keywords: Human-Centered AI, Network Traffic Classification, XAI, Semantic Model Validation

1

1 Introduction Network traffic classification is essential for identifying the applications and services generating network traffic, enabling effective network management, security monitoring, anomaly detection, and quality-of-service provisioning. Over time, traffic classification techniques have evolved from simple port-based approaches, which identify applications based on well-known transport-layer port numbers, and statistical methods to increasingly sophisticated ML and deep learning (DL) models [35, 41]. Recent advances in these methods have significantly improved predictive performance, leading research efforts to focus primarily on the development of more accurate classification models. While existing evaluation practices provide insights into how well a model performs, they provide limited support for understanding why a model performs well. Consequently, models can achieve high predictive accuracy, while they may rely on dataset-specific shortcut features, spurious correlations, protocol-specific artifacts, dataset biases, or other unintended patterns unrelated to the underlying classification task [17]. Such dependencies often remain hidden during development and may lead to substantial performance degradation when models are deployed in evolving network environments, where traffic characteristics and data distributions continuously change [28, 36]. Current traffic classification pipelines predominantly assess predictive performance. They provide little insight into what the model has actually learned. Semantic validation addresses this limitation by determining whether the learned decision strategy is based on meaningful traffic characteristics. At the same time, research fields such as explainable artificial intelligence (XAI), human-centered artificial intelligence (HCAI), and visual analytics (VA) increasingly emphasize the importance of understanding, validating, and contextualizing model behavior through human expertise [3, 7, 14, 21]. Despite these advances, systematic approaches for integrating explanation-based validation and expert reasoning into network traffic classification remain largely unexplored. Motivated by this gap, this work builds upon the knowledge gained from our previous research on explainable traffic classification, including the design, implementation, and evaluation of human-centered VA system for explanation analysis and distribution shift assessment [10]. That work highlighted that robust traffic classification requires more than accurate prediction. Models must remain reliable under evolving network conditions, while enabling experts to understand, validate, and improve their behavior. The practical experience obtained from that work revealed recurring challenges that extend beyond individual explanation techniques and highlighted the need for a systematic framework that integrates explainability, expert reasoning, semantic model validation, and robustness assessment throughout the ML lifecycle. Ultimately, the goal is to actively involve domain experts in the continuous evaluation and refinement of ML models, supporting traffic classifiers that remain robust, reliable, and trustworthy. The resulting insights may reveal deficiencies in the learned decision strategy, thereby informing revisions to earlier stages of the pipeline. Consequently, explainability is viewed not merely as a means of understanding model behavior, but as

2

a mechanism for continuously improving data, feature representations, models, and evaluation strategies. To support semantic model validation in network traffic classification, we adopt the knowledge generation model for VA proposed by Sacha et al. [39], as it already provides the core stages of iterative, human-centered knowledge generation. We adapt the framework by incorporating explainability, stakeholder-specific objectives, and explanation-based evaluation into the workflow. The resulting framework extends conventional performance-centric evaluation with semantic model validation, supporting the development of traffic classification models that are not only accurate but also transparent, robust, and correct for the right reasons. Overall, our contributions are:

• We identify recurring shortcomings in current network traffic classification workflows, including dataset artifacts, performance-centric evaluation, and limited integration of explainability and expert reasoning, and translate these findings into a practical framework for trustworthy model development. • We apply the knowledge generation model for VA into a human-centered framework for explainable traffic classification that connects user, tasks, goals, explanation requirements, and evaluation strategies to support trustworthy model development and assessment. • The resulting framework establishes a structured design space for human-centered XAI in traffic classification by organizing explanation generation and evaluation according to user roles, analytical goals, explanation levels, explanation requirements, and semantic validation criteria. • Based on the proposed framework, we derive design recommendations for trustworthy XAI evaluation by incorporating domain experts and semantic validation.

2 Related Work Network traffic classification has evolved considerably over the past decades. Early approaches relied on well-known port numbers and deep packet inspection (DPI), where applications were identified through predefined ports or protocol signatures. However, port-based classification became unreliable because applications increasingly employed dynamic and shared ports, while the growing adoption of encryption substantially reduced the effectiveness of DPI by concealing application-layer payloads [35]. These limitations motivated a shift toward ML-based traffic classification. Initial research employed classical supervised learning algorithms, such as decision trees, naı̈ve Bayes classifiers, support vector machines, k-nearest neighbors, random forests, and ensemble methods, which significantly improved classification performance compared with port-based classification and DPI. As larger benchmark datasets became available and computational resources advanced, deep learning (DL) emerged as the dominant paradigm, shifting the focus from manually engineered features to representation learning [4, 23, 42]. Neural networks automatically extract hierarchical features from raw traffic data, allowing increasingly complex spatial and temporal traffic patterns to be learned without explicit feature design [22]. Numerous neural network architectures have since been explored, including multilayer perceptrons, convolutional neural networks (CNNs),

3

recurrent neural networks (RNNs), long short-term memory (LSTM) networks, autoencoders, graph neural networks (GNNs), and transformer-based models [1, 12]. As these models have become increasingly complex, XAI has emerged as a complementary research direction for investigating and interpreting their decision-making processes [14, 24, 34]. Recent work demonstrates that explainability provides value beyond increasing the transparency of black-box models. Garcia et al. [16] employed explanation techniques to understand model predictions, debug DL classifiers, and simplify model architectures. Nascita et al. [31] analyzed feature importance to assess trustworthiness and improve multimodal traffic classification models, while their subsequent work used explanations to investigate the behavior of incrementally trained classifiers under evolving traffic scenarios [32]. Luis-Bisbé et al. [27] applied GradCAM to reveal shortcut learning, data leakage, and limitations of existing evaluation protocols. Collectively, these works demonstrate that explainability can support domain-specific reasoning and guide model refinement through expert interpretation. One of the early DL approaches to encrypted traffic classification is deep packet, proposed by Lotfollahi et al. [26]. The authors introduced a one-dimensional CNN that automatically learns feature representations directly from raw packet bytes and demonstrated strong classification performance on the ISCX VPN-nonVPN benchmark dataset [13]. Building upon this architecture, our subsequent work [8] incorporated class activation maps (CAM), an explainability technique originally developed for computer vision, to visualize the packet regions contributing to individual predictions and thereby enable the inspection of the model’s decision-making process [52]. The work adopted a design study methodology to elicit domain experts’ explanations of data and model behavior, informing the iterative design of explanation visualizations aligned with established network analysis workflows and familiar analytical representations [43]. Building upon this work, the approach was further extended by supporting the transition from local explanation analysis to subset-based aggregation, allowing analysts to investigate the overall decision strategy learned by the model for specific application classes [9]. Our preceding work investigated data drift and out-of-distribution detection [10], two major challenges in deployed ML systems, where changes in the underlying data distribution may lead to degraded predictive performance over time [28]. To support the investigation of such distributional changes, the proposed VA approach combines XAI with interactive visual analysis, enabling domain experts to detect, interpret, and validate data shifts using their domain knowledge. Collectively, these studies demonstrate that explainability alone is insufficient to support informed decision-making. Meaningful interpretation of both model behavior and data characteristics requires the active involvement of domain experts, whose expertise is essential for validating, contextualizing, and reasoning about the resulting insights. These observations motivated increased attention to sense-making, emphasizing the role of explanations in supporting expert reasoning and informed decision-making throughout the ML lifecycle [19, 29, 30, 30, 45, 50]. Similar principles have long been established in the field of VA, where interactive visual interfaces are designed to combine automated analysis with human reasoning in

4

order to facilitate knowledge generation [7, 20]. Rather than replacing human expertise, VA frameworks explicitly model iterative processes of exploration, hypothesis generation, verification, and knowledge acquisition through close interaction between computational methods and domain experts. Among these, the knowledge generation model proposed by Sacha et al. [39] provides a particularly suitable foundation, as it explicitly describes how human reasoning and computational analysis interact through iterative exploration and verification cycles to transform analytical findings into domain knowledge. These characteristics closely align with the objectives of explanation-based model analysis, where explanations serve not only to expose model behavior but also to support expert reasoning and informed decision-making. Recent surveys on XAI for network traffic analysis report similar observations. Although existing approaches have demonstrated the value of explainability for understanding ML-based traffic classifiers, they also identify the need for more humancentered explanation strategies, improved support for different stakeholder groups, and tighter integration of explainability into the ML lifecycle [33]. These findings further reinforce the importance of combining explainability with interactive analysis and expert-driven sense-making.

3 Limitations of Traffic Classification Workflows The proposed workflow is grounded in findings gathered throughout this research, including a systematic literature analysis, the development and evaluation of ML models for network traffic classification, the application of XAI techniques, empirical studies, expert interviews, and practitioner discussions [8–10]. This evidence consistently reveals recurring shortcomings in benchmark datasets (3.1), current model evaluation practices (3.2), and the integration of explainability and expert knowledge into model development (3.3), motivating the proposed workflow.

3.1 Dataset-Related Limitations A recurring observation across network traffic classification workflows is that model development is often emphasized more strongly than systematic analysis of the underlying datasets. While many studies remove obvious identifiers such as ethernet headers or source and destination IP addresses to reduce information leakage, less apparent sources of shortcut information often remain unnoticed. They may only become apparent through detailed dataset inspection or by analyzing model behavior using XAI techniques. Such analyses provide insights into the model’s decision-making process and reveal whether predictions are driven by meaningful traffic characteristics or unintended shortcut features. To investigate this issue, we analyzed the ISCX VPN-nonVPN dataset, one of the widely adopted benchmark datasets for encrypted network traffic classification [18, 44]. The dataset was constructed by generating representative real-world network activities across a diverse set of applications, including web browsing, streaming, VoIP, chat, file transfer, email, and peer-to-peer communication [13]. For each application scenario, traffic was captured both as regular network traffic and while routed through a VPN,

5

Table 1 Flow-level separability of the dataset. Unique flows is the fraction of an application’s flow identifiers (src port, dst port, protocol) that appear in no other application; Unambiguous packets is the fraction of its packets carried by those unique flows. High values indicate that the class is trivially separable from the flow identifier alone, before any payload inspection. Application

Number of flows

Number of packets

Unique flows

Unambiguous packets

AIM chat Email Facebook FTPS Hangouts ICQ Netflix SCP SFTP Skype Spotify Torrent VoipBuster Vimeo Youtube Average

399 3057 1463 551 1466 404 286 26 144 4671 299 629 2436 551 803 1146

4974 51 098 556 912 2 220 823 1 869 579 7476 732 836 183 251 423 560 1 470 613 97 470 269 096 1 459 926 366 694 268 310 665 508

0.64 0.84 0.87 0.97 0.87 0.66 0.88 0.96 0.91 0.88 0.93 0.98 0.78 0.91 0.96 0.87

0.49 0.71 0.98 0.99 0.99 0.56 0.23 0.99 0.99 0.98 0.25 0.98 0.99 0.86 0.69 0.78

resulting in fourteen traffic categories that represent a broad spectrum of encrypted communication patterns. Data-induced artifacts represent an often overlooked source of bias in network traffic classification datasets. Such artifacts unintentionally correlate with the target classes and may enable ML models to exploit shortcut features instead of learning meaningful characteristics of network traffic. To better understand the presence of such artifacts and account for them during the design of future benchmark datasets, we examine specific structural properties of the dataset from an artifact-oriented perspective. A network flow is a fundamental abstraction in network traffic analysis that represents a sequence of packets belonging to the same communication session between two communicating endpoints. For the purpose of our analysis, each flow is represented solely by its transport-layer identifier, i.e., the tuple (source port, destination port, protocol), deliberately excluding IP addresses and TCP sequence numbers. One potential source of data-induced artifacts arises from transport-layer flow identifiers that occur exclusively within a single application. Such application-specific flow identifiers may unintentionally encode the target class and therefore enable shortcut learning. To quantify this effect, we measure the fraction of unique flows, i.e., flow identifiers occurring exclusively within a single application, and the fraction of unambiguous packets, i.e., packets belonging to such unique flows (Table 1). Together, these metrics assess the extent to which transport-layer identifiers alone distinguish application classes. The analysis reveals that transport-layer identifiers alone provide substantial discriminative information. On average, 87% of all flow identifiers are unique to a single application, while 78% of all packets belong to such unique flows. For several applications, including Torrent, FTPS, SCP and SFTP, nearly all traffic can be distinguished solely by the transport-layer identifier before considering any payload. Consequently, a model can

6

achieve high predictive performance by implicitly learning application-specific port and protocol combinations rather than meaningful characteristics contained within the application-layer data carried in the packet payload. TORRENT (629) AIM (399)

EMAIL (3,057) well-known (<1024) (3,078)

:5355 (6,494)

FACEBOOK (1,463) FTPS (551)

registered (1024-49151) (3,848)

:443 (1,766)

HANGOUT (1,466) ICQ (404)

:10505 (1,556)

NETFLIX (286) SCP (26) SFTP (144)

:80 (437)

SKYPE (4,671)

ephemeral (>=49152) (10,259)

:15685 (261) :465 (250) :49539 (227) :3478 (159) :13000 (70) :61009 (66) :51413 (64) :40273 (63)

SPOTIFY (299) VOIPBUSTER (2,436)

other (5,772)

VIMEO (551) YOUTUBE (803)

Fig. 1 Sankey diagram of the per-application flow structure. Each flow, characterized by its signature (src port, dst port, protocol), is traced through three stages: source-port band (left), application class (centre), and destination port (right). Ribbon width is proportional to the number of distinct flow signatures and is coloured by application; source ports are grouped into the three IANA ranges, and the twelve most frequent destination ports are shown individually while the remainder are aggregated as other. Node labels give the number of flows.

The analysis further demonstrates that shortcut behavior is highly applicationdependent. For applications such as Netflix and Spotify, most flow identifiers remain unique, whereas the majority of packets are transmitted through a small number of shared, high-volume flows. Consequently, flow-level and packet-level separability differ substantially, indicating that the structural properties of individual application classes

7

can strongly influence the composition of the training and test sets. This observation is particularly relevant because many existing studies use random packet-level train-test splits. As a result, packets from the same network flow can appear in both the training and test sets. Since packets within a flow share common communication characteristics, this introduces information leakage and may lead to overly optimistic estimates of classification performance. Figure 1 provides a holistic view of the flow structure of the ISCX VPN-nonVPN dataset by visualizing every unique flow signature as a three-stage Sankey diagram. Each flow is represented by its transport-layer identifier, i.e., the tuple (src port, dst port, protocol), connecting the source-port range, the corresponding application class, and the destination port. Source ports are grouped according to the three IANA ranges (well-known, registered, and ephemeral), while the most frequent destination ports are shown individually and less frequent ports are aggregated into an other category. Ribbon width corresponds to the number of distinct flow signatures, enabling the structural relationships between transport-layer identifiers and application classes to be examined at a glance. The visualization reveals that the flow structure is highly application-dependent and far from uniformly distributed. Although most flows originate from ephemeral source ports, substantial numbers also originate from well-known and registered ports. On the destination side, the distribution is dominated by a small number of ports, most notably ports 5355 and 443, while many remaining ports occur only rarely. More importantly, several destination ports are shared across multiple applications, whereas others are almost exclusively associated with a single application. Applications whose transport-layer identifiers predominantly map to unique destination ports can be distinguished largely from their flow identifiers alone, whereas applications sharing common ports require additional discriminative information. Consequently, the intrinsic classification difficulty varies considerably between application classes before any packet payload is analyzed. These observations complement the quantitative results presented in Table 1 and demonstrate that the dataset itself already contains structural properties capable of biasing model learning toward transport-layer artifacts rather than characteristics of the encrypted traffic. Beyond transport-layer identifiers, additional shortcut sources may arise from the data collection process itself. For example, if traffic belonging to a particular application is captured during a specific time period or under consistent experimental conditions, temporal or environmental characteristics may become correlated with the target class, creating sampling-induced shortcuts that are unrelated to the actual application behavior. Such correlations can originate from the order of data collection, network configuration, operating system state, background traffic, or other experimental factors that are unintentionally preserved in the dataset. These artifacts are often difficult to identify through conventional preprocessing alone and typically require systematic dataset analysis or explanation-based inspection to become apparent. The observations presented are further supported by insights gathered through interviews and discussions with domain experts conducted during our evaluations, and through conversations with network analysis practitioners at conferences such as SharkFest.

8

Table 2 Overall performance (%, mean ± std). Packets are represented by the first 600 bytes from the IP header (TCP/UDP payload packets only). The CNN is initially trained with source and destination IP addresses masked. Tests 1–3 evaluate the same model under progressively stronger masking at test time (IP → IP+ports → IP+ports+TCP sequence/acknowledgement). The retrained model is trained and evaluated with all four fields masked. CNN results are averaged over 5 random seeds; As a complementary experiment, we trained a HistGradientBoostingClassifier (HGB) over 10 seeds using only the source port, destination port, and TCP sequence number (set to zero for UDP). Setting

Test-time masking

Accuracy

Macro-F1

CNN, Test 1 (train mask {ip}) CNN, Test 2 (train mask {ip}) CNN, Test 3 (train mask {ip}) CNN, retrained masked

{ip} {ip,seq,ack} {ip,seq,ack,port} {ip,seq,ack,port}

93.62 ± 0.06 87.94 ± 1.96 50.21 ± 1.58 91.12 ± 0.15

93.71 ± 0.06 87.92 ± 2.18 45.86 ± 1.92 91.20 ± 0.15

Tabular (HGB, ports+seq only)

—

93.28 ± 0.17

93.57 ± 0.16

3.2 Limitations of Performance-Centric Evaluation The performance and trustworthiness of ML models are inherently constrained by the quality of the data used for training. Regardless of the complexity of the learning algorithm, models can only learn from the information present in the dataset. Consequently, biases, artifacts, and unintended correlations introduced during data collection or preprocessing may become part of the learned decision strategy and remain unnoticed during conventional model evaluation. As demonstrated in the previous subsection 3.1, these representations may still contain capture metadata, lower-layer protocol information, and other dataset-specific artifacts that unintentionally encode class labels, thereby allowing models to exploit shortcut features instead of learning meaningful characteristics of encrypted traffic. Table 2 illustrates the impact of transport-layer identifiers on model performance. While progressively masking these fields leads to a substantial degradation in CNN accuracy, the identifier-only baseline achieves nearly the same performance using only source port, destination port, and TCP sequence number. Together, these results indicate that much of the predictive performance is driven by shortcut features rather than application-specific characteristics. The choice of evaluation metric is equally important, particularly for the highly imbalanced datasets commonly encountered in network traffic classification. For this reason, most studies appropriately report the macro F1-score, which weights all classes equally and therefore provides a more balanced assessment of classification performance [37, 53]. However, the reported F1-score is often insufficiently specified. Many studies simply state that precision, recall, and F1-score are computed from the numbers of true positives (TP), false positives (FP), and false negatives (FN), without indicating whether the reported F1-score corresponds to the micro, macro, or weighted average in the multiclass setting. This ambiguity complicates the interpretation and comparison of reported results. The use of the micro F1-score is dominated by majority classes and may therefore conceal poor performance on underrepresented applications [48, 51]. Aceto et al. additionally report the G-mean, which combines sensitivity

9

and specificity to provide a more balanced evaluation under class imbalance [2]. Some studies construct artificially balanced datasets by sampling an equal number of instances per class. In these circumstances, metric such as accuracy are appropriate. Beyond aggregated performance metrics, confusion matrices are widely used as an intuitive visual tool for analyzing class-specific prediction behavior, making systematic misclassifications and confusion between individual classes readily identifiable. Predictive performance alone provides only a limited assessment of model quality. High classification performance does not necessarily indicate that a model has learned meaningful characteristics of network traffic. Rather, performance metrics alone cannot determine whether predictions are based on semantically meaningful traffic patterns or on the exploitation of dataset-specific shortcut features.

3.3 Expert Perspectives and Explainability High predictive performance alone is insufficient for the practical deployment of classification models. Beyond achieving competitive benchmark results, practitioners require transparent and interpretable models that enable them to understand, verify, and validate model predictions before they can be trusted in operational environments. Identified shortcomings may indicate the need for improved dataset construction, revised preprocessing strategies, alternative model architectures, or modified evaluation procedures. This observation is consistent with recent literature advocating explainability as a prerequisite for trustworthy ML and was further reinforced by our own experiences during model evaluation, interviews with domain experts, and discussions with network analysis practitioners at technical conferences. Consequently, expert-guided interpretation should not be viewed as an isolated post-hoc analysis but rather as an iterative process that supports continuous improvement throughout the entire model development and deployment lifecycle. XAI provides an important mechanism for addressing this gap. By exposing the reasoning underlying model predictions, XAI enables experts to validate learned decision patterns against established domain knowledge, identify misleading or spurious features, and assess whether the model relies on semantically meaningful traffic characteristics. This not only increases confidence in model predictions but also facilitates a deeper understanding of model limitations and failure cases that remain hidden when relying solely on predictive performance metrics. The integration of XAI benefits from an iterative design study process [43]. Such a process combines the development of visual explanation techniques, continuous collaboration with domain experts using their familiar analytical tools, and repeated cycles of validation and reflection. Throughout these iterations, explainability evolves from a post-hoc interpretation tool into an integral component of model validation, dataset inspection, and iterative model improvement. In addition, it enables the extraction of domain insights and strengthens experts’ trust in both the learned models and their explanations. A first observation is that explanations inherently require context. While explanation methods can identify influential bytes, protocol fields, or packet regions, these highlighted features have little meaning in isolation. Their relevance can only be assessed by relating them to protocol semantics, application behavior, and networking expertise. Consequently, explanations cannot be interpreted automatically 10

but require human reasoning to determine whether influential features correspond to meaningful characteristics of traffic or merely reflect implementation details, protocol structures, or unintended dataset artifacts. Both expert studies emphasize that explainability acts primarily as an interface between machine learning developers and network analysts [8, 10].

Fig. 3 Global explanations of application classes obtained by aggregating local attribution maps across multiple samples. The visualization enables experts to identify consistent attribution patterns at the class level [11].

Fig. 2 Local explanation of an individual network traffic sample. The attribution map highlights the contribution of each input byte to the model’s prediction, allowing experts to inspect the decision rationale for a single classification [8].

Applying XAI revealed that only the initial bytes of the encrypted payload consistently contribute to the model’s predictions, whereas the majority of the payload receives negligible attribution scores. This behavior is illustrated in Figures 2 and 3, which show the attribution values for an exemplary packet classification and the corresponding aggregated attribution values at the application-class level, respectively. This indicates that large portions of the encrypted payload should not contain discriminative information for the classification task, since correctly encrypted payload bytes are intended to be indistinguishable from random noise. Consequently, substantially smaller input representations may be sufficient, reducing computational complexity while maintaining predictive performance. This observation is supported by previous work [10], which demonstrated that restricting the model input to the informative payload region does not degrade classification performance. The interpretation of such findings requires confirmation by domain experts before they are incorporated into feature selection, preprocessing, or model refinement. Another important lesson concerns the aggregation of explanations. While local explanations provide insights into individual predictions, practitioners are often interested in understanding the overall decision strategy learned for a subset of data (e.g., an application class, illustrated in Figure 3). This requires representative global explanations. Numerous mathematical aggregation strategies exist, each emphasizing different properties of the underlying explanation values. Selecting an appropriate aggregation strategy is consequently essential for obtaining meaningful global explanations. Moreover, the resulting explanations require expert interpretation and should

11

support interactive exploration of different data subsets. Besides selecting an appropriate aggregation strategy, particular attention must be paid to the intuitive visual representation of global explanations, as aggregating explanations from many samples introduces additional visualization and interpretation challenges than explaining individual predictions. Collectively, these findings motivate the need for a systematic workflow that supports explanation-based model validation and semantic assessment. Section 4 introduces the proposed workflow.

4 Specific Workflow for Network Traffic Classification Computer Visualization

Expert Roles Action

Hypothesis Knowledge

Data

Verification loop Exploration loop

Model XAI

Finding

Knowledge loop

Insight

Sensemaking loop

Fig. 4 Adaptation of the knowledge generation model for VA to network traffic classification. Compared with the original model [39], the framework explicitly incorporates XAI into the computational pipeline and refines the human component into multiple expert roles with stakeholder-specific objectives, supporting explanation-based exploration, verification, and knowledge generation

Motivated by the findings presented in the previous section, we apply the knowledge generation model for VA proposed by Sacha et al. [39, 40] to support the development and semantic validation of network traffic classification models. This model provides the most suitable conceptual foundation for our work because it explicitly integrates computational analysis with human cognitive reasoning. Existing process models differ in their objectives. Data science methodologies, such as KDD [15], CRISP-DM [6], and CRISP-ML(Q) [47], focus on the engineering lifecycle of ML systems, including data preparation, model development, evaluation, and deployment. They primarily assess models through predictive performance and operational quality, without incorporating XAI, interactive visual analytics, or expert-driven semantic assessment. Visualization and visual analytics models, such as those proposed by van Wijk [49], and Keim et al. [20], move closer to our objective by integrating data, analytical models, visualization, and user interaction to support knowledge generation. Nevertheless, these frameworks do not explicitly represent the human reasoning process and support the semantic assessment of ML models. The knowledge generation model proposed by Sacha et al. consists of two tightly coupled components, as illustrated in Figure 4. The computer system comprises

12

network traffic data, ML models, and interactive visualizations. The human component represents the analyst’s cognitive processes, including interpretation, reasoning, hypothesis generation, and decision-making based on visualized data, model information, and classification explanations. In the following, we describe how these components are instantiated and adapted to the domain of network traffic classification.

4.1 Data The data component comprises the structured network traffic data used for model training and assessment. Since the framework addresses application classification, the dataset is assumed to be labeled. Although the data follows a structured representation, its characteristics depend on the underlying network protocols, each providing protocol-specific fields and semantics. From a functional perspective, the data must support appropriate preprocessing and transformation steps prior to model development. This includes the identification and removal or masking of known sources of bias and dataset-specific artifacts that could lead to shortcut learning rather than meaningful feature extraction. Common examples include masking IP addresses, port numbers, sequence and acknowledgment numbers, and removing Ethernet frame information. Padding or truncation may also be required as part of the preprocessing pipeline to obtain a homogeneous input representation for the model. In addition, the data should satisfy several non-functional requirements. It should accurately reflect real-world application traffic to ensure meaningful semantic evaluation, be sufficiently complete with minimal missing or corrupted values, and provide representative coverage of the application classes. These properties are essential to ensure that the generated explanations correspond to genuine traffic characteristics rather than artifacts of the dataset itself. The data partitioning strategy should preserve the independence of the training, validation, and test sets. All packets belonging to the same flow must be placed exclusively in either the training, validation, or test set and must never be distributed across multiple partitions. Otherwise, the model may exploit flow-specific characteristics shared between packets of the same flow, leading to information leakage, overly optimistic performance estimates, and an unreliable assessment of its generalization capability.

4.2 Visualization The visualization component represents the interface through which analytical results are transformed into actionable knowledge. Rather than serving solely as a means of presenting data, visualizations should support the exploration, interpretation, and validation of both the underlying data and the generated explanations. Consequently, visualization acts as the primary communication channel between automated analysis and human reasoning. From a functional perspective, visualizations should present information in a manner that is intuitive and consistent with the analytical task. Since explanations can be

13

Table 3 Stakeholder-specific objectives and explanation requirements.

Aspect

ML Experts

Network Experts

Primary objective

Understand, validate, and improve the behavior of the ML model.

Expert knowledge

ML optimization, model architectures, explainability methods, feature engineering.

Typical tasks

analytical

Primary focus

evaluation

Model debugging, shortcut detection, robustness analysis, explanation validation, comparison of alternative models. Faithfulness, robustness, consistency, sensitivity, and explanation reliability. Knowledge about model behavior, explanation reliability, and opportunities for improving the learning process. Model refinement, preprocessing improvements, feature engineering, training strategy.

Validate whether the learned decision strategies correspond to meaningful network patterns. Network protocols, packet structures, traffic analysis, application behavior, communication semantics. Protocol validation, semantic interpretation, artifact detection, deployment assessment, traffic investigation. Semantic correctness, usefulness and transferability.

Knowledge generated

Feedback workflow

to

the

Knowledge about semantic validity, protocol behavior, dataset quality, and deployment suitability. Dataset refinement, semantic validation, feature engineering, data collection, deployment decisions.

provided at both local and global levels, the visualization should support the exploration of individual traffic instances as well as aggregated explanations describing application classes or the classifier as a whole. Different user groups interact with XAI systems with fundamentally different objectives (Table 3), which directly influence their explanation requirements. ML practitioners primarily seek explanations to debug and improve models, identify erroneous or biased decision behavior, understand which information the model utilizes, and assess its strengths and limitations. In contrast, domain experts are typically interested in understanding and validating model predictions within the application context, integrating model outputs into downstream decision-making, learning new domain knowledge, assessing prediction reliability, or contesting model decisions when necessary. Consequently, the visualization should not adopt a one-size-fits-all approach but instead provide user-centered explanation interfaces that adapt the presented information and level of detail to the analytical goals and expertise of the intended users (context). Moreover, explanations should remain coherent with networking concepts and protocol semantics to support meaningful interpretation [38] (Table 4). Furthermore, visualizations should facilitate interactive exploration by allowing analysts to inspect different subsets of traffic, compare explanations across applications or models, and investigate alternative scenarios (see controllability in Table 4). Such interaction enables analysts to iteratively formulate and validate hypotheses, thereby supporting the knowledge generation process rather than merely presenting static results. Finally, the visualization should minimize cognitive effort by emphasizing relevant information while avoiding unnecessary visual complexity. Explanation representations

14

Table 4 Presentation- and user-related evaluation requirements for XAI explanations in network traffic classification [34].

Category Presentation

User

Characteristic Objective Compactness Explanations should remain concise by highlighting only the most relevant features and avoiding redundant information. Composition Explanations should be presented using intuitive and well-structured visual representations that are appropriate for the explanation type (e.g., local or global) and support efficient analytical workflows. Context Explanations should provide information relevant to the analytical objectives and domain knowledge of network analysts. Coherence Explanations should be consistent with networking knowledge and protocol semantics while avoiding dataset-specific artifacts. Controllability The explanation interface should support interactive exploration, comparison of traffic subsets, and investigation of alternative scenarios.

should remain compact. Their composition should employ intuitive, well-structured visual representations that are appropriate for the explanation type. Furthermore, explanations should provide sufficient context by relating highlighted features to the analyst’s objectives and domain knowledge, while maintaining coherence with networking concepts and protocol semantics (Table 4).

4.3 Model The model component represents the ML model responsible for learning patterns from the input data and correctly classifying network traffic into the corresponding classes. Its primary goal is to learn meaningful decision patterns that generalize beyond the training data rather than memorizing dataset-specific characteristics. Besides achieving high predictive performance, the model should demonstrate good generalization to unseen data and robustness against variations in the input. Accordingly, model evaluation should include cross-validation to obtain reliable performance estimates, multiple training runs with different random initializations to assess stability, and robustness testing on unseen or perturbed data to evaluate the reliability and generalizability of the learned decision function [25]. The model should be evaluated using metrics appropriate for the underlying class distribution. For imbalanced datasets, metrics such as macro F1-score, precision, and recall provide a more informative assessment than accuracy alone [46]. For balanced datasets, accuracy is also an appropriate performance measure. In addition, confusion matrices provide valuable insights into the model’s prediction behavior by revealing application classes that are frequently confused and therefore require further investigation.

15

Table 5 Content-related evaluation requirements for XAI explanations in network traffic classification [34].

Category

Characteristic Objective

Example Evaluation

Correctness

Feature deletion/insertion, perturbation analysis.

Content

Consistency

Continuity

Explanations should faithfully represent the reasoning of the underlying classifier rather than producing only plausible feature attributions. Identical traffic samples should produce identical explanations independent of implementation details. Similar network flows or packets should produce similar explanations, demonstrating robustness to small variations in the input.

Repeated explanations for identical inputs, invariance across equivalent models, consistency across different random initializations Stability under slight input perturbations, explanation similarity for neighboring flows/packets, fidelity under small traffic variations.

4.4 XAI Methods XAI comprises a broad range of methods that provide insights into the decision-making process of machine learning models. Existing approaches can be broadly categorized into intrinsic and post-hoc methods, with the latter being most commonly applied in network traffic classification because they can explain already trained black-box models without modifying their architecture [5]. Post-hoc methods can further be distinguished into feature attribution techniques, concept-based explanations, surrogate models, example-based explanations, and counterfactual explanations [14, 34]. Feature attribution methods, such as SHAP, LIME, Integrated Gradients, Layer-wise relevance propagation (LRP), and CAM, are among the most widely used approaches because they identify the input features that contribute most strongly to individual predictions [5, 14]. Because explanations serve as the basis for semantic model validation and expert reasoning, their quality must also be assessed. Therefore, we incorporate an XAI evaluation component into the framework (Figure 4). This component provides a structured approach for assessing the quality of explanations generated for network traffic classification. Unlike predictive performance, explanation quality cannot be adequately characterized by a single metric. The evaluation requirements adopted in this framework are summarized in Tables 5 and are adapted from established explanation quality dimensions proposed in the XAI literature [34]. From the content-related characteristics, we focus on correctness, consistency, and continuity because they evaluate whether explanations faithfully represent the model’s reasoning. The content-related characteristics are primarily evaluated through technical, functionally grounded evaluation methods that assess the behavior of the explanation algorithm independently of human judgment. Examples include perturbation analyses, feature deletion and insertion experiments, repeated explanation generation, and robustness analyses under small input variations.

16

4.5 Exploration Loop The exploration loop describes how analysts iteratively investigate explanation results to identify meaningful patterns in model behavior and network traffic. Exploration may begin with a suspicious prediction, an unexpected feature attribution, or a hypothesis regarding protocol-specific behavior. Analysts can inspect local explanations for individual traffic instances, aggregate explanations across traffic subsets or application classes, compare different models, and switch between local and global perspectives to identify recurring explanation patterns and potential anomalies. The resulting findings provide insights into the classifier’s decision-making process and may reveal discriminative traffic characteristics, protocol-specific behaviors, unexpected feature attributions, shortcut learning, or dataset-specific artifacts. These findings represent candidate explanations of the observed model behavior and form the basis for subsequent verification.

4.6 Verification Loop The verification loop systematically evaluates candidate findings identified during exploration to determine whether they reflect semantically meaningful classifier behavior or originate from shortcut learning, dataset artifacts, or limitations of the explanation method. Verification should therefore combine complementary sources of evidence, including alternative explanation methods, quantitative evaluation metrics, additional datasets, different model architectures, and expert assessment. Depending on the investigated hypothesis, verification may examine whether highlighted protocol fields correspond to known communication behavior, whether similar explanation patterns are consistently observed across different explanation methods, models, or datasets, and whether domain experts consider the identified decision strategy plausible and consistent with protocol semantics. Quantitative analyses provide evidence regarding the reproducibility and robustness of the observed explanation patterns, while expert interpretation determines whether these patterns are meaningful from a networking perspective. Only findings that are consistently supported by complementary technical evidence and expert interpretation should be regarded as validated knowledge. Such knowledge can subsequently guide refinements of the dataset, feature representation, explanation strategy, or the underlying ML model, thereby supporting the development of more robust and semantically meaningful traffic classifiers. For example, if technical analyses consistently show that a classifier relies on the destination port to identify encrypted Spotify traffic, a network expert may conclude that this behavior reflects a datasetspecific artifact rather than semantically meaningful traffic characteristics. This insight may motivate masking the destination port or redesigning the feature representation before retraining the model.

4.7 Knowledge Generation Loop The knowledge generation loop consolidates validated findings into reusable knowledge that can support future model development and evaluation. Rather than focusing on individual explanations or specific observations, it captures general insights about 17

Table 6 Design recommendations for trustworthy network traffic classification derived from the adapted knowledge generation framework

Guideline category

Recommendation

Purpose

Data

Remove artifacts and ensure independent data partitions Evaluate predictive performance using appropriate metrics and assess robustness, stability, and generalization. Evaluate explanation correctness, consistency, continuity, and comprehensibility Design interactive visualizations tailored to stakeholderspecific explanation requirements Explore explanation patterns across traffic subsets, application classes, and model variants. Verify candidate findings using complementary technical evidence and expert assessment Capture validated findings as reusable design recommendations and best practices.

Reduces shortcut learning and information leakage. Provides evidence of the model’s predictive performance and reliability.

Model

XAI Visualization

Exploration

Verification

Knowledge

Ensures that explanations can be reliably interpreted for semantic model analysis. Supports efficient exploration and interpretation of explanation results. Identifies candidate findings, anomalies, and potential shortcut learning. Determines whether observed patterns are reproducible and semantically meaningful. Supports future dataset design, model development, explanation design, and evaluation.

network traffic characteristics, model behavior, explanation methods, and the overall development process. The generated knowledge may describe reliable protocol-specific communication patterns, discriminative traffic characteristics, common sources of shortcut learning, dataset limitations, or recurring explanation patterns observed across different models and datasets. It may also identify effective explanation techniques, visualization strategies, or evaluation procedures that consistently support semantic model validation. Unlike the verification loop, which determines whether individual findings are supported by sufficient evidence, the knowledge generation loop generalizes these validated findings into reusable insights and recommendations. The resulting knowledge can subsequently guide future dataset construction, feature engineering, explanation design, visualization development, model evaluation, and ML model development. As this knowledge is reused in subsequent analyses, it influences new analytical objectives, supports more efficient exploration, and contributes to the continuous improvement of explainable network traffic classification workflows. To facilitate the practical application of the proposed framework, Table 6 summarizes the main design guidelines associated with each framework component.

18

5 Discussion and Future Work The proposed framework provides a structured mechanism for capturing knowledge generated during explanation-based model analysis and feeding it back into earlier stages of the development process. Beyond supporting the analysis of individual models, this enables recurring findings, such as protocol-specific shortcut features, dataset artifacts, explanation patterns, and effective preprocessing strategies, to be systematically documented and reused across subsequent studies. Such accumulated knowledge may improve the design of future benchmark datasets by increasing their transparency, documenting known sources of information leakage, and motivating standardized preprocessing and partitioning strategies. It also improves the reproducibility of analyses by explicitly documenting the rationale behind dataset modifications and model design decisions. Because the proposed framework separates explanation generation from explanation analysis, it provides a common evaluation process in which different XAI methods can be assessed under consistent evaluation criteria and analytical tasks. This enables future work to compare explanation techniques not only with respect to their technical characteristics but also according to their practical utility during semantic validation. Furthermore, it facilitates the investigation of explanation-specific design choices, such as how local attribution values should be aggregated into representative global explanations and which aggregation strategies are most suitable for different analytical objectives. For example, aggregating explanations across an entire application class or only across the most similar samples within a class. A limitation of the proposed framework is the additional effort required to integrate explanation generation, interactive analysis, and expert-driven verification throughout the ML pipeline. Depending on the application scenario, incorporating all workflow components may not always be feasible because of constraints in time, computational resources, or domain expertise. Nevertheless, the framework is designed to support incremental adoption. Even the integration of individual components establishes a consistent process for documenting findings and feeding validated insights back into subsequent development iterations. As additional workflow components are incorporated, the accumulated knowledge base expands, enabling the framework to be continuously refined through practical application and facilitating the transfer of validated insights across different datasets, models, and traffic classification scenarios.

6 Conclusion This paper addressed the limitations of performance-centric evaluation in network traffic classification by emphasizing the importance of semantic model validation. Building on findings synthesized from the literature, empirical analyses, practical experience with explainable traffic classification, and expert feedback, we identified key considerations for integrating explanation-based analysis into the model development process. These findings motivated the adaptation of the knowledge generation model for VA to the domain of network traffic classification, resulting in a human-centered framework that integrates data, ML, explainability, visualization, and expert reasoning into

19

an iterative workflow. Finally, the framework was translated into a set of design recommendations that provide practical guidance for the trustworthy development and evaluation of network traffic classification models.

Acknowledgements. This research work was supported by the National Research Center for Applied Cybersecurity ATHENE. ATHENE is funded jointly by the German Federal Ministry of Research, Technology and Space and the Hessian Ministry of Science and Research, Arts and Culture. Author Contributions. Igor Cherepanov was primarily responsible for the conception of the work and the preparation of the manuscript. All authors have read and approved the final manuscript and agree to be accountable for all aspects of the work. Data Availability. Not applicable. Code Availability. Not applicable.

7 Declarations Competing Interests. The authors have no relevant financial or non-financial competing interests to disclose. Research Involving Human Participants and/or Animals. No animals were involved in this research. Informed Consent. Participants taking part in this research have done so freely and is based on their willingness to share information crucial for pursuing the purpose and objective of the paper. Open Access. This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution, and reproduction in any medium or format, as long as appropriate credit is given to the original author(s) and the source, a link to the Creative Commons licence is provided, and any changes made are indicated. The images or other third-party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and the intended use is not permitted by statutory regulation or exceeds the permitted use, permission must be obtained directly from the copyright holder. To view a copy of this licence, visit https://creativecommons.org/licenses/by/4.0/.

References [1] Abbasi M, Shahraki A, Taherkordi A (2021) Deep learning for network traffic monitoring and analysis (ntma): A survey. Computer Communications 170:19–41. https://doi.org/https://doi.org/10.1016/j.comcom.2021.01.021 [2] Aceto G, Ciuonzo D, Montieri A, et al (2019) Mobile encrypted traffic classification using deep learning: Experimental evaluation, lessons learned, and challenges. IEEE Transactions on Network and Service Management 16(2):445–458. https: //doi.org/10.1109/TNSM.2019.2899085 20

[3] Amershi S, Weld D, Vorvoreanu M, et al (2019) Guidelines for human-ai interaction. In: Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery, New York, NY, USA, CHI ’19, p 1–13, https://doi.org/10.1145/3290605.3300233 [4] Azab A, Khasawneh M, Alrabaee S, et al (2024) Network traffic classification: Techniques, datasets, and challenges. Digital Communications and Networks 10(3):676–692. https://doi.org/https://doi.org/10.1016/j.dcan.2022.09.009 [5] Barredo Arrieta A, Dı́az-Rodrı́guez N, Del Ser J, et al (2020) Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai. Information Fusion 58:82–115. https://doi.org/https://doi.org/10. 1016/j.inffus.2019.12.012 [6] Chapman P (2000) Crisp-dm 1.0: Step-by-step data mining guide. URL https: //api.semanticscholar.org/CorpusID:59777418 [7] Chatzimparmpas A (2025) Visual analytics for explainable and trustworthy artificial intelligence. IEEE Computer Graphics and Applications 45(2):100–111. https://doi.org/10.1109/MCG.2025.3533806 [8] Cherepanov I, Ulmer A, Joewono JG, et al (2022) Visualization of class activation maps to explain ai classification of network packet captures. In: 2022 IEEE Symposium on Visualization for Cyber Security (VizSec), pp 1–11, https: //doi.org/10.1109/VizSec56996.2022.9941392 [9] Cherepanov I, Sessler D, Ulmer A, et al (2023) Towards the visualization of aggregated class activation maps to analyse the global contribution of class features. In: Longo L (ed) Explainable Artificial Intelligence. Springer Nature Switzerland, Cham, pp 3–23, https://doi.org/10.1007/978-3-031-44067-0 [10] Cherepanov I, Sessler D, Feil A, et al (2026) Model-aware visual analytics for aligning data shift in network traffic classification. In: Proceedings of the 21st International Conference on Computer Graphics, Interaction and Visualization Theory and Applications - GRIVAPP, INSTICC. SciTePress, pp 63–74, https: //doi.org/10.5220/0014328100004728 [11] Cherepanov I, Sessler D, Ulmer A, et al (2026) Interactive analysis of global explanations using aggregated class activation maps for network data. URL https: //arxiv.org/abs/2608.13575, arXiv:2608.13575 [12] Dong W, Yu J, Lin X, et al (2025) Deep learning and pre-training technology for encrypted traffic classification: A comprehensive review. Neurocomputing 617:128444. https://doi.org/https://doi.org/10.1016/j.neucom.2024.128444

21

[13] Draper-Gil G, Lashkari AH, Mamun MSI, et al (2016) Characterization of encrypted and vpn traffic using time-related. In: Proceedings of the 2nd international conference on information systems security and privacy (ICISSP), pp 407–414, https://doi.org/10.5220/0005740704070414 [14] Dwivedi R, Dave D, Naik H, et al (2023) Explainable ai (xai): Core ideas, techniques, and solutions. ACM Comput Surv 55(9). https://doi.org/10.1145/ 3561048 [15] Fayyad U, Piatetsky-Shapiro G, Smyth P (1996) The kdd process for extracting useful knowledge from volumes of data. Communications of the ACM 39(11):27 – 34. https://doi.org/10.1145/240455.240464 [16] Garcia L, Bartlett G, Ravi S, et al (2022) Explaining deep learning models for per-packet encrypted network traffic classification. In: 2022 IEEE International Symposium on Measurements & Networking (M&N), pp 1–6, https://doi.org/10. 1109/MN55117.2022.9887744 [17] Geirhos R, Jacobsen JH, Michaelis C, et al (2020) Shortcut learning in deep neural networks. Nature Machine Intelligence 2(11):665–673. https://doi.org/10. 1038/s42256-020-00257-z [18] Han S, Zhang H, Lu M, et al (2026) A comprehensive survey on encrypted network traffic classification. Computer Networks 287:112524. https://doi.org/https: //doi.org/10.1016/j.comnet.2026.112524 [19] Hoffman RR, Mueller ST, Klein G, et al (2018) Metrics for explainable ai: Challenges and prospects. arXiv preprint arXiv:181204608 [20] Keim DA, Mansmann F, Schneidewind J, et al (2008) Visual analytics: Scope and challenges. In: Simoff SJ, Böhlen MH, Mazeika A (eds) Visual Data Mining: Theory, Techniques and Tools for Visual Analytics. Springer Berlin Heidelberg, Berlin, Heidelberg, pp 76–90, https://doi.org/10.1007/978-3-540-71080-6 [21] Kucher K, Zohrevandi E, Westin CAL (2025) Towards visual analytics for explainable ai in industrial applications. Analytics 4(1). https://doi.org/10.3390/ analytics4010007 [22] LeCun Y, Bengio Y, Hinton G (2015) Deep learning. nature 521(7553):436–444. https://doi.org/10.1038/nature14539 [23] Lim HK, Kim JB, Heo JS, et al (2019) Packet-based network traffic classification using deep learning. In: 2019 International Conference on Artificial Intelligence in Information and Communication (ICAIIC), pp 046–051, https://doi.org/10. 1109/ICAIIC.2019.8669045

22

[24] Linardatos P, Papastefanopoulos V, Kotsiantis S (2021) Explainable ai: A review of machine learning interpretability methods. Entropy 23(1). https://doi.org/10. 3390/e23010018 [25] Lones MA (2024) Avoiding common machine learning pitfalls. Patterns 5(10):101046. https://doi.org/10.1016/j.patter.2024.101046 [26] Lotfollahi M, Jafari Siavoshani M, Shirali Hossein Zade R, et al (2020) Deep packet: A novel approach for encrypted traffic classification using deep learning. Soft Computing 24(3):1999–2012. https://doi.org/10.1007/s00500-019-04030-2 [27] Luis-Bisbé E, Morales-Gómez V, Perdices D, et al (2024) No pictures, please: Using explainable artificial intelligence to demystify cnns for encrypted network packet classification. Applied Sciences 14(13). https://doi.org/10.3390/ app14135466 [28] Malekghaini N, Akbari E, Salahuddin MA, et al (2022) Data drift in dl: Lessons learned from encrypted traffic classification. In: 2022 IFIP Networking Conference (IFIP Networking), pp 1–9, https://doi.org/10.23919/IFIPNetworking55013. 2022.9829791 [29] Miller T (2019) Explanation in artificial intelligence: Insights from the social sciences. Artificial intelligence 267:1–38. https://doi.org/10.1016/j.artint.2018.07. 007 [30] Mohseni S, Zarei N, Ragan ED (2021) A multidisciplinary survey and framework for design and evaluation of explainable ai systems. ACM Trans Interact Intell Syst 11(3–4). https://doi.org/10.1145/3387166 [31] Nascita A, Montieri A, Aceto G, et al (2021) Xai meets mobile traffic classification: Understanding and improving multimodal deep learning architectures. IEEE Transactions on Network and Service Management 18(4):4225–4246. https: //doi.org/10.1109/TNSM.2021.3098157 [32] Nascita A, Cerasuolo F, Aceto G, et al (2023) Explainable mobile traffic classification: the case of incremental learning. In: Proceedings of the 2023 on Explainable and Safety Bounded, Fidelitous, Machine Learning for Networking. Association for Computing Machinery, New York, NY, USA, SAFE ’23, p 25–31, https://doi.org/10.1145/3630050.3630178 [33] Nascita A, Aceto G, Ciuonzo D, et al (2024) A survey on explainable artificial intelligence for internet traffic classification and prediction, and intrusion detection. IEEE Communications Surveys & Tutorials 27(5):3165–3198. https: //doi.org/10.1109/COMST.2024.3504955 [34] Nauta M, Trienes J, Pathak S, et al (2023) From anecdotal evidence to quantitative evaluation methods: A systematic review on evaluating explainable ai. ACM

23

Comput Surv 55(13s). https://doi.org/10.1145/3583558 [35] Nguyen TT, Armitage G (2008) A survey of techniques for internet traffic classification using machine learning. IEEE Communications Surveys & Tutorials 10(4):56–76. https://doi.org/10.1109/SURV.2008.080406 [36] Pacheco F, Exposito E, Gineste M, et al (2019) Towards the deployment of machine learning solutions in network traffic classification: A systematic survey. IEEE Communications Surveys & Tutorials 21(2):1988–2014. https://doi.org/10. 1109/COMST.2018.2883147 [37] Peng L, Xie X, Huang S, et al (2024) Ptu: Pre-trained model for network traffic understanding. In: 2024 IEEE 32nd International Conference on Network Protocols (ICNP), pp 1–12, https://doi.org/10.1109/ICNP61940.2024.10858503 [38] Rong Y, Leemann T, Nguyen TT, et al (2024) Towards human-centered explainable ai: A survey of user studies for model explanations. IEEE Transactions on Pattern Analysis and Machine Intelligence 46(4):2104–2122. https://doi.org/10. 1109/TPAMI.2023.3331846 [39] Sacha D, Stoffel A, Stoffel F, et al (2014) Knowledge generation model for visual analytics. IEEE Transactions on Visualization and Computer Graphics 20(12):1604–1613. https://doi.org/10.1109/TVCG.2014.2346481 [40] Sacha D, Sedlmair M, Zhang L, et al (2017) What you see is what you can change: Human-centered machine learning by interactive visualization. Neurocomputing 268:164–175. https://doi.org/https://doi.org/10.1016/j.neucom.2017.01.105, advances in artificial neural networks, machine learning and computational intelligence [41] Salman O, Elhajj IH, Kayssi A, et al (2020) A review on machine learning– based approaches for internet traffic classification. Annals of Telecommunications 75(11):673–710. https://doi.org/10.1007/s12243-020-00770-7 [42] Salman O, Elhajj IH, Kayssi A, et al (2020) A review on machine learning– based approaches for internet traffic classification. Annals of Telecommunications 75(11):673–710. https://doi.org/10.1007/s12243-020-00770-7 [43] Sedlmair M, Meyer M, Munzner T (2012) Design study methodology: Reflections from the trenches and the stacks. IEEE Transactions on Visualization and Computer Graphics 18(12):2431–2440. https://doi.org/10.1109/TVCG.2012.213 [44] Sharma A, Lashkari AH (2025) A survey on encrypted network traffic: A comprehensive survey of identification/classification techniques, challenges, and future directions. Computer Networks 257:110984. https://doi.org/https://doi.org/10. 1016/j.comnet.2024.110984

24

[45] Shneiderman B (2022) Human-centered ai: ensuring human control while increasing automation. In: Proceedings of the 5th Workshop on Human Factors in Hypertext. Association for Computing Machinery, New York, NY, USA, HUMAN ’22, https://doi.org/10.1145/3538882.3542790 [46] Sokolova M, Lapalme G (2009) A systematic analysis of performance measures for classification tasks. Information Processing & Management 45(4):427–437. https://doi.org/https://doi.org/10.1016/j.ipm.2009.03.002 [47] Studer S, Bui TB, Drescher C, et al (2021) Towards crisp-ml(q): A machine learning process model with quality assurance methodology. Machine Learning and Knowledge Extraction 3(2):392–413. https://doi.org/10.3390/make3020020 [48] Wang T, Xie X, Wang W, et al (2024) Netmamba: Efficient network traffic classification via pre-training unidirectional mamba. In: 2024 IEEE 32nd International Conference on Network Protocols (ICNP), IEEE, pp 1–11, https://doi.org/10. 1109/ICNP61940.2024.10858569 [49] van Wijk J (2005) The value of visualization. In: VIS 05. IEEE Visualization, 2005., pp 79–86, https://doi.org/10.1109/VISUAL.2005.1532781 [50] Xu W (2019) Toward human-centered ai: a perspective from human-computer interaction. Interactions 26(4):42–46. https://doi.org/10.1145/3328485 [51] Zhao R, Zhan M, Deng X, et al (2023) Yet another traffic classifier: A masked autoencoder based traffic transformer with multi-level flow representation. In: Proceedings of the AAAI Conference on Artificial Intelligence, pp 5420–5427, https://doi.org/10.1609/aaai.v37i4.25674 [52] Zhou B, Khosla A, Lapedriza A, et al (2016) Learning deep features for discriminative localization. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) [53] Zhou G, Guo X, Liu Z, et al (2025) Trafficformer: An efficient pre-trained model for traffic data. In: 2025 IEEE Symposium on Security and Privacy (SP), pp 1844–1860, https://doi.org/10.1109/SP61157.2025.00102

25

Record · ID 919289 · SHA-256 6c23291bb9ae5e79
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.