Conceptio › Archive › arXiv CS
arXiv CSopen access

TRIPROBE: Probing Task Separability Beyond Classification for XAI

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

arXiv:2609.18525v1 [cs.AI] 16 Sep 2026

TRIPROBE: Probing Task Separability Beyond Classification for XAI 1st Amirhossein Sadough

2nd Freek Hens

3rd Aleksa Bokšan

Machine Learning and Neural Computing Radboud University Nijmegen, Netherlands ORCID:0009-0005-5647-2888

Machien Learning and Neural Computing Radboud University Nijmegen, Netherlands ORCID:0009-0005-0405-0560

Delft University of Technology Delft, Netherlands ORCID:0009-0003-3099-4472

4th Mohammad Mahdi Dehshibi

5th Mahyar Shahsavari

Unconventional Computing Lab University of the West of England (UWE) Bristol, United Kingdom ORCID: 0000-0001-8112-5419

Machine Learning and Neural Computing Radboud University Nijmegen, Netherlands ORCID:0000-0002-9671-0917

Abstract—Modern evaluation of learning pipelines often reduces to downstream accuracy, leaving open the question of why tasks succeed or fail. TriProbe addresses this gap with a multi-level probing framework for explainable diagnosis of task separability. Rather than treating models as black boxes, TriProbe traces how separability evolves across inputs, learned features, and final classifiers. It decomposes multi-task problems into binary subtasks and applies three complementary probes: a Foundational Probe on input spaces, a Latent Probe on feature representations, and a Final Probe on classifier outputs. Using Maximum Fisher’s Discriminant Ratio as a principled separability metric, TriProbe identifies bottlenecks and affected task pairs. Experiments on the Roshambo sEMG benchmark show how TriProbe reveals hidden breakdowns, guiding data collection, validation, and architecture design. Index Terms—Explainable AI, task separability, probing framework, Fisher’s discriminant ratio, learning pipeline

I. I NTRODUCTION Understanding why learning models succeed or fail in distinguishing between tasks remains a fundamental challenge in machine learning and signal processing. Across diverse application areas, including speech, EEG/MEG, and sEMG, task separability is affected not only by intrinsic data complexity but also by architectural choices across the learning pipeline [1]. Existing evaluation methods typically focus on downstream accuracy, providing little insight into where separability bottlenecks emerge [2]–[4]. This lack of interpretability complicates both data collection and model design, motivating the need for systematic and explainable tools [5], [6]. Model probing has emerged as a promising methodology for understanding how representations evolve in modern learning pipelines. Early work introduced diagnostic classifiers to test whether embeddings captured linguistic or structural information [7]–[9]. Complementary explainable AI approaches, such as Shapley values and integrated gradients, provide attribution scores but typically stop short of tracing separability across multiple stages of the pipeline [10], [11]. In parallel, studies

in physiological signals, such as sEMG gesture recognition, have emphasized the importance of transfer learning and representation quality for task separability [12], [13]. However, these methods primarily aim at boosting accuracy rather than systematically diagnosing where separability is gained or lost across stages. To address this gap, we introduce TriProbe, a multi-level probing framework that explains the root causes of task separability difficulties by tracing them across successive stages of learning. The central premise of TriProbe is that class separability may either deteriorate or improve throughout different stages of the learning pipeline. Understanding these dynamics requires not only identifying where representational bottlenecks arise but also determining which task pairs are most affected. TriProbe decomposes a multi-task classification problem into a set of binary sub-problems and systematically analyzes separability across the processing hierarchy. The framework first examines the discriminative potential of the original input data and handcrafted feature representations through a Foundational Probe. It then evaluates the quality of the learned latent representations produced by the feature extractor using a Latent Probe. Finally, a Final Probe assesses the separability achieved in the network’s output space, enabling a comprehensive diagnosis of how discriminative information evolves from the input to the final decision layer. The contributions of this work are threefold. First, we propose TriProbe as a stage-consistent diagnostic framework for explainable evaluation of task separability across learning stages. Second, we show how it provides actionable insights by diagnosing task difficulty, localizing bottlenecks, and revealing raw-level data complexity that informs both collection strategies and interpretation of results. Third, we demonstrate TriProbe on an sEMG gesture recognition task, where it uncovers bottlenecks consistent with downstream performance and prior studies, while exposing how separability evolves across stages.

Foundational Probe

Raw Data Feature Extractor

Final Probe

Latent Probe

Body Final

Actual

Evaluation

Predicted

Fig. 1: TriProbe: a framework for explainable separability analysis across learning stages.

II. P ROPOSED MULTI - LEVEL PROBING and the overall score is F1(ci ) = maxk F1k (ci ). For pairwise analysis between two classes (ci , cj ), the same formulation This work introduces TriProbe, a multi-level probing frameapplies by replacing the “rest” statistics with those of class work for diagnosing the root causes of task separability c . Higher F1 values indicate stronger separability, while low j difficulty across successive stages of learning. The key idea values suggest intrinsic overlap. This metric is particularly useis that separability may degrade or improve at different ful because it highlights the single most discriminative feature points in the pipeline, so identifying bottlenecks requires both dimension, allowing us to assess whether poor separability localizing the stage and pinpointing the most affected class stems from intrinsic data overlap at the raw level. pairs. TriProbe decomposes a multi-task problem into binary sub-tasks, preventing models from exploiting indirect cues from B. Latent Probe unrelated classes and ensuring that analysis reflects the intrinsic This probe evaluates task separability in the learned repredifficulty of each pair. As illustrated in Figure 1, TriProbe sentation space. The feature extractor can be any contemporary employs three complementary probes : (i) a Foundational Probe architecture, which makes the approach general. In this work, to assess raw data and hand-crafted features, (ii) a Latent we employ an autoencoder (AE) and use the encoder output Probe to evaluate feature extractor representations, and (iii) a after ensuring satisfactory reconstruction quality, while discardFinal Probe to analyze final-level separability. By combining ing the decoder. This choice avoids bias from downstream task binary decomposition with multi-level probing, TriProbe offers objectives such as classification, while still capturing rich dataa principled and interpretable diagnosis of whether limitations driven features. We train a dedicated AE for each class, enabling arise from data, representation, or final stages. fine-grained diagnosis and preventing the AE to exploit indirect A. Foundational Probe cues from unrelated classes during reconstruction. This ensures Let X ∈ Rn×m denote the raw multi-channel input data, that the encoder learns features solely from its own class where n is the number of time samples and m the number of distribution, thereby preserving class-specific representation input channels. For each class ci , let Xci ∈ Rni ×m represent quality. Concretely, the encoder maps each input sample x into the set of raw samples belonging to that class, with ni denoting a latent feature vector f ∈ Rd . Collecting samples of class the number of samples. In addition to the raw space, we also ci forms a feature matrix Fci ∈ Rni ×d , and similarly Fcj consider a hand-crafted feature space. A feature extraction for class cj , which are then compared to measure pairwise function Φ(·) maps each raw input x ∈ Rn×m into a feature separability in the latent space. vector f = Φ(x) ∈ Rd , where d is the number of engineered features (e.g., entropy measures, zero-crossings, RMS, etc.). C. Final Probe Collecting across all samples of class ci yields the feature This probe is placed immediately before the final separation matrix Fci ∈ Rni ×d . stage (e.g., the last classifier layer). It aims to evaluate how Maximum Fisher’s Discriminant Ratio (F1): A central effectively the body network prepares discriminative representaanalytical tool in our framework is Maximum Fisher’s Discrim- tions for the final decision. To maintain generality, we employ inant Ratio (F1), drawn from data complexity analysis [14]. a simple multilayer perceptron (MLP) as the classifier body, This metric quantifies separability between two distributions positioned after AE’s latent representation (encoder output). along individual feature dimensions and identifies the single The MLP is chosen deliberately as a generic architecture, free most discriminative feature for distinguishing a target class from domain-specific inductive biases. Therefore, the probe from a reference. In the one-vs-rest setting, where a target class reflects the intrinsic quality of the representations rather than ci is compared against all others, the Fisher score of feature k advantages of a specialized classifier. For each input sample, is defined as the probe captures the representation vector at the penultimate (µi,k − µrest,k )2 layer (i.e., the input to the final layer). These vectors form F1k (ci ) = , 2 + σ2 the basis for evaluating task separability at the separation σi,k rest,k

boundary. This enables assessment of whether difficulty arises from insufficiently discriminative representations at the body level, as opposed to intrinsic data limitations (Foundational Probe) or bottlenecks in feature extraction (Latent Probe).

simple multilayer perceptron (MLP). To avoid task-objective bias, the AE was kept frozen (decoder discarded, encoder weights fixed), and only the MLP was trained. The initial MLP layers are treated as the body, while the last layer defines the final separator, with the probe placed at their interface. This D. Interpretation design preserves architectural neutrality while enabling the TriProbe provides a stage-wise view of how task sepa- probe to identify whether separability limitations arise from rability evolves through the learning pipeline. Low scores insufficient body representations or from the final decision rule. at the Foundational Probe signal intrinsic data complexity, while discrepancies between later probes expose architectural B. Results limitations. A key guideline is that if separability is strong in Figure 2 illustrates the application of TriProbe on the the latent space but weak at the final stage, the issue lies in the Roshambo dataset. We first validate the use of F1 for task intervening architecture rather than the data. Thus, TriProbe separability evaluation and show its consistency with both not only diagnoses task difficulty via binary decomposition downstream performance and prior results on this dataset. This and localizes bottlenecks, but also informs data quality, guides establishes the analytical core of TriProbe as a robust evaluator, architectural refinement, and supports the rational interpretation capable of supporting multiple interpretations when deployed of results. By pinpointing where separability is lost or preserved, at different stages of the learning pipeline. Second, we report TriProbe advances the goals of Explainable AI with a practical the experiment specific findings. tool for both analysis and design. F1 Validation: F1 applied at the Foundational probe shows that the P–S sub-task has the weakest separability, with III. E XPERIMENT substantial overlap in both raw data and hand-crafted feature A. Setup assessments. In contrast, R–S and R–P consistently display We evaluate the proposed TriProbe framework on the stronger separability. A closer look reveals a minor discrepancy: Roshambo dataset [15], a benchmark known for its low raw data suggests R–P is slightly easier than R–S, while handsignal-to-noise ratio. The dataset contains recordings from ten crafted features suggest the opposite. At the Latent probe, participants using a Myo armband with eight sEMG channels F1 results align with the raw-level findings, and the Final sampled at 200 Hz. Participants performed three gestures — probe largely preserves this pattern, except for a noticeable Rock (R), Paper (P ), Scissors (S) — plus a control (Rest). degradation in R–S. Confusion matrices further confirm that Each participant completed three sessions, with five trials of P–S is the hardest pair to discriminate, consistent with both our 3 seconds (s) per gesture, yielding 450 trials in total. To isolate probes and prior results from [16] (see Fig. 3). Interestingly, steady-state activity, the first and last 600 ms of each trial were their two model variants diverge slightly: one favors R–S, the discarded. A sliding window of 400 samples (2 s) with 50% other R–P, echoing our observation that such small gaps are overlap was then applied, yielding 296 samples per class of sensitive to the feature space on which the learner operates. Two key insights emerge: (i) inherent data complexity size 8×400 (channels × time). Foundational Probe: Each 8 × 400 sample was transformed observed at the raw level propagates through the pipeline, into a 72-D feature vector (8 channels × 9 features) using meaning downstream stages should not contradict these trends, a modality-agnostic set of hand-crafted features including and (ii) F1 can flag sub-task separability difficulties early, even Shannon entropy, sample entropy, zero crossings, waveform at the data collection stage, making it a practical diagnostic length, root mean square (RMS), slope sign changes, median tool. Taken together, these results validate F1 as an effective frequency, wavelet energy, and fractal dimension. This yields probe of task separability, allowing researchers to peak into Fci ∈ R296×72 per class, capturing temporal, spectral, and the black box. complexity characteristics while avoiding domain-specific bias. TriProbe Analysis: With F1 validated as a reliable separability evaluator, we now leverage it to analyze different stages of Latent Probe: Three class-specific AE (R-only, P-only, the learning pipeline. By decomposing the multi-task problem S-only) were trained to construct latent spaces for pairwise into binary sub-tasks, TriProbe enables pairwise separability separability analysis. When referring to an AE in the experi- analysis and reveals which sub-tasks act as bottlenecks from ments, we specifically use the architecture detailed in Table I. raw data through to downstream performance. Our results The encoder output, once achieving satisfactory reconstruction consistently indicate that separability is most challenging quality, was taken as the latent representation. Each class ci for P–S, as reflected in both our binary confusion matrices produces Fci ∈ R296×392 , with 392 being the latent dimension. and prior multi-task evaluations in [16]. This shows that Training used 300 epochs with Huber loss (δ = 0.25), Adam TriProbe can diagnose the root cause of multi-task difficulty optimizer (lr = 5×10−4 , weight decay 10−5 ), and all available already at the raw data level. Accordingly, P–S emerges as the samples, aiming the robust representation learning rather than primary bottleneck pair, drawing attention to where learning downstream classification. pipeline design should be strengthened, especially in multi-task Final Probe: For each binary sub-task (R–P, R–S, P–S), we settings. Beyond bottleneck detection, TriProbe also reveals trained a pairwise AE whose encoder output was fed into a how separability evolves across stages, supporting architectural

TABLE I: The utilized AE. k: kernel, s: stride, BN: BatchNorm. Layer Input Conv1 Conv2 Conv3 Conv4 Latent Deconv1 Deconv2 Deconv3 Deconv4 Output

Raw Data

Config – 1 → 128, k = 3×3, s = 1, ReLU, BN, Dropout 128 → 256, k = 3×3, s = 1, Leaky-ReLU, BN, Dropout 256 → 512, k = 3×3, s = 1, ELU, BN, Dropout 512 → 1, k = 2×3, s = 1 – 1 → 512, k = 2×3, s = 1, ReLU, BN, Dropout 512 → 256, k = 3×3, s = 1, Softplus, BN, Dropout 256 → 128, k = 3×3, s = 1, ELU, BN, Dropout 128 → 1, k = 3×3, s = 1 –

Feature Extractor

Body

Final

Evaluation P 79

36

R 119

7

R 110

8

S 28

93

S 24

86

P 14

104

P

S

R

S

R

P

Actual

Rock-Paper Rock-Scissors Paper-Scissors

Output 1 × 8 × 400 128 × 6 × 398 256 × 4 × 396 512 × 2 × 394 1 × 1 × 392 392 512 × 2 × 394 256 × 4 × 396 128 × 6 × 398 1 × 8 × 400 1 × 8 × 400

Latent

Foundational Probe

Predicted Latent Probe

Predicted

Predicted

Final Probe

Fig. 2: Experimenting the proposed TriProbe framework on the Roshambo dataset, an sEMG-based gesture recognition task with three classes (Rock, Paper, and Scissors). The framework decomposes the problem into binary sub-tasks (Rock–Paper, Rock–Scissors, Paper–Scissors) and evaluates separability across three levels: (i) Foundational Probe measuring raw data and hand-crafted feature discriminability using F1; (ii) Latent Probe assessing latent representations with F1 and visualized using UMAP scatter plots, with covariance ellipses summarizing class spread; and (iii) Final Probe analyzing separability at the penultimate classifier layer. Evaluation confusion matrices are shown for each binary sub-task.

Predicted

R 0.940.040.02 P 0.000.830.17 0.5 S 0.000.120.88 R P S 0.0

Actual

Actual

R 0.950.000.04 P 0.010.740.25 S 0.000.150.85 R P S

Predicted

Fig. 3: Normalized confusion matrices of two Roshambo model variants from [16].

search. For instance, R–S shows good separation in the Latent probe but degrades at the Final probe, suggesting ineffective learning in the MLP body. Redesigning this stage could mitigate the loss. Similarly, P–S not only appears as the hardest pair overall but also undergoes degradation from Foundational to Latent probe, underscoring the need for stronger feature extraction. Thus, TriProbe not only diagnoses bottlenecks but also tracks the propagation of separability across the pipeline, enabling principled evaluation of whether architectural modifications improve or degrade performance relative to a baseline.

IV. D ISCUSSION TriProbe provides a structured methodology for diagnosing task separability across the full learning pipeline. Unlike existing probing approaches that focus on isolated representations, TriProbe localizes where separability is gained or lost throughout the entire learning pipeline. While validated here on an sEMG gesture recognition dataset, its formulation is architecture-agnostic and task-independent, suggesting broad applicability across domains such as speech, EEG/MEG, and sensor networks. While TriProbe builds upon established components such as Fisher’s Discriminant Ratio, autoencoders, and probing concepts, its novelty lies in their integration into a unified diagnostic framework for explainable task separability analysis. Rather than introducing a new classifier or separability metric, TriProbe provides a stage-consistent methodology that systematically traces how discriminative information evolves from the raw input space, through learned latent representations, to the final decision space. By combining binary task decomposition with a common separability criterion across all stages, the framework enables direct localization of representational bottlenecks and distinguishes whether limitations originate from the data itself, the feature extractor, or the downstream classifier. This unified perspective is not provided by existing probing or attribution

methods, which typically analyze only a single stage or explain individual predictions. A promising use case for TriProbe is early-stage evaluation during data collection. By probing separability at the raw and hand-crafted feature level, TriProbe serves as an early warning system, allowing pilot studies and assess data quality before committing to large-scale collection or model training. This is particularly valuable in domains where acquisition is costly. Another application lies in verifying the rationality of experimental findings. If probe results at early stages contradict downstream model performance, TriProbe can highlight potential issues such as overfitting, data leakage, or ineffective architectural choices. In this way, TriProbe complements traditional performance metrics by providing an interpretable explanation of where separability is lost or preserved across stages. Beyond these general applications, our findings also suggest dataset-specific improvements. For example, the observed degradation of the R–S pair between the latent and final probes indicates that the current MLP body may be ineffective, motivating experiments with alternative architectures. Such targeted refinements illustrate how TriProbe can guide iterative design by localizing weaknesses in the pipeline. Extending this idea, future work should evaluate TriProbe across different datasets and modalities, not only to confirm its generalizability but also to explore how separability evolves under varying data complexities and learning architectures. V. C ONCLUSION In conclusion, TriProbe contributes to Explainable AI by providing a unified framework for tracing the root causes of task separability difficulty across raw data, learned representations, and decision spaces. Rather than relying solely on downstream performance metrics, it offers stage-wise insights into how discriminative information evolves throughout the learning pipeline, enabling the localization of bottlenecks and the identification of challenging class pairs. Consequently, TriProbe serves not only as an early diagnostic tool for assessing data quality, but also as a means of validating experimental findings and guiding architecture refinement through interpretable separability analysis. Although we demonstrated TriProbe here on an sEMG gesture recognition task, the framework is designed to be architecture-agnostic and readily applicable to other machine learning domains. Future work will evaluate TriProbe across additional modalities and datasets, investigate alternative separability measures, and integrate the framework into adaptive learning pipelines, with the long-term goal of establishing a general-purpose, separability-driven methodology for interpretable and reliable AI. R EFERENCES [1] Shimon Fridkin and Michael Bendersky, “Interpretable machine learning: A comprehensive review of foundations, methods, and the path forward,” WIREs Data Mining and Knowledge Discovery, vol. 16, no. 1, pp. e70075, 2026. [2] Ričards Marcinkevičs and Julia E. Vogt, “Interpretable and explainable machine learning: A methods-centric overview with concrete examples,” WIREs Data Mining and Knowledge Discovery, vol. 13, no. 3, pp. e1493, 2023.

[3] Mateo Espinosa Zarlenga, Pietro Barbiero, Gabriele Ciravegna, Giuseppe Marra, Francesco Giannini, Michelangelo Diligenti, Zohreh Shams, Frederic Precioso, Stefano Melacci, Adrian Weller, Pietro Lio, and Mateja Jamnik, “Concept embedding models: beyond the accuracy-explainability trade-off,” in Proceedings of the 36th International Conference on Neural Information Processing Systems, Red Hook, NY, USA, 2022, NIPS ’22, Curran Associates Inc. [4] Pantelis Linardatos, Vasilis Papastefanopoulos, and Sotiris Kotsiantis, “Explainable AI: A review of machine learning interpretability methods,” Entropy, vol. 23, no. 1, 2021. [5] Kim Huat Goh, Le Wang, Adrian Yong Kwang Yeow, Hermione Poh, Ke Li, Joannas Jie Lin Yeow, and Gamaliel Yu Heng Tan, “Artificial intelligence in sepsis early prediction and diagnosis using unstructured data in healthcare,” Nature Communications, vol. 12, no. 1, pp. 711, 2021. [6] Meike Nauta, Jan Trienes, Shreyasi Pathak, Elisa Nguyen, Michelle Peters, Yasmin Schmitt, Jörg Schlötterer, Maurice van Keulen, and Christin Seifert, “From anecdotal evidence to quantitative evaluation methods: A systematic review on evaluating explainable AI,” ACM Comput. Surv., vol. 55, no. 13s, July 2023. [7] Guillaume Alain and Yoshua Bengio, “Understanding intermediate layers using linear classifier probes,” arXiv preprint arXiv:1610.01644, 2016. [8] John Hewitt and Christopher D Manning, “A structural probe for finding syntax in word representations,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 2019, pp. 4129–4138. [9] Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni, “What you can cram into a single vector: Probing sentence embeddings for linguistic properties,” 2018. [10] Mukund Sundararajan and Amir Najmi, “The many shapley values for model explanation,” in International conference on machine learning. PMLR, 2020, pp. 9269–9278. [11] Mukund Sundararajan, Ankur Taly, and Qiqi Yan, “Axiomatic attribution for deep networks,” in International conference on machine learning. PMLR, 2017, pp. 3319–3328. [12] Ulysse Côté-Allard, Cheikh Latyr Fall, Alexandre Drouin, Alexandre Campeau-Lecours, Clément Gosselin, Kyrre Glette, François Laviolette, and Benoit Gosselin, “Deep learning for electromyographic hand gesture signal classification using transfer learning,” IEEE Transactions on Neural Systems and Rehabilitation Engineering, vol. 27, no. 4, pp. 760–771, 2019. [13] Mr. Amol Pandurang Yadav and Dr. Sandip.R. Patil, ““optimizing semg gesture recognition with stacked autoencoder neural network for bionic hand”,” MethodsX, vol. 14, pp. 103207, 2025. [14] Tin Kam Ho and M. Basu, “Complexity measures of supervised classification problems,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 24, no. 3, pp. 289–300, 2002. [15] Elisa Donati, “EMG from forearm datasets for hand gestures recognition,” May 2019. [16] Nikhil Garg, Ismael Balafrej, Yann Beilliard, Dominique Drouin, Fabien Alibart, and Jean Rouat, “Signals to Spikes for Neuromorphic Regulated Reservoir Computing and EMG Hand Gesture Recognition,” in International Conference on Neuromorphic Systems 2021. 2021, Association for Computing Machinery.

Record · ID 965470 · SHA-256 57bd87ad7828e429
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.