Conceptio › Archive › arXiv CS
arXiv CSopen access

Rethinking XAI Evaluation: A Human-Centered Audit of Shapley Benchmarks in High-Stakes Settings

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

arXiv:2604.22662v1 [cs.LG] 24 Apr 2026

Rethinking XAI Evaluation: A Human-Centered Audit of Shapley Benchmarks in High-Stakes Settings Inês Oliveira e Silva

Sérgio Jesus

Iker Perez

Rita P. Ribeiro

University of Porto, Feedzai Porto, Portugal [email protected]

Feedzai Porto, Portugal [email protected]

Feedzai London, UK [email protected]

University of Porto Porto, Portugal [email protected]

Carlos Soares

Hugo Ferreira

Pedro Bizarro

University of Porto Porto, Portugal [email protected]

Feedzai Lisbon, Portugal [email protected]

Feedzai Lisbon, Portugal [email protected]

Abstract Shapley values are a cornerstone of explainable AI, yet their proliferation into competing formulations has created a fragmented landscape with little consensus on practical deployment. While theoretical differences are well-documented, evaluation remains reliant on quantitative proxies whose alignment with human utility is unverified. In this work, we use a unified amortized framework to isolate semantic differences between eight Shapley variants under the low-latency constraints of operational risk workflows. We conduct a large-scale empirical evaluation across four risk datasets and a realistic fraud-detection environment involving professional analysts and 3,735 case reviews. Our results reveal a fundamental misalignment: standard quantitative metrics, such as sparsity and faithfulness, are decoupled from human-perceived clarity and decision utility. Furthermore, while no formulation improved objective analyst performance, explanations consistently increased decision confidence, signaling a critical risk of automation bias in high-stakes settings. These findings suggest that current evaluation proxies are insufficient for predicting downstream human impact, and we provide evidence-based guidance for selecting formulations and metrics in operational decision systems.

Keywords Explainable AI (XAI), Shapley values, Feature attribution, Evaluation methodology, Human-in-the-loop, Automation bias

1

Introduction

Machine learning (ML) systems are increasingly deployed in highstakes domains such as fraud detection, credit assessment, and healthcare, where predictions have direct human and financial consequences [32, 37]. In these settings, model outputs rarely constitute final decisions. Instead, predictions are reviewed by human decision-makers operating under time, attention, and regulatory constraints. As a result, explanations are viewed as indispensable for accountability and oversight, and have become a core requirement in operational ML deployments [8, 9, 47]. Yet, despite widespread adoption, the practical value of explanations in human-in-the-loop workflows remains poorly understood, often assumed rather than empirically established [6, 8, 27]. Among explanation methods, local approaches grounded in cooperative game theory, most notably Shapley values, have emerged

as the de facto standard for feature attribution [5], by providing an axiomatic decomposition of model predictions into feature-level contributions [34]. However, the framework has fragmented into competing formulations based on divergent assumptions about the semantics of feature absence, realized in popular implementations such as KernelSHAP, TreeSHAP, and related tools [4, 11, 14, 33, 43, 58, 60]. This raises a critical evaluation question for practitioners: does the choice of formulation matter to the end-user, and do standard evaluation procedures anticipate the impact? Modern XAI evaluation relies on theoretical analysis and quantitative proxies [11, 15, 35], as well as mathematical distinctions between “faithfulness” to the model or “truthfulness” to the data [12]. Yet, systematic evidence regarding how explanation methods perform against human-centered benchmarks remains scarce. Existing evaluations often focus on isolated functional (model-based) properties and rarely stress-test these metrics against human behavior under realistic operational constraints [10, 41]. Furthermore, comparisons are often confounded by implementation choices, which mask the true semantic differences of the definitions themselves. In this work, we treat XAI evaluation as a scientific object of study. We use Shapley values as a representative class of feature attribution, and audit the alignment between quantitative evaluation benchmarks and human utility through a comprehensive study of 3,735 instances of human-AI interaction. We use a unified amortized framework to eliminate implementation confounders, and instantiate eight Shapley definitions under identical computational conditions. Our scope reflects high-throughput, high-stakes tabular data environments in low-latency financial, healthcare, e-commerce, or logistics deployments [23, 49, 55]. We conduct this audit across 5 datasets in a real production-grade fraud-detection environment involving 37 participants, including professional analysts. Our analysis reveals a fundamental decoupling between metrics and utility: standard quantitative proxies are poor predictors of human-perceived clarity or decision confidence. Crucially, while explanations fail to improve objective analyst performance, they consistently boost confidence. This exposes a major deployment risk: current evaluation practices may favor explanations that encourage automation bias without improving decision quality. In Figure 1 we provide an overview of the landscape explored in this paper. Our contributions are:

Oliveira e Silva et al.

Insertion AUC

0.90

Better Referenced Empirical Filt. Cond. Counterfactual Bs. Zero

0.85

Bs. Mean

0.80

Cond. Countf. Marginal 0.75 Uniform Jt. Marg. Worse 0.70 6 5 4 3 Sparsity Ranking

1.8

User Confidence-Clarity Trade-off Better Jt. Marg. Countf.

Marginal

1.0 Uniform

0.6

Marginal Cond. Jt. Marg.

Cond.

Bs. Mean

Bs. Mean Countf.

Filt. Cond.

Uniform Filt. Cond.

Bs. Zero

Worse

2

Cross Agreement Between Formulations Average Spearman Correlation

0.95

Insertion AUC-Sparsity Trade-off

Clarity (Odds Ratio)

1.00

1.0 1.3 Confidence (Odds Ratio)

1.8

Bs. Zero

M. C. J.M. B.M. CF. U. F.C. B.Z.

1

0

-1

Figure 1: Overview of the empirical XAI audit in this paper. Left: Quantitative benchmarks reveal systematic functional variation (e.g., sparsity, faithfulness) across Shapley formulations. Center: Controlled analyst studies show differences in perceived clarity and decision confidence. Right: Pairwise agreement between Shapley formulations reveals structured alignment. • In Section 2, we describe a unified framework to isolate semantic explanation effects of Shapley formulations, eliminating algorithmic confounders for fair comparison. • In Sections 3-5, we demonstrate through expert-driven reviews that common XAI evaluation metrics fail to align with human decision-making outcomes. • In Section 6, we document a critical failure mode: the promotion of decision confidence without accuracy gains. We further release 3,735 granular human-AI interaction measurements to support the development of behaviorally grounded XAI benchmarks.

2

Shapley Value Foundations

Shapley values [50] provide an additive decomposition of a model’s output 𝑓 : X ↦→ R (e.g., risk scores or regression estimates) into feature-level contributions. Their axiomatic foundation and status as de facto standard for tabular applications is ideal for auditing alignment between XAI evaluation proxies and human utility. Let F = {1, . . . , 𝑑 } denote the full feature set and S ⊆ F a coalition of “present” features. For an input instance 𝒙 ∈ X, a value function 𝑣 𝒙 (S) specifies the expected model output when only features in S are observed:   𝑣 𝒙 (S) = E𝑿 F\S ∼𝑝 (·) 𝑓 (𝒙 S , 𝑿 F\S ) , (1) where the background distribution 𝑝 (·) defines the semantics of feature absence. The Shapley contribution of feature 𝑖 to a model output is defined as ∑︁ |S|!(𝑑 − |S| − 1)!   𝜙𝑖 = 𝑣 𝒙 (S ∪ {𝑖}) − 𝑣 𝒙 (S) , (2) 𝑑! S ⊆ F\{𝑖 }

which satisfies efficiency, symmetry, and dummy axioms [34, 53]. A Taxonomy of Semantic Variants. Equation (1) highlights that Shapley values are a family of definitions induced by the choice of background distribution 𝑝 (·). This choice reflects the tension between being “true to the model” versus “true to the data” [12]

and can produce divergent attribution patterns [4, 16, 26, 33] with direct implications for human-in-the-loop workflows [11]. While causal [24] or in-manifold [54] variants also exist, they require generative models or assumptions often incompatible with the low-latency, high-throughput requirements of production tabular systems. We therefore focus our audit on computationally feasible variants prevalent in production settings, summarized in Table 1.

2.1

Amortization as an Experimental Control

Exact computation of Equation (2) requires evaluating 2𝑑 feature coalitions, which is intractable for all but very small 𝑑. The SHAP framework [34] addresses this by recasting Shapley estimation as fitting an additive surrogate model ∑︁ 𝑔(𝑆) = 𝑣 𝒙 (∅) + 𝜙𝑖 , 𝑖 ∈𝑆

subject to an efficiency constraint 𝑓 (𝒙) = 𝑔(F ). In this setting, KernelSHAP [34] estimates 𝜙𝑖 for a given instance 𝒙 ∈ X by solving a weighted least-squares regression over sampled coalitions: ∑︁ 2 arg min 𝜋 (𝑆) 𝑣 𝒙 (𝑆) − 𝑔(𝑆) , (3) 𝜙 1 ,...,𝜙𝑑 𝑆 ⊆ F

with kernel weights 𝜋 (𝑆) =

𝑑 −1 𝑑  |𝑆 | |𝑆 |(𝑑 − |𝑆 |)

emphasizing small and large coalitions. In practice, Monte Carlo approximations of Equation (3) introduce variance and bias [22]. Consequently, a fragmented landscape of model-specific heuristic optimizations has emerged [33, 40, 58, 59], which often confound comparisons between Shapley definitions with algorithmic noise. To isolate semantic effects from implementation artifacts, we utilize surrogates or amortizers [13, 29] as a unified experimental control. Instead of solving (3) independently for each 𝒙, a parametric universal approximator 𝜙ˆ𝜃 (𝒙) is trained to minimize the expected

A Human-Centered Audit of Shapley Value Benchmarks in High-Stakes Workflows

Table 1: Taxonomy of Shapley value formulations categorized by the representation of feature absence. Category

Definition

Sampling Distribution

Semantics / Assumptions

Referenced

Fixed Baseline

𝛿 (𝑿 F\S = 𝒙 ′F\S )

Absent features fixed to a reference; efficient but highly sensitive to baseline selection.

Uniform

𝑢 (𝑿 F\S )

Samples uniformly from a hyperbox; ignores data distribution and feature correlations.

Marginal

𝑝 (𝑿 F\S )

Samples absent features from joint empirical distribution; preserves marginals but breaks feature dependencies with 𝒙 S .

Joint-Marginal

Î

Samples from univariate marginals independently; removes all feature dependencies.

Conditional

𝑝 (𝑿 F\S | 𝑿 S = 𝒙 S )

Search Counterfactual

) ) 𝑝 (𝑿 (𝑐 | 𝑓 (𝑿 (𝑐 ) ≈𝑦 ∗ ) F\S F

Filtered Conditional

𝑝 (𝑿 F\S | 𝑓 (𝑿 F ) ∈ Y)

Empirical

Counterfactual

𝑗 ∈ F\S 𝑝 (𝑋 𝑗 )

attribution loss over the data distribution: " # ∑︁ 2 min E𝑿 𝜋 (𝑆) 𝑣 𝑿 (𝑆) − 𝑔ˆ𝜃 (𝑆; 𝑿 ) , 𝜃

(4)

where ∑︁

Uses counterfactual perturbations as baselines; high variance and optimization-intensive. Conditions on background samples with specific model outputs; blends attribution with contrastive logic.

3.1

𝑆⊆F

𝑔ˆ𝜃 (𝑆; 𝒙) = 𝑣 𝒙 (∅) +

Preserves empirical dependencies via conditioning; difficult to estimate in high-dimensional settings.

𝜙ˆ𝑖,𝜃 (𝒙)

Evaluation Metrics and Proxies

We use a compact set of metrics to capture functional properties, cross-formulation agreement, and downstream analyst behavior. Metrics are defined per instance 𝒙, and grounded in established notions such as sensitivity, rank agreement, uncertainty, or sparsity. Results are aggregated across datasets, models, and experimental repetitions.

𝑖 ∈𝑆

Í is an additive model satisfying 𝑖 𝜙ˆ𝑖,𝜃 (𝒙) = 𝑓 (𝒙) − 𝑣 𝒙 (∅). This reduces inference to a single forward pass, enabling the millisecondlevel Service Level Agreements (SLAs) required in production risk systems [13, 19, 36]. Crucially, this separation of computational mechanics from theoretical semantics enables like-for-like comparisons across Shapley formulations under identical conditions.

3

Auditing XAI Metric Alignment

Prior research has established broad desiderata for explanation quality [18, 48, 56], converging on functionally grounded properties that characterize faithfulness to the predictive model, robustness to perturbations, selectivity, or truthfulness (see Appendix B). However, these properties are frequently treated as ends in themselves [10]. In this work, we treat such metrics as quantitative proxies subject to audit, and seek to explore their alignment with actual human utility in high-stakes workflows. We evaluate Shapley attributions along two complementary axes: (i) model-based quantitative proxies, and (ii) utility in a large-scale human-in-the-loop study. To isolate the effect of the underlying Shapley formulation as the primary independent variable, we hold the explanation interfaces and interaction mechanisms strictly fixed. Our study spans classification tasks on benchmark datasets [2, 17, 20, 25] and real-world financial transactions reviewed by professional analysts. Non-proprietary code and data are released for reproducibility.1 1 https://github.com/feedzai/SHAP-Value-Function-Evaluation

3.1.1 Quantitative Evaluation. We evaluate whether attributions exhibit behaviors theoretically associated with faithful and usable explanations, independent of human judgment. Deletion AUC assesses faithfulness by measuring how predictive uncertainty increases as features are removed in order of importance [44]. Let 𝒙 (𝑘 ) denote an input with top-𝑘 features removed, and 𝐻 (·) the predictive entropy. Faithful explanations identify features whose removal quickly degrades model confidence: 1 ∑︁ 𝐻 (𝒙 (𝑘 ) ) − 𝐻 (𝒙 (0) ) . 𝑑 𝐻 (𝒙 (𝑑 ) ) − 𝐻 (𝒙 (0) ) 𝑑

1−

(5)

𝑘=1

An analogous Insertion AUC measures confidence gains as features are reintroduced. Perturbation Sensitivity measures local stability relative to model output changes under perturbations 𝜖 > 0:  ∥𝜙ˆ𝜃 (𝒙 + 𝜖) − 𝜙ˆ𝜃 (𝒙)∥ 2 [|𝑓 (𝒙 + 𝜖) − 𝑓 (𝒙)| + 𝛿] , (6) where 𝛿 > 0 ensures numerical stability. Lower values indicate smoother and robust explanations. Counterfactual Contrastivity evaluates if explanations meaningfully differ across perturbed instances 𝒙 ′ that cross the decision boundary:  (7) ∥𝜙ˆ𝜃 (𝒙) − 𝜙ˆ𝜃 (𝒙 ′ )∥ 2 [|𝑓 (𝒙) − 𝑓 (𝒙 ′ )| + 𝛿] . Higher values indicate explanations that better reflect decisionrelevant model changes.

Oliveira e Silva et al.

Sparsity (L1 /L2 ) quantifies the concentration of attribution mass across features, a property assumed to reduce cognitive load: ∥𝜙ˆ𝜃 (𝒙)∥ 1



∥𝜙ˆ𝜃 (𝒙)∥ 2 .

(8)

A lower ratio indicates selective attributions concentrated on fewer features. 3.1.2 Amortizer Alignment and Agreement. To validate our experimental control, we measure • Attribution Error: The mean squared error (MSE) between amortized 𝜙ˆ𝑗,𝜃 (𝒙) and a high-sample KernelSHAP ground truth 𝜙 𝑗 (𝒙); and • Recall@𝑘: The overlap across the top-𝑘 most influential features between amortized and ground truth attributions. In all cases, KernelSHAP references are tailored to each specific Shapley definition. Finally, we measure Cross-Agreement (Spearman correlation) between different Shapley formulations to quantify how the choice of definition alters the explanation rank-order. 3.1.3 Human-in-the-Loop Utility. To assess operational impact, we measure analyst behavior in tasks where model predictions are paired with different Shapley formulations. We record Decision ˆ and Decision Time (log(1 + 𝑡)) to measure effiAccuracy (I{𝑦 = 𝑦}) ciency. In addition, we collect self-reported Confidence and Clarity scores. By analyzing the gap between Confidence and Accuracy, we specifically audit for automation bias, where explanations increase trust without improving decision quality.

3.2

Shapley formulations (Table 1); and (iv) 37 analysts, including professional fraud-review experts. A total of 3735 case-review measurements are released publicly2 (details in Appendices C–E).

User Interface and Analyst Workflow

Experiments use a standardized interface (Figure 2) grounded in academic conventions [34, 57] and mirroring professional risk tools [27]. It presents a model score/percentile, a Shapley attribution bar chart, and natural-language reason codes derived from the attributions. Visual encodings and interaction affordances are identical across all experimental conditions to ensure observed differences arise solely from the content of the Shapley formulation. Further interface details are discussed in Appendix E. Analysts submit a binary decision (risk/no risk) and self-report confidence (weak/moderate/strong) and clarity (confusing/clear). Each review produces a structured log of analyst metadata, decision outcome, confidence, timestamps, case features, model predictions, and the Shapley formulation shown. We utilize two deployments: an open-source frontend for benchmarks and a sandboxed system integrated into real-world fraud review pipelines.

4.1

To enable meaningful comparison across heterogeneous tasks, we distinguish between scale-invariant and scale-dependent metrics. Scale-invariant metrics (e.g., Deletion AUC, Recall@𝑘) are averaged directly across dataset–model pairs as they share a consistent interpretation. Conversely, scale-dependent metrics (e.g., sensitivity, sparsity) are converted to relative explainer rankings within each pair before aggregation to ensure cross-task commensurability. All results are reported as point estimates with standard errors obtained via bootstrapping (50 subsamples per pair).

4.2

Methodology

To establish a rigorous foundation for our audit, we proceed in two stages: (1) a multi-dimensional evaluation of Shapley formulations across functional and behavioral axes, and (2) a synthesis testing whether quantitative proxies reliably predict human utility. Our evaluation spans an unusually large-scale experimental matrix: (i) 5 risk datasets, including Maternal [2], Credit [25], HELOC [20], Adult [17], and a proprietary real-world fraud dataset; (ii) 2 industry-standard predictive models in low-latency systems (Logistic Regression and LightGBM [31]); (iii) 8 distinct

Human-in-the-Loop Experimental Design

Experiments follow a blinded, randomized within-subjects design with a no-explanation control. Participants review cases balanced across operational decision thresholds. By holding interface layout and interaction constraints constant, we isolate the Shapley formulation’s semantic content as the primary behavioral driver. Inferential Modeling. To isolate explanation effects from confounders like case difficulty, analyst expertise, or repeated exposure, we employ mixed-effects modeling [21, 42, 45]. We fit a global model across datasets, predictive models, and analysts, to account for structured heterogeneity and maximize statistical power for generalizable inference. We model effects relative to a no-explanation baseline for accuracy (logistic), confidence (ordinal), and time (loglinear), directly auditing for automation bias where explanations inflate trust without improving decisions. Perceived clarity is modeled relative to the population mean, excluding no-explanation cases. Fixed effects include the Shapley formulation, model type, dataset, model entropy, prediction error, and analyst experience, with further controls for self-reported subject domain, ML, and Shapley familiarity. Results are reported as odds ratios or multiplicative effects with 95% confidence intervals (details in Appendix E). Synthesis. To bridge quantitative and qualitative metrics, we extend the mixed-effects framework at the case level. We use quantitative explanation properties (e.g., sparsity, AUC) as predictors for perceived clarity and confidence, applying the same rigorous confounder controls. Thus, we statistically determine if offline proxies are reliable predictors of human utility in operational settings.

5 4

Aggregation of Quantitative Metrics

Results

Table 2 summarizes quantitative proxies, while Table 3 reports behavioral outcomes from our analyst study. All formulations are evaluated under matched data, models, and computational budgets. Fixed formulations use zero or mean baselines, while the remainder formulations rely on background resampling; conditional variants use proximity-weighted sampling, and counterfactual formulations employ optimization-based generation using DiCE [39].

2 https://github.com/feedzai/SHAP-Value-Function-Evaluation/tree/main/data

A Human-Centered Audit of Shapley Value Benchmarks in High-Stakes Workflows

Figure 2: The case review interface integrates model scores, Shapley attributions, and contextual data summaries while holding visual encodings fixed across Shapley formulations.

5.1

Quantitative Benchmarks

The benchmarks in Table 2 reveal structured trade-offs across the Shapley landscape. No single formulation dominates across all metrics. High sparsity and contrastivity improve insertion AUC, reflecting selective attributions, but increase sensitivity to small input perturbations. Conversely, low-sensitivity formulations are stable and perform better on deletion tasks, but produce less selective explanations. Across formulations, amortization achieves high fidelity to the KernelSHAP ground truth reference, achieving low error and high

Recall@𝑘 for top-ranked features. Thus, observed properties and differences reflect semantic formulation choices rather than approximation artifacts. Also, cross-alignment analysis in Figure 1 shows that formulations cluster into families. Key trends include: • Fixed baselines: The Zero baseline yields the strongest Deletion AUC and Recall@𝑘, but remains unaligned with empirical formulations. While the Uniform variant trades stability for reduced sparsity and contrastivity. The Mean baseline performs poorly across both both stability and selectivity metrics.

Oliveira e Silva et al.

Table 2: Average quantitative results across scale-invariant and scale-dependent metrics (bootstrapped standard deviations). Scale-dependent metrics use relative Rank statistics. Bold indicates the best outcome per metric.

Value F.

Rank Scale Dependent Metrics Sparsity

Sensitivity

Contrast.

Attr. Error

Scale Invariant Metrics Del. AUC

Ins. AUC

Recall@1

Recall@3

Recall@5

Base. Zero 3.97 ± 2.74 4.10 ± 2.86 4.40 ± 2.99 3.41 ± 2.49 0.11 ± 0.08 0.88 ± 0.20 0.90 ± 0.07 0.90 ± 0.06 0.91 ± 0.05 Base. Mean 4.39 ± 2.39 5.21 ± 1.84 4.65 ± 1.56 3.46 ± 2.15 0.19 ± 0.13 0.83 ± 0.13 0.83 ± 0.11 0.85 ± 0.06 0.87 ± 0.07 Uniform 5.25 ± 2.30 2.81 ± 1.73 6.09 ± 2.20 5.00 ± 1.88 0.15 ± 0.24 0.76 ± 0.14 0.86 ± 0.10 0.85 ± 0.06 0.86 ± 0.06 Marginal 4.40 ± 1.34 3.66 ± 1.57 4.14 ± 1.93 3.57 ± 1.17 0.16 ± 0.10 0.75 ± 0.09 0.80 ± 0.12 0.84 ± 0.07 0.86 ± 0.08 Joint-Marg. 4.98 ± 1.95 3.51 ± 1.50 5.38 ± 1.22 3.15 ± 1.28 0.16 ± 0.10 0.75 ± 0.09 0.82 ± 0.08 0.82 ± 0.07 0.84 ± 0.07 Conditional 5.68 ± 1.65 4.73 ± 1.66 3.92 ± 1.09 6.24 ± 1.52 0.16 ± 0.11 0.77 ± 0.09 0.63 ± 0.24 0.72 ± 0.10 0.78 ± 0.09 Counterf. 5.17 ± 2.23 5.21 ± 2.66 5.94 ± 1.97 7.46 ± 0.86 0.20 ± 0.08 0.77 ± 0.09 0.53 ± 0.10 0.63 ± 0.10 0.70 ± 0.13 Filt. Cond. 2.15 ± 1.31 6.77 ± 1.42 1.47 ± 0.69 3.73 ± 2.00 0.18 ± 0.08 0.95 ± 0.05 0.90 ± 0.08 0.88 ± 0.06 0.89 ± 0.06 Lower is better (↓)

• Empirical formulations: Marginal and Joint-Marginal variants show mutual agreement and the most balanced performance, i.e. low sensitivity, moderate sparsity and contrastivity. Conditional Shapley departs from this pattern, producing dense, sensitive attributions that reflect data correlations rather than model behavior [12]. • Counterfactual formulations: The Filtered Conditional variant maximizes sparsity, contrastivity, and Insertion AUC, but suffers from instability under perturbation. Fully counterfactual explanations are dense and unstable.

5.2

Human-in-the-Loop Utility

Analyst behavior results in Table 3 and Figure 3 reveal a critical asymmetry: formulation choice has no meaningful impact on objective performance (decision accuracy, review time) but strongly influences subjective outcomes (clarity, confidence). Explanations primarily shape how analysts interpret and commit to decisions rather than improving decision quality. Across formulations: • Objective performance: No formulation reliably improves accuracy or complex case (p97.5) decision time. Effects are driven by case-level factors such as model uncertainty, error or analyst exposure. Coupled with increased confidence, it highlights a severe risk of automation bias. • Perceived clarity is sensitive to formulation choice. JointMarginal explanations significantly improve metrics, while the Zero baseline, Uniform, and Filtered Conditional formulations are consistently rated as confusing. • Confidence shows strong explainer effects. Conditional, Counterfactual, and Mean formulations substantially inflate analyst confidence relative to no explanation. Ultimately, clarity and confidence remain constrained by core signals of case difficulty, i.e. model uncertainty and prediction error. Explanations modulate analyst perception, but do not override task difficulty, aligning with recent literature on XAI and complementary performance [7, 46, 51].

Higher is better (↑)

5.3

Alignment: Proxies vs. Human Outcomes

Finally, we audit the alignment between quantitative proxies and human outcomes by using functional properties as predictors for clarity and confidence (Figure 4). In all cases, we control for case difficulty, prediction error, analyst experience, and expertise. Results reveal a decoupling between quantitative metrics and perceived utility: standard criteria do not show positive associations with clarity or confidence. Notably, sparsity is negatively associated with confidence, contradicting the assumption that lowdimensional explanations enhance intelligibility [10]. Thus, widely used benchmarks capture structural properties but fail to predict human perception or use. While valuable for debugging, they are not substitutes for behavioral evaluation in operational settings.

Feedback

Accuracy

Confidence

Decision Time

Bs. Zero Bs. Mean Uniform Marginal Jt. Marg. Cond. Countf. Filt. Cond. Bs. Zero Bs. Mean Uniform Marginal Jt. Marg. Cond. Countf. Filt. Cond.

Performance

0.5

1.0

Clarity

1.8

0.5

1.0

1.8

Figure 3: Effect sizes on human outcomes, with 95% confidence intervals, across accuracy, decision time (𝑝 97.5 ), confidence and clarity. X-axis is in log-scale.

A Human-Centered Audit of Shapley Value Benchmarks in High-Stakes Workflows

Table 3: Fixed effects on human outcomes. Decision time is reported as multiplicative effects; Accuracy, Clarity, and Confidence as Odds Ratios. Standard errors retrieved via the Delta Method. Bold values indicate significance (𝑝 < 0.05). Decision Time (↓)

Value Func.

Odds Ratio (↑)

p2.5

p50

p97.5

Accuracy

Clarity

Confidence

Base. Zero Base. Mean Uniform Marginal Joint-Marg. Conditional Counterf. Filt. Cond.

1.29 ± 0.09 0.99 ± 0.07 1.11 ± 0.08 0.97 ± 0.07 1.17 ± 0.08 0.98 ± 0.07 1.08 ± 0.07 1.03 ± 0.07

1.16 ± 0.05 1.03 ± 0.04 1.11 ± 0.05 1.01 ± 0.04 1.06 ± 0.05 1.05 ± 0.04 1.00 ± 0.04 1.04 ± 0.05

1.03 ± 0.07 0.88 ± 0.06 1.02 ± 0.08 1.08 ± 0.08 0.99 ± 0.07 0.94 ± 0.07 1.15 ± 0.08 0.98 ± 0.07

1.04 ± 0.20 1.05 ± 0.21 0.84 ± 0.16 0.76 ± 0.14 1.19 ± 0.23 0.99 ± 0.19 1.02 ± 0.20 0.79 ± 0.15

0.57 ± 0.06 1.17 ± 0.14 0.79 ± 0.09 1.18 ± 0.14 1.52 ± 0.19 1.23 ± 0.14 1.25 ± 0.16 0.68 ± 0.08

1.20 ± 0.15 1.49 ± 0.19 1.02 ± 0.13 1.11 ± 0.14 1.20 ± 0.16 1.58 ± 0.20 1.55 ± 0.20 1.30 ± 0.16

Log Count Model Entropy Score Error Professional Analyst

0.84 ± 0.02 1.25 ± 0.03 1.00 ± 0.02 0.85 ± 0.05

0.79 ± 0.01 1.23 ± 0.02 1.00 ± 0.02 0.88 ± 0.04

0.80 ± 0.02 1.07 ± 0.03 0.99 ± 0.02 1.10 ± 0.07

0.93 ± 0.06 0.95 ± 0.08 0.19 ± 0.02 0.94 ± 0.17

1.01 ± 0.07 0.68 ± 0.05 0.86 ± 0.05 1.08 ± 0.17

0.94 ± 0.05 0.44 ± 0.05 0.80 ± 0.03 1.10 ± 0.13

6

Conclusion

This paper has conducted a unified large-scale empirical study and audit of Shapley value formulations, testing the alignment between functional quantitative proxies and human utility in high-stakes risk workflows. By evaluating eight semantic variants under strictly matched data, model, interfaces, and computational conditions, we proved fundamental decoupling in XAI evaluation: standard quantitative benchmarks capture structural properties of explanations but fail to predict how they are perceived or used by final users. Furthermore, while no Shapley formulation improved objective

Quantitative vs Qualitative Metrics Better

Confidence (Odds Ratio)

1.2

Contrastivity

Insertion AUC

Deletion AUC

Sparsity

Worse

0.8 0.8

1.0 Clarity (Odds Ratio)

Evidence-Based Guidance. The results have direct implications for both research and practice. Based on our audit, we offer the following guidance for the deployment and evaluation of XAI: • Decouple Algorithm from User: Practitioners must treat proxy metrics (e.g., deletion AUC, sensitivity) as debugging tools for model-faithfulness, not as proxies for human interpretability or trust. • System designers should prioritize empirical formulations for clarity (Marginal, Joint-Marginal, Conditional), while remaining wary of Conditional or Counterfactual variants that inflate confidence without commensurate gains in accuracy. Limitations and Scope. Our audit focused on low-latency, tabular risk models as a backbone to financial, e-commerce or medical decision-making [55]. However, results may differ in vision or language domains where feature semantics follow different dynamics. Additionally, our experiments were conducted in controlled settings and cannot capture longer-term effects such as learning, adaptation, or changes in institutional decision norms. Ultimately, our results motivate a shift in the XAI community toward behaviorally grounded evaluation; until metrics are proven to align with human outcomes, they cannot serve as substitutes for rigorous user studies in high-stakes production systems.

Sensitivity

1.0

accuracy, several significantly inflated analyst confidence. This uncovers a systemic risk of automation bias where explanations encourage over-reliance without providing cognitive gains.

1.2

Figure 4: Effect of quantitative explanation metrics on clarity and decision confidence.

References [1] Kjersti Aas, Martin Jullum, and Anders Løland. 2021. Explaining individual predictions when features are dependent: More accurate approximations to Shapley values. Artificial Intelligence 298 (2021), 103502. [2] Marzia Ahmed, Mohammod Abul Kashem, Mostafijur Rahman, and Sabira Khatun. 2020. Review and Analysis of Risk Factor of Maternal Health in Remote

Oliveira e Silva et al.

Area Using the Internet of Things (IoT). In InECCE2019. Springer Singapore, 357–365. [3] Takuya Akiba, Shotaro Sano, Toshihiko Yanase, Takeru Ohta, and Masanori Koyama. 2019. Optuna: A Next-generation Hyperparameter Optimization Framework. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. Association for Computing Machinery, 9 pages. [4] Emanuele Albini, Jason Long, Danial Dervovic, and Daniele Magazzeni. 2022. Counterfactual Shapley Additive Explanations. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’22). Association for Computing Machinery, 1054–1070. [5] Sajid Ali, Tamer Abuhmed, Shaker El-Sappagh, Khan Muhammad, Jose M. AlonsoMoral, Roberto Confalonieri, Riccardo Guidotti, Javier Del Ser, Natalia DíazRodríguez, and Francisco Herrera. 2023. Explainable Artificial Intelligence (XAI): What we know and what is left to attain Trustworthy Artificial Intelligence. Information Fusion 99 (2023), 101805. [6] Kasun Amarasinghe, Kit T. Rodolfa, Sérgio Jesus, Valerie Chen, Vladimir Balayan, Pedro Saleiro, Pedro Bizarro, Ameet Talwalkar, and Rayid Ghani. 2024. On the Importance of Application-Grounded Experimental Design for Evaluating Explainable ML Methods. Proceedings of the AAAI Conference on Artificial Intelligence 38 (2024), 20921–20929. [7] Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel Weld. 2021. Does the whole exceed its parts? the effect of ai explanations on complementary team performance. In Proceedings of the 2021 CHI conference on human factors in computing systems. 1–16. [8] Umang Bhatt, Alice Xiang, Shubham Sharma, Adrian Weller, Ankur Taly, Yunhan Jia, Joydeep Ghosh, Ruchir Puri, José M. F. Moura, and Peter Eckersley. 2020. Explainable machine learning in deployment. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency. Association for Computing Machinery, 648–657. [9] Michael Bücker, Gero Szepannek, Alicja Gosiewska, and Przemyslaw Biecek. 2022. Transparency, auditability, and explainability of machine learning models in credit scoring. Journal of the Operational Research Society 73, 1 (2022), 70–90. [10] Dulce Canha, Sylvain Kubler, Kary Främling, and Guy Fagherazzi. 2025. A Functionally-Grounded Benchmark Framework for XAI Methods: Insights and Foundations from a Systematic Literature Review. ACM Comput. Surv. 57, 12, Article 320 (2025), 40 pages. [11] Hugh Chen, Ian C. Covert, Scott M. Lundberg, and Su-In Lee. 2023. Algorithms to estimate Shapley value feature attributions. Nature Machine Intelligence 5, 6 (2023), 590–601. [12] Hugh Chen, Joseph D Janizek, Scott Lundberg, and Su-In Lee. 2020. True to the model or true to the data? arXiv preprint arXiv:2006.16234 (2020). [13] Ian Covert, Chanwoo Kim, Su-In Lee, James Zou, and Tatsunori Hashimoto. 2024. Stochastic amortization: a unified approach to accelerate feature and data attribution. In Proceedings of the 38th International Conference on Neural Information Processing Systems (NIPS ’24). Curran Associates Inc., Article 143, 50 pages. [14] Ian Covert and Su-In Lee. 2021. Improving KernelSHAP: Practical Shapley Value Estimation Using Linear Regression. In Proceedings of The 24th International Conference on Artificial Intelligence and Statistics (Proceedings of Machine Learning Research, Vol. 130). PMLR, 3457–3465. [15] Ian Covert, Scott Lundberg, and Su-In Lee. 2021. Explaining by Removing: A Unified Framework for Model Explanation. Journal of Machine Learning Research 22, 209 (2021), 1–90. [16] Anupam Datta, Shayak Sen, and Yair Zick. 2016. Algorithmic Transparency via Quantitative Input Influence: Theory and Experiments with Learning Systems. In 2016 IEEE Symposium on Security and Privacy (SP). IEEE Computer Society, 598–617. [17] Frances Ding, Moritz Hardt, John Miller, and Ludwig Schmidt. 2021. Retiring Adult: New Datasets for Fair Machine Learning. In Advances in Neural Information Processing Systems, Vol. 34. Curran Associates, Inc., 6478–6490. [18] Finale Doshi-Velez and Been Kim. 2017. Towards A Rigorous Science of Interpretable Machine Learning. arXiv:1702.08608 [stat.ML] [19] Joel Dyer, Nicholas Bishop, Yorgos Felekis, Fabio Massimo Zennaro, Anisoara Calinescu, Theodoros Damoulas, and Michael Wooldridge. 2024. Interventionally Consistent Surrogates for Complex Simulation Models. In Advances in Neural Information Processing Systems, Vol. 37. Curran Associates, Inc., 21814–21841. [20] FICO. 2025. FICO Expands Educational Analytics Challenge Program with Three New Historically Black Colleges and Universities to Educate Aspiring Data Scientists. FICO. Accessed: 2026-01-02. [21] Garrett M Fitzmaurice, Nan M Laird, and James H Ware. 2012. Applied longitudinal analysis. John Wiley & Sons. [22] Jeremy Goldwasser and Giles Hooker. 2024. Stabilizing Estimates of Shapley Values with Control Variates. In Explainable Artificial Intelligence. Springer Nature Switzerland, 416–439. [23] Léo Grinsztajn, Edouard Oyallon, and Gaël Varoquaux. 2022. Why do tree-based models still outperform deep learning on typical tabular data? Advances in neural information processing systems 35 (2022), 507–520.

[24] Tom Heskes, Evi Sijben, Ioan Gabriel Bucur, and Tom Claassen. 2020. Causal Shapley Values: Exploiting Causal Knowledge to Explain Individual Predictions of Complex Models. In Advances in Neural Information Processing Systems, Vol. 33. Curran Associates, Inc., 4778–4789. [25] Hans Hofmann. 1994. Statlog (German Credit Data). UCI Machine Learning Repository. [26] Dominik Janzing, Lenon Minorics, and Patrick Bloebaum. 2020. Feature relevance quantification in explainable AI: A causal problem. In Proceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics (Proceedings of Machine Learning Research, Vol. 108). PMLR, 2907–2916. [27] Sérgio Jesus, Catarina Belém, Vladimir Balayan, João Bento, Pedro Saleiro, Pedro Bizarro, and João Gama. 2021. How can I choose an explainer? An Applicationgrounded Evaluation of Post-hoc Explanations. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency. Association for Computing Machinery, 11 pages. [28] Sérgio Jesus, Pedro Saleiro, Inês Oliveira e Silva, Beatriz M. Jorge, Rita P. Ribeiro, João Gama, Pedro Bizarro, and Rayid Ghani. 2024. Aequitas Flow: Streamlining Fair ML Experimentation. Journal of Machine Learning Research 25 (2024), 1–7. [29] Neil Jethani, Mukund Sudarshan, Ian Connick Covert, Su-In Lee, and Rajesh Ranganath. 2022. FastSHAP: Real-Time Shapley Value Estimation. In International Conference on Learning Representations. [30] Amir-Hossein Karimi, Bernhard Schölkopf, and Isabel Valera. 2021. Algorithmic Recourse: from Counterfactual Explanations to Interventions. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’21). Association for Computing Machinery, 353–362. [31] Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie-Yan Liu. 2017. LightGBM: A Highly Efficient Gradient Boosting Decision Tree. In Advances in Neural Information Processing Systems, Vol. 30. Curran Associates, Inc. [32] Farhina Sardar Khan, Syed Shahid Mazhar, Kashif Mazhar, Dhoha A. AlSaleh, and Amir Mazhar. 2025. Model-agnostic explainable artificial intelligence methods in finance: a systematic review, recent developments, limitations, challenges and future directions. Artificial Intelligence Review 58, 8 (2025), 232. [33] Scott M. Lundberg, Gabriel G. Erion, and Su-In Lee. 2019. Consistent Individualized Feature Attribution for Tree Ensembles. arXiv:1802.03888 [cs.LG] [34] Scott M. Lundberg and Su-In Lee. 2017. A unified approach to interpreting model predictions. In Proceedings of the 31st International Conference on Neural Information Processing Systems. Curran Associates Inc., 4768–4777. [35] Luke Merrick and Ankur Taly. 2020. The Explanation Game: Explaining Machine Learning Models Using Shapley Values. In International Cross-Domain Conference for Machine Learning and Knowledge Extraction. Springer International Publishing, 17–38. [36] Lucas Thibaut Meyer, Marc Schouler, Robert Alexander Caulk, Alejandro Ribes, and Bruno Raffin. 2023. Training Deep Surrogate Models with Large Scale Online Learning. In Proceedings of the 40th International Conference on Machine Learning, Vol. 202. PMLR, 24614–24630. [37] Ibomoiye Domor Mienye, George Obaido, Nobert Jere, Ebikella Mienye, Kehinde Aruleba, Ikiomoye Douglas Emmanuel, and Blessing Ogbuokiri. 2024. A survey of explainable artificial intelligence in healthcare: Concepts, applications, and challenges. Informatics in Medicine Unlocked 51 (2024), 101587. [38] Christoph Molnar, Giuseppe Casalicchio, and Bernd Bischl. 2020. Interpretable Machine Learning – A Brief History, State-of-the-Art and Challenges. In ECML PKDD 2020 Workshops. Springer International Publishing, 417–431. [39] Ramaravind K. Mothilal, Amit Sharma, and Chenhao Tan. 2020. Explaining machine learning classifiers through diverse counterfactual explanations. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency. Association for Computing Machinery, New York, NY, USA, 11 pages. [40] Maximilian Muschalik, Fabian Fumagalli, Barbara Hammer, and Eyke Hüllermeier. 2024. Beyond TreeSHAP: efficient computation of any-order shapley interactions for tree ensembles. In Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence (AAAI’24). AAAI Press, Article 1604, 9 pages. [41] Meike Nauta, Jan Trienes, Shreyasi Pathak, Elisa Nguyen, Michelle Peters, Yasmin Schmitt, Jörg Schlötterer, Maurice van Keulen, and Christin Seifert. 2023. From Anecdotal Evidence to Quantitative Evaluation Methods: A Systematic Review on Evaluating Explainable AI. ACM Comput. Surv. 55, 13s (2023). [42] John Ashworth Nelder and Robert WM Wedderburn. 1972. Generalized linear models. Journal of the Royal Statistical Society Series A: Statistics in Society 135 (1972), 370–384. [43] Lars Henry Berge Olsen and Martin Jullum. 2026. Improving the Weighting Strategy in KernelSHAP. In Explainable Artificial Intelligence. Springer Nature Switzerland, 194–218. [44] Iker Perez, Piotr Skalski, Alec Barns-Graham, Jason Wong, and David Sutton. 2022. Attribution of predictive uncertainties in classification models. In Proceedings of the Thirty-Eighth Conference on Uncertainty in Artificial Intelligence (Proceedings of Machine Learning Research, Vol. 180). PMLR, 1582–1591. [45] José C Pinheiro and Douglas M Bates. 2000. Mixed-effects models in S and S-PLUS. Springer.

A Human-Centered Audit of Shapley Value Benchmarks in High-Stakes Workflows

[46] Forough Poursabzi-Sangdeh, Daniel G Goldstein, Jake M Hofman, Jennifer Wortman Vaughan, and Hanna Wallach. 2021. Manipulating and measuring model interpretability. In Proceedings of the 2021 CHI conference on human factors in computing systems. 1–52. [47] Cynthia Rudin. 2019. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence 1, 5 (2019), 206–215. [48] Waddah Saeed and Christian Omlin. 2023. Explainable AI (XAI): A systematic meta-survey of current challenges and future opportunities. Knowledge-Based Systems 263 (2023), 110273. [49] Maria Sahakyan, Zeyar Aung, and Talal Rahwan. 2021. Explainable artificial intelligence for tabular data: A survey. IEEE access 9 (2021), 135392–135422. [50] Lloyd S Shapley. 1953. A Value for n-Person Games. In Contributions to the Theory of Games II. Princeton University Press, Princeton, 307–317. [51] Venkatesh Sivaraman, Leigh A Bukowski, Joel Levin, Jeremy M Kahn, and Adam Perer. 2023. Ignore, trust, or negotiate: understanding clinician acceptance of AI-based treatment recommendations in health care. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–18. [52] Erik Štrumbelj and Igor Kononenko. 2010. An Efficient Explanation of Individual Classifications using Game Theory. Journal of Machine Learning Research 11, 1 (2010), 1–18. [53] Mukund Sundararajan and Amir Najmi. 2020. The Many Shapley Values for Model Explanation. In Proceedings of the 37th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 119). PMLR, 9269–9278. [54] Muhammad Faaiz Taufiq, Patrick Blöbaum, and Lenon Minorics. 2023. Manifold Restricted Interventional Shapley Values. In Proceedings of The 26th International Conference on Artificial Intelligence and Statistics (Proceedings of Machine Learning Research, Vol. 206). PMLR, 5079–5106. [55] Boris Van Breugel and Mihaela Van Der Schaar. 2024. Position: Why Tabular Foundation Models Should Be a Research Priority. In Proceedings of the 41st International Conference on Machine Learning, Vol. 235. PMLR, 48976–48993. [56] Giulia Vilone and Luca Longo. 2021. Notions of explainability and evaluation approaches for explainable artificial intelligence. Information Fusion 76 (2021), 89–106. [57] Oskar Wysocki, Jessica Katharine Davies, Markel Vigo, Anne Caroline Armstrong, Dónal Landers, Rebecca Lee, and André Freitas. 2023. Assessing the communication gap between AI models and healthcare professionals: Explainability, utility and trust in AI-driven clinical decision-making. Artificial Intelligence 316 (2023), 103839. [58] Jilei Yang. 2021. Fast TreeSHAP: Accelerating SHAP Value Computation for Trees. ArXiv abs/2109.09847 (2021). [59] Peng Yu, Albert Bifet, Jesse Read, and Chao Xu. 2022. Linear tree shap. In Advances in Neural Information Processing Systems, Vol. 35. Curran Associates, Inc., 25818–25828. [60] Artjom Zern, Klaus Broelemann, and Gjergji Kasneci. 2023. Interventional SHAP Values and Interaction Values for Piecewise Linear Regression Trees. Proceedings of the AAAI Conference on Artificial Intelligence 37, 9 (2023), 11164–11173.

A

Extended Overview of Shapley Value Formulations

The definition of the value function 𝑣 𝒙 (S) hinges on how one represents feature absence. These formulations are not interchangeable; they encode fundamental assumptions about the data-generating process and the intended use of the explanation. Table 1 summarizes the formulations considered in this work. Here, we extend their definitions with a discussion on use cases and limitations.

A.1

Fixed Baseline Shapley Values

The fixed baseline formulation replaces absent features with values from a predefined reference input 𝒙 ′ [34, 53]. The background distribution is: 𝑝 (·) = 𝛿 (𝑿 F\S = 𝒙 ′F\S ). While computationally efficient (requiring only one model evaluation per coalition), this method is highly sensitive to the choice of 𝒙 ′ . As there is rarely a "neutral" reference point in tabular data [11], the choice of zero, mean, or median can fundamentally shift the attribution logic.

A.2

Marginal (Interventional) Shapley Values

In the marginal formulation [26, 53], absent features are sampled from their global empirical distribution: 𝑝 (·) = 𝑝 (𝑿 F\S ). This corresponds to do-intervention under a causal graph with no feature dependencies [24, 26]. While it isolates the model’s response to specific features, it frequently generates off-manifold inputs, i.e. combinations of values that could never occur in reality. This may lead to misleading attributions in highly correlated datasets [35].

A.3

Joint-Marginal Shapley Values

The joint-marginal variant treats every absent feature as independent, sampling each from its univariate marginal [16], i.e. Ö 𝑝 (·) = 𝑝 (𝑋 𝑗 ). 𝑗 ∈ F\S

By enforcing independence, it simplifies computation but amplifies the off-manifold problem. In our study, this variant was rated as the clearest by analysts, suggesting that "true-to-the-model" logic may be easier for humans to parse than complex distributions.

A.4

Conditional Shapley Values

The conditional formulation preserves empirical dependencies by conditioning on observed coalition values [1, 11]: 𝑝 (·) = 𝑝 (𝑿 F\S | 𝑿 S = 𝒙 S ). This produces realistic in-manifold imputations but is computationally intensive and difficult to estimate in high dimensions, where practical implementations rely on approximate matching, nearest neighbours, or parametric density models. Our audit shows this variant significantly inflates analyst confidence without improving accuracy.

A.5

Uniform Shapley Values

The uniform formulation replaces absent features with samples drawn uniformly from a predefined hyperbox [35, 52]: 𝑝 (·) = 𝑢 (𝑿 F\S ). While useful for exploring the model’s global behavior, it often results in noisy, confusing explanations for human reviewers.

A.6

Search Counterfactual Shapley Values

This formulation identifies the minimal changes needed to flip a model’s prediction [4]. It replaces absent features with values from a vector 𝑿 (𝑐 ) that reaches a target outcome 𝑦 ∗ : ) ) 𝑝 (·) = 𝑝 (𝑿 (𝑐 | 𝑓 (𝑿 (𝑐 ) ≈𝑦 ∗ ), F\S F

This emphasizes actionability is but is computationally demanding [30], and can be sensitive to chosen distance metrics or plausibility constraints.

A.7

Filtered Conditional Shapley Values

This variant restricts the background distribution to samples where model outputs fall within a specific range Y: 𝑝 (·) = 𝑝 (𝑿 F\S | 𝑓 (𝑿 F ) ∈ Y).

Oliveira e Silva et al.

Table 4: Overview of key functional properties of explanation quality [54], with example attributes. Property [54]

What it captures

Example quantitative / qualitative attributes

Representativeness

What explanation is provided.

Local vs. global, model portability or data coverage.

Structure

How is the explanation provided.

Expressive power, graphical integrity, morphological clarity.

Selectivity

Size of the explanation.

How concise. Sparsity and top-𝑘 relevance.

Contrastivity

Relational differences.

Sensitivity to counterfactuals, benchmarks or perturbations.

Interactivity

User control.

Interaction options, adjustable parameters or interfaces.

Fidelity

Surrogate consistency

Approximation accuracy and model assumptions.

Faithfulness

How reliable for the black-box.

Causal correctness, feature removal and white-box checks.

Truthfulness

Domain alignment.

Expert plausibility; bias exposure.

Stability

Behavioural consistency.

Similarity across perturbations or resampling.

(Un)certainty

How transparent is it?

Confidence disclosure, stochastic awareness.

Speed

Computational efficiency.

Latency; throughput under fixed budgets.

In our audit, this produced the most sparse and contrastive explanations, yet was consistently rated as confusing by professional analysts.

A.8

Excluded Formulations

We excluded causal [24] and generative in-manifold [54] variants due to their high computational overhead and requirement for strong causal assumptions. These are currently incompatible with the millisecond-level latency required in production risk systems and are thus outside the scope of our operational audit.

B

Properties of Explanation Quality

Recent taxonomies of explainability [10] identify a broad set of functionally grounded properties that characterize a "good" explanation. These span computational, representational, and human-centered dimensions. This appendix refines these concepts for the Shapley value framework, distinguishing between properties that are inherent to the framework and those that are formulation-dependent and thus central to our audit. Table 4 provides a high-level overview.

B.1

Properties Invariant Across Shapley Configurations

The following properties are inherent to the Shapley framework itself, and remain constant across all eight variants in our study. Thus, they do not serve as discriminators in our evaluation. Representativeness. Shapley values are, by construction, local feature attributions that are model-agnostic [1, 34]. All formulations share this scope; therefore, representativeness does not vary between definitions. Structure. The structural form of a Shapley explanation is always additive. Consequently, the visual format (e.g., bar charts) is independent of the value-function choice. All variants in our study were presented using the exact same visual encoding to ensure structural parity.

Interactivity. Interaction depends on the UI (e.g., toggles, sliders) rather than the mathematical definition. By holding the interface fixed, we ensure interactivity remains a constant across our experimental matrix. (Un)certainty. While uncertainty estimates can be attached to explanations (e.g., via ensemble bootstraps [15]), these quantify the stochasticity of the underlying model or the estimation procedure, not the conceptual choice of the value function. Speed. While training costs vary, the inference speed is standardized by our use of amortized surrogates. This ensures that the operational feasibility of the explanations is consistent across all formulations during the analyst review process.

B.2

Properties Relevant for Shapley Evaluation

These properties depend directly on the choice of value function and form the basis of our empirical audit. Selectivity. This captures how concentrated attributions are across features, and determines the amount of information users must process, affecting cognitive load [10]. Different formulations produce different sparsity profiles. We quantify this via the Sparsity metric (8), testing the assumption that "less is more" for human interpretability. Contrastivity. It captures whether explanations express differences relative to a meaningful reference or baseline [10]. Certain Shapley definitions (e.g., counterfactual, causal) are inherently contrastive. Other definitions may conflate average and specific contributions, leading to sensitivity to arbitrary baselines. We operationalize this through sensitivity to counterfactual perturbations (7). Fidelity. Fidelity reflects how accurately the additive decomposition of Shapley attributions reconstruct a black-box prediction. Specifically, how accurately an amortized surrogate reconstructs the "ground truth" for its specific definition. We use Attribution Error and Recall@k as control measures to ensure that our surrogates are mathematically sound.

A Human-Centered Audit of Shapley Value Benchmarks in High-Stakes Workflows

Faithfulness. It captures whether the explanation preserves the causal influence of features on predictions [38]. Because different value functions define different intervention semantics, faithfulness varies significantly. We evaluate this using Deletion and Insertion AUC (5). Truthfulness. Truthfulness connects the explanation to domain reality and expert priors [10]. Some value functions can yield explanations that contradict established domain knowledge. We assess this through our human-in-the-loop experiments in Section 3.1.3, measuring whether an explanation helps or hinders an analyst’s decision-making process. Stability. Stability denotes reproducibility under small input or model perturbations. Formulations that integrate over correlated feature distributions tend to produce smoother attributions. We operationalize this via the Perturbation Sensitivity metric (6).

HELOC. Credit bureau attributes from FICO [20] used to predict risk for home equity lines of credit. This dataset is a cornerstone for XAI evaluation in the financial sector due to its high-dimensional numerical feature set. Fraud Risk. This is a proprietary, large-scale dataset of card payment transactions collected and downsampled from a real-world financial risk assessment system. Each instance is represented by a heterogeneous set of numerical and categorical features capturing transactional, behavioral, and contextual information available at decision time. The dataset is highly imbalanced, reflecting real operational fraud prevalence.

C.1

Dataset Preprocessing

Table 5: Summary statistics for experimental datasets; including sample sizes for training, validation, and A/B testing with amortisers. Prevalence refers to the proportion of positive (risk) instances.

All datasets undergo a standardized preprocessing pipeline. Categorical features are ordinally encoded to preserve fixed dimensionality across models and explainers. Numerical features are used as provided. Dataset-specific preprocessing is minimal, except for HELOC, which contains missing or invalid sentinel values; these are handled via feature filtering, instance removal, and targeted imputation. Each dataset is divided into training, validation, and test sets using stratified sampling to preserve class proportions. Validation and test sets are capped at 10% of the dataset size or a maximum of 200 instances each for A/B Testing, whichever is smaller. All splits are generated using a fixed random seed to ensure reproducibility. The final number of features, label prevalence, and split sizes for each dataset are reported in Table 5.

Dataset

D

C

Datasets

We evaluate five datasets commonly used in risk assessment and explainability research, summarized in Table 5. These span healthcare, credit scoring, income prediction, and fraud detection, and vary in feature types, class imbalance, and task complexity.

Maternal Risk German Credit Adult HELOC Fraud Risk*

Feat. 6 20 10 22 54

Cat. 0 11 6 0 15

Prev. 26.8% 30.0% 63.1% 52.1% 9.1%

Train 812 800 998,300 7,029 16,496

Val. 101 100 200 200 200

Test 101 100 200 200 200

*Heavily downsampled to achieve a 10-1 risk ratio before amortization.

Maternal Risk. Clinical measurements used to predict maternal health outcomes [2]. The task is to predict whether a pregnancy is associated with elevated health risk. This represents a low-dimensional, fully continuous healthcare task. German Credit. The German Credit dataset [25] consists of an anonymized financial and demographic set of features for loan applicants. It combines numerical and categorical attributes describing credit history and employment, frequently used to study the intersection of XAI and algorithmic fairness. Adult. The Adult (Income) dataset [17] contains demographic and employment-related attributes used to predict income thresholds. We use a curated subset of features from the Aequitas package [28] to focus on socio-economic indicators common in automated decision-making.

Extended Experimental Setup

Our audit utilizes a standardized training and optimization pipeline to ensure that variations in explanation quality are not confounded by model performance or optimization noise. We train two predictive binary classification risk models per dataset: Logistic Regression and LightGBM [31]. Hyperparameters were optimized via Optuna [3] with a Tree-structured Parzen Estimator (TPE) sampler and median pruning. We use 3-fold stratified cross-validation and 20 trials per configuration. The optimization objective was set to validation AUC. Final model configurations are summarized in Table 6. Amortizer Architecture and Training. To support the low-latency requirements of the audit, we deploy feed-forward neural network amortizers. • Preprocessing: Numerical features are scaled using robust scaling (IQR-based) with tanh saturation to mitigate outliers. Categorical features are passed through learned embedding layers. • Architecture: Layers include fully connected blocks of varying dimensionality, Layer Normalization, LeakyReLU activations, and Dropout (0.1). • Optimization: We employ a two-phase schedule: (1) Adam optimization with a linear warmup, followed by (2) SGD with momentum and cosine decay. • Parameters: Amortizers are trained for 1000 epochs (100 for Adult) with a learning rate of 0.001. We sample four

Oliveira e Silva et al.

Table 6: Risk model hyperparameter configurations.

Table 7: Summary of participant profiles in the A/B testing experiments.

Hyperparameter Maternal Credit Adult HELOC Fraud Total participants Professional analysts

LightGBM Num. Leaves Learning Rate N-Estimators Max. Depth Min. Child Samples Subsample Bagging Frac. Feature Frac. Colsample/tree Early Stopping

11 0.27 258 4 25 0.76 0.59 0.65 0.72 72

82 0.10 121 4 22 0.65 0.68 0.86 0.61 132

98 0.14 285 6 90 0.60 0.96 0.54 0.52 78

59 0.05 187 3 100 0.61 0.99 0.76 0.89 179

6 0.10 50 3 500 0.50 0.50 0.50 0.50 20

0.13 0.00 𝑆1 648 Bal. T 0.66

0.36 0.01 𝑆1 237 Bal. F 0.82

0.60 0.00 𝑆2 374 Bal. T 0.31

0.59 0.00 𝑆1 102 Bal. T 1.53

0.09 0.00 𝑆3 785 Bal. F 0.19

Logistic Regression C (Reg. Strength) Tol Solver Max Iterations Class Weight Fit Intercept Intercept Scaling

𝑆 1 : newton-cholesky, 𝑆 2 : sag, 𝑆 3 : newton-cg, Bal.: Balanced

37 5 (13.5%) Moderate 10 (27.0%)

High 18 (48.6%)

Shapley knowledge

Yes 25 (67.6%)

No 12 (32.4%)

Domain knowledge Maternal Risk UCI German Credit Adult FICO HELOC Fraud Risk

Yes 3 (12.5%) 4 (13.3%) 3 (12.5%) 2 (10.0%) 8 (38.1%)

No 21 (87.5%) 26 (86.7%) 21 (87.5%) 18 (90.0%) 13 (61.9%)

ML knowledge

E.2

Low 9 (24.3%)

Experimental Interface and Task Flow

The interface was designed to resemble operational risk analysis tools commonly used in practice. All visual components, layouts, and interaction patterns were held fixed across datasets, models, and Shapley value formulations.

Shapley masks per input during training, with an efficiency loss weight of 0.1.

E.2.1 Overview and task flow. The experiment followed a strictly linear workflow to prevent backtracking or cross-case comparison:

Implementation of Value Functions. We amortize eight distinct value functions (summarized in Appendix A). Sampling-based formulations (Marginal, Conditional, etc.) utilize a background set of 100 instances specifically generated according to the respective sampling logic. Fixed baseline formulations are implemented using both Zero and Mean reference points to contrast their performance against empirical variants.

(1) Onboarding: Dataset selection and completion of the analyst profile. (2) Contextualization: Review of standardized, dataset-specific task summaries. (3) Review Loop: Sequential case analysis and decision-making.

E

Human-in-the-Loop A/B Testing Framework Details

To ensure the validity of our audit, we recruited a diverse pool of participants and professional analysts to interact with a standardized risk-assessment interface. By holding the UI, model outputs, and task context constant, we isolated the choice of Shapley formulation as the sole independent variable.

E.1

Participant Profiles

Table 7 summarizes the 37 individuals who contributed to the 3,735 human–AI interactions recorded in this study. The pool includes 5 professional fraud analysts (13.5%) and a high proportion of participants with significant ML and Shapley-specific knowledge, ensuring that the observed "confidence-inflation" was not merely a result of user naivety. Here, domain knowledge reflects self-assessed familiarity with the specific dataset being analyzed, presented separately for each dataset included in the study.

Participants could only advance after submitting a final decision, ensuring that recorded interactions were independent. E.2.2 Dataset Selection and Analyst Profile. Upon selecting a dataset from a drop-down menu, participants used a panel to self-report their expertise in three areas: domain knowledge for the specific data, general ML familiarity, and prior experience with Shapleybased explanations. These attributes are collected exclusively for post-hoc analysis and stratification and do not influence task assignment, model selection, or explanation configuration. A screenshot of the dataset selection and analyst profile form is shown in Figure 5. E.2.3 Task summary panel. A collapsible summary panel provided a constant reference for the prediction objective, feature definitions (including categorical levels), and the operational meaning of risk. This ensured that all participants, regardless of prior expertise, operated under a shared conceptual framework. Figure 6 shows an example of the task summary panel.

A Human-Centered Audit of Shapley Value Benchmarks in High-Stakes Workflows

Figure 5: A screenshot showing the initial onboarding form with dropdowns for expertise levels. Data explorer. A read-only explorer allows participants to inspect the raw feature values of the current instance. Numerical features were accompanied by univariate distributions, while categorical features included prevalence and historical risk statistics. This ensured analysts could verify the data without manipulating it. E.2.5 Decision and feedback collection. The feedback loop (Figure 7) was designed to capture the core outcome variables of our audit. After reviewing a case, analysts submitted: • Decision: A binary "Risk" or "No Risk" judgment. • Confidence: Self-reported certainty (Weak, Moderate, Strong). • Clarity: A subjective assessment of the explanation (Clear or Confusing).

Figure 6: Task summary panel providing analysts with dataset and task context. E.2.4 Case-review screen. The interface (detailed in Figure 2 of the main text) comprised three primary information blocks: Model output. The risk score is displayed as a value in [0, 1], together with its empirical percentile within the dataset. A score distribution plot provides additional context by situating the instance relative to the overall population. Explanation panel. Explanations featured a horizontal bar chart of the four features with the highest absolute Shapley attributions, using consistent color, scale, and ordering conventions across all Shapley formulations. To mirror real-world compliance requirements, the interface displays automatically generated reason codes, i.e. short natural-language statements deterministically derived from the Shapley values (e.g., “Age value of 70 is high”).

Figure 7: Sequence of questions asked to the analyst for each reviewed case. No additional controls, filters, or explanation parameters are exposed. This design ensures that observed differences in decision behavior, confidence, and response time arise from explanation content rather than interface affordances or user-driven customization.

Record · ID 134537 · SHA-256 353c3934c5b62dff
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.