Why Model Credibility Isn’t Enough: - Rethinking Trust in Simulation Architectures Romain Barbedienne 1) Adeline Lanugue 1) Rim Kaddah 1) Julien Silande 1,2) Anthony Levillain 3) Cedric Leclerc 4) Maxime Hayet 5) Boussaad Soualmi 1) Cristian Maxim 1) 1) IRT SystemX, 2 Bd Thomas Gobert, 91120 Palaiseau, France (E-mail: [email protected]) 2) Keysight Technology, La janais 3 rue Pierre et marie curie, 35131 Chartres-de-Bretagne, France 3) OPmobility Alphatech, ZAC du Bois de Plaisance, 214 Av. de la Mare Gessart, 60280 Venette, France 4)Renault Technocentre, 1 avenue du Golf, 78280 Guyancourt, France 5)Stellantis Green-Campus, 43 Rue Jean Pierre Timbaud, 78300 Poissy, France ABSTRACT: Credibility of a simulation model is an important topic. Several approaches try to quantify the credibility of simulation. However, models are mostly assembled within a simulation architecture. Can the credibility of a simulation architecture be assessed based on the credibility of the models that comprise it? This paper aims to address this issue by providing an overview of the current state of the art in the field of assembly credibility. It will compare sensitivity analysis techniques, qualitative analysis by experts, explainability in AI, and networks. Finally, an assessment of the proposed approaches, based on criteria such as rigor, generalization, and resource requirements, will reveal the strengths and weaknesses of each approach. KEY WORDS: Credibility, Model Architecture, Sensitivity Analysis 1.INTRODUCTION
These submodels can sometimes even be delivered by suppliers
Designing a system requires ensuring that it operates
[1]. In this context, is it possible to evaluate the credibility of a
correctly while guaranteeing safety appropriate for its intended
simulation architecture based on the credibility of the models that
use. In recent years, computer simulation has become widely used
comprise it?
in the industrial sector to validate these functional and safety
This paper is a analyze different solution that aimed at
requirements. Indeed, for complex systems, simulation is less
evaluating scientific approaches to addressing this issue. The first
expensive than conducting real-life tests. Thereby, simulation
part of this paper is focus on the methods that allows to evaluate
leads to anticipate and test a system as soon as possible to enhance
credibility of a simulation model. We will discuss the direct
it. Moreover, it allows to realize a lot of tests in a short time and
application of these methods to the evaluation of simulation
quickly improve the system. However, to be pertinent and prevent
architectures. The second part will explore the methods that can
errors in decision making, these tests should be done in a model
be used to evaluate the impact of a model in a simulation
which is enough credible and robust.
architecture. The third part is an analysis and a comparison of
In the M&S domain, credibility could be defined as the
different methods.
degree of confidence that a stakeholder places in a product for a specific purpose based on established independent evidence. To
2.Evaluation of simulation model credibility
ensure the credibility of a model, it is necessary to conduct efforts
The credibility of simulation model is characterized by the
to ensure that the system is properly designed and that its
degree of confidence that a stakeholder places in a product for a
verification and validation have been carried out correctly. The
specific purpose based on established independent evidence. Thus,
credibility of a simulation model is essential to ensure the right
assessing the credibility of a model is subjective; it cannot be
decision. Improving the credibility of a simulation model is an
measured using precise metrics. However, several studies have
essential step, but it can be very time-consuming. It is therefore
focused on rationalizing credibility.
necessary to strike the right balance between improving the model's credibility and the time spent on this task. In practice, for decisions, it is necessary to combine
2.1.Method for ensuring the credibility of models
models from different teams. Engineers build model architectures
NASA has introduced Standard 7009B [2]. This standard
composed of several submodels. These models are connected to
defines uniform technical and engineering requirements for the
one another via interfaces. Each interface contains one or more
processes, procedures, practices, and methods recognized as
variables that are exchanged between the submodels. The
standards in modeling and simulation. One of the key elements in
submodels constituting this architecture have their own credibility.
assessing modeling and simulation capabilities is the set of
recommendations covering the following topics: data traceability,
concern the credibility assessment when using virtual chains tools
verification, validation, technical review, and product/process
to validate vehicle. Some rules are linked with the experiences of
management. Each requirement must be met throughout a model’s
the team, others about proofs, and the traceability in a
lifecycle, which adds complexity to the design process since the
documentation. To respect these rules, as soon as possible, in the
requirements must be reviewed with every modification. This
V-Model enhance the credibility of the model and the system.
enhances the model’s credibility but is particularly suited for high-
2.4. Comparison and discussion
criticality simulation models.
The five approaches have the common objective of rationalizing credibility assessment but differ in scope and rigor.
2.2.Frameworks for credibility assessment
PCMM and NASA-STD-7009B are the most scientifically
Oberkampf et al. [3] introduced the Predictive Capability
grounded, addressing respectively the technical foundations of a
Maturity Model (PCMM). This framework evaluates the maturity
model. The Prostep ivip and IRT SystemX frameworks adopt a
of a simulation model across several dimensions: representation
more pragmatic stance, offering structured tools that facilitate
and geometric fidelity, physics and material model fidelity, code
communication between stakeholders in industrial contexts.
verification,
solution
uncertainty
quantification
verification,
model
validation,
and
However, these approaches remain largely declarative: credibility
sensitivity
analysis.
Each
relies on self-reported information, with limited mechanisms to
dimension is defined for different maturity levels. PCMM allows
independently verify the claims made — a significant weakness
for a quick assessment of a model’s strengths and weaknesses.
when models are delivered by external suppliers. The European
One limitation of this approach is that it is primarily applicable to
Commission regulation provides a useful normative baseline but
3D models.
stays general and qualitative, oriented toward certification
and
Prostep ivip publishes a white paper [4] where they explain the
compliance rather than scientific rigor.
emphasizes the need to standardize the assessment of simulation
A critical limitation shared by all five frameworks is that they
credibility by incorporating criteria related to quality, accuracy,
are designed to assess standalone simulation models, and their
and criticality, to strengthen stakeholders’ confidence in the
extension to a simulation architecture raises unresolved scientific
results obtained. This document introduces a framework for
challenges.
assessing the credibility of processes. Within this framework, the
experimental reference data, which is often unavailable at the
Wheel of V&V and the Credibility Spider Chart help highlight the
architecture level — particularly in early design phases or in
strengths and weaknesses of a simulation model, enabling
OEM–supplier relationships where test data cannot be shared.
stakeholders to quickly assess the model’s credibility. This highly
Furthermore, none of these frameworks provides a systematic
promising work needs to be extended to other areas of verification
method to address compositional aspects: even if each submodel
and validation; indeed, the credibility of a model also depends on
is individually credible, the propagation of uncertainties across
its design and its specifications.
interfaces and emergent behaviors arising from model coupling
Most
validation
methods
require
access
to
Finally, IRT SystemX publishes a white paper [5] to improve
remain largely uncharacterized. Addressing these issues requires
the interaction between OEM and supplier in the context of
developing methods capable of aggregating submodel credibility
simulation. They provide a questionnaire designed to assess the
and reducing reliance on declarative assessments in favor of more
credibility of simulation models across five key areas: robustness
objective, verifiable evidence. In order to assess the credibility of
and sensitivity, model uncertainty and margin, verification,
a model assembly based on its submodels, the remainder of this
validation, and use. This questionnaire is summarized in a memo
paper will focus on methods for evaluating the impact of a
that provides an overview of the level of rigor applied to each
simulation model within a model assembly.
model. A minimum score is required for each axis, based on the criticality of the simulation, the system design cycle, and whether or not the model needs to be integrated into a simulation architecture.
3.Evaluation of impact of a model in an Assembly 3.1.Problem statement A simulation architecture is composed of submodels connected through interfaces, each submodel carrying its own credibility
2.3. Regulations
level Figure 1. Even when each submodel is found to be
The European commission has established a regulation aimed
sufficiently credible individually, this does not imply that the
at establishing guidelines for the certification of vehicles capable
whole assembly is credible, as allowable errors from submodels
of operating without human intervention [6]. The four parts
may accumulate to unacceptable levels at the system level. Rather
than attempting a direct compositional aggregation of model
3.2.Sensitivity analysis
credibility an alternative strategy consists in evaluating the impact
Sensitivity analysis is a broad family of methods aimed at
of each submodel on the architecture's output variables of interest.
determining how variations in model inputs or parameters
This impact-based approach reframes the problem: instead of
propagate to model outputs. Depending on whether the analysis is
asking whether a submodel is credible in isolation, it asks how
conducted at a single operating point or across the entire parameter
much a submodel's uncertainty or error propagates to the decision-
space, these methods are generally classified as either local or
relevant outputs of the architecture.
global. The following subsections review both approaches, discussing their respective theoretical foundations, practical applicability, and limitations in the context of simulation model assemblies. 3.2.1.
Local sensitivity analysis Local sensitivity analysis characterizes how small
perturbations in model parameters affect model outputs in the vicinity of a nominal operating point [7]. Formally, given a model Figure 1 Credibility of simulation architecture This analysis is inherently context dependent. The validity of a model is defined over a specific domain of model form, inputs, parameters, and responses, which effectively limits its use to the particular application for which it was validated. Thus, the
F:ℝ𝑝 → ℝ𝑞 mapping a parameter vector Λ ∈ ℝ𝑝 to an output vector 𝑌 ∈ ℝ𝑞 , the local sensitivity is defined as the partial derivative of model output 𝑦𝑖 related to parameter 𝜆𝑗 evaluated at a nominal point Λ1 (where 𝑖 ∈ ⟦1, 𝑞⟧ and 𝑗 ∈ ⟦1, 𝑝⟧): 𝜕𝑦𝑖 | Sij = 𝜕𝜆𝑗 Λ=Λ1
contribution of a given submodel to the architecture's output
In practice, when analytical derivatives are unavailable, this
depends on the scenarios and parameter sets used for decision-
quantity is approximated numerically using finite difference
making. Two submodels may have very different impacts
schemes [8].
depending on the operating conditions simulated. This implies that
From a computational standpoint, local sensitivity analysis is
the impact analysis must be conducted specifically for the
highly efficient. For a model with 𝑝 parameters, a full sensitivity
scenarios and parameter ranges used for the decision making.
matrix requires only 𝑝 + 1 model evaluations using a forward
Furthermore, a model may be valid for one set of experimental
difference scheme, or at most 2p evaluations with a centered
conditions and invalid in another; therefore, during the impact
difference scheme. This makes the approach tractable even for
assessment, input variables and parameters of each submodel must
models with a large number of parameters, provided each
remain strictly within their respective validation domains to
evaluation is not prohibitively expensive. When derivatives can be
ensure that the conclusions drawn are meaningful. Finally, one last
computed analytically or via automatic differentiation, the cost
consideration to keep in mind is that simulation models are often
reduces further, requiring only a single forward pass through the
compiled in the Functional Mock-up Interface (FMI) format. This facilitates co-simulation, but also allows models to be exchanged
model alongside its linearization. However, local sensitivity analysis rests on a fundamental
between suppliers and OEMs without the risk of sharing
assumption: that the model behavior is approximately linear in the
proprietary know-how.
neighborhood of the nominal point Λ1 . This assumption carries
To systematically evaluate the contribution of each submodel
several important consequences. First, the method is strictly valid
to the architecture's outputs, two complementary families of
only when parameter variations remain small relative to the scale
methods will be explored in this paper. Sensitivity analysis
over which the model's behavior changes. Second, and more
methods are useful for testing the robustness of results, identifying
critically, local sensitivity indices are entirely dependent on the
model inputs that cause significant impact in the output. Moreover,
choice of nominal point, meaning that the conclusions drawn may
statistical methods offer a complementary perspective by
not be representative of the model's behavior across its full
characterizing the probabilistic distribution of outputs and
parameter range. In the presence of nonlinearities, the sensitivity
quantifying uncertainty propagation across the architecture. These two approaches will be explored in the following sections.
structure can vary dramatically across the parameter space, rendering point-wise estimates potentially misleading. Another example of limitation of local sensitivity is the failure to distinguish between a local minimum and an inflection point. This
difference is, however, crucial for determining the robustness of a model Figure 2.
The Morris method is frequently used as a first step to identify the most influential parameters before proceeding to a more resource-intensive analysis. At this stage, it is critical not to underestimate the influence of any parameter, as a false negative (incorrectly classifying an influential parameter as negligible) would propagate through all subsequent analyses. To address this risk, Sohier et al. [10] proposed an adaptive variant of the Morris method in which the number of elementary effect samples is increased selectively for parameters with low estimated influence,
𝑑𝐹
Figure 2 : Both cases yield 𝑑𝜆𝑖 (𝜆1 ) ≈ 0 so local analysis alone cannot distinguish them
thereby reducing the probability of misclassification while controlling the total number of model evaluations. Two further screening approaches are worth mentioning. One-
3.2.2.
Global sensitivity analysis
At-a-Time (OAT) analysis [11] varies each parameter individually
Unlike local sensitivity analysis, which characterizes model
across its full range while holding all others fixed, and plots the
behavior at a single nominal point, global sensitivity analysis
resulting output response to reveal the form of each parameter's
evaluates the influence of parameters over their entire admissible
influence (including nonlinearities and threshold effects). It is
range. By exploring the full parameter space rather than a local
simpler than the Morris method but shares its inability to detect
neighborhood, this family of methods provides a more robust and
interactions, and the fraction of parameter space explored shrinks
representative characterization of model behavior, in particular
super-exponentially with the number of parameters, making it
when nonlinearities or parameter interactions are present. Global
unsuitable for high-dimensional problems. Sequential bifurcation
sensitivity analysis methods are generally organized into three
[12] takes a different approach: parameters are grouped and tested
main families: screening-based approaches, variance-based
in aggregate, with groups progressively subdivided until
approaches, and metamodel-based approaches
individual influential parameters are isolated. It is both effective and efficient, requiring relatively few simulations run, but
Screening-based approaches. The objective of screening
assumes that the model response can be approximated by a first-
methods is to rank parameters by order of influence at low
order polynomial and that the signs of parameter effects are known
computational cost, typically as a preliminary step before a more
a priori, which limits its applicability to FMI models.
expensive analysis. The most widely used method in this family is the Morris method [9]. The Morris method proceeds as follows.
Variance-based
approaches.
Variance-based
methods
Starting from a randomly sampled initial configuration in the
provide a quantitative decomposition of output variance into
parameter space, each parameter is perturbed individually by a
contributions attributable to each parameter and to their
fixed step, and the corresponding elementary effect (defined as the
interactions. The most rigorous framework in this family is that of
ratio of the output variation to the parameter perturbation) is
Sobol indices [13]. This framework decomposes the variance of
recorded. This procedure is repeated for r independently sampled
the output of the model into fractions that can be attributed to
configurations, yielding a distribution of r elementary effects per
parameters or set of parameters.
parameter. Two statistics of this distribution are then used as
Sobol indices yield rigorous, interpretable results but require
sensitivity measures: the mean of the absolute values estimates the
two conditions: parameters must be mutually independent, and the
overall influence of the parameter on the output, while the
problem must admit a scalar output. When the model produces
standard deviation serves as an indicator of nonlinear behavior and
time-series outputs, it remains possible to apply Sobol indices to
interactions with other parameters. The method requires only
scalar quantities of interest derived from the trajectory, such as the
r(p+1) model evaluations in total, which makes it tractable even
maximum, minimum, or time-averaged value. The principal
for models with a large number of parameters. It is also
practical limitation is computational cost: estimating Sobol
straightforward to implement for any black-box model. Its
indices requires a structured experimental design, typically Monte
principal limitation is that it does not quantify parameter
Carlo sampling or Latin Hypercube Sampling, and the required
interactions precisely, and its conclusions are qualitative rather
number of model evaluations scales as 𝑂(𝑟(𝑝 + 2)) for 𝑝
than quantitative.
parameters, making the analysis potentially expensive for highdimensional problems. Alternative variance-based methods, such
as the Fourier Amplitude Sensitivity Test (FAST) [13] and its
analysis between different variables can be an applied to evaluate
extended variant eFAST [13], offer reduced computational cost by
the influence between different models in an assembly.
estimating sensitivity indices through spectral analysis of the model output.
Statistical correlation approaches. Granger introduced a method in the context of economy [15]. The idea is to evaluate if two time series (𝑋𝑡 , 𝑌𝑡 ) variable are causal from a statistical point
Metamodel-based approaches. When the computational cost
of view (where 𝑡 ∈ [𝑇0 , 𝑇1 ]). This method combined a regression
of direct sampling is prohibitive, metamodel-based methods
analysis, and a statistical test; the F-test. The regression analysis
construct a surrogate that approximates the original simulator at a
consists of Find 𝐴𝑖,𝑘 (where 𝑖 ∈ {1,2}, 𝑘 ∈ ⟦1, 𝑑⟧) such as:
fraction of the cost. Sensitivity analysis is then performed on the
𝑌𝑡+1 = ∑𝑑𝑘=1 𝐴1,𝑘 𝑌𝑡−𝑘 + 𝐴2,𝑘 𝑋𝑡−𝑘
surrogate rather than on the original model. Gaussian process
Where 𝑑 is the depth. Then, If 𝐴2,𝑘 is predominant in front
regression, also known as Kriging [14], is among the most widely
of 𝐴1,𝑘 it mean that 𝑌𝑡+1 depend mainly on 𝑋𝑡−𝑘 . A F-test is
used surrogates for this purpose: it provides not only a point
performed to ensure that the variables are related. Granger
prediction but also a variance estimate that quantifies prediction
causality can be applied in the context of model assembly. Each
uncertainty. Once the surrogate is calibrated on a limited number
model on the simulation assembly, Granger-type statistical
of model evaluations, sensitivity indices can be estimated
correlation can be assessed between each input and output of the
analytically or at negligible additional cost.
model (Error! Reference source not found.).
Limitations in the context of model assembly. Applying global sensitivity analysis in a model assembly context raises two fundamental challenges that go beyond computational cost. First, most sensitivity analysis methods are formulated for scalar or static outputs and are not directly applicable to time-series models; when dealing with dynamic systems, the reduction to scalar metrics necessarily entails a loss of information about the temporal structure of the response. Second, and more fundamentally, all methods described above quantify the influence of individual parameters on an output variable. This raises a conceptual question: if every parameter of a given sub-model has some influence on the decision variable, does that imply that the sub-model itself is influential? Conversely, what if only a small subset of its parameters drives the output, does the remaining sub-model contribute meaningfully? Standard sensitivity analysis provides no direct answer to these questions, because it operates at the parameter level rather than at the model level. Addressing the contribution of a sub-model as a whole requires a different conceptual framework, one in which the object of analysis is the sub-model rather than its individual parameters. Furthermore, standard sensitivity analysis requires varying parameters across their admissible range, whereas the approach sought here aims to assess sub-model influence while keeping all
Figure 3 Example of statistical correlation study of submodel3 in a model assembly The main advantage of this method is than it requires only one simulation to get data, then the granger correlation is applied. Thus, simulation time is really limited. However, the granger causality is adapted for variables linearly correlated. which is not always the case for simulation models. In the context of regression modeling, Tibshirani introduced lasso (Least Absolute Shrinkage and Selection Operator) regression [16]. This method is a linear regression technique for time series. Let 𝑌𝑡 the variable of interest and 𝑋𝑗,𝑡 the variables dependent on 𝑌 where 𝑡 ∈ ⟦𝑇1 , 𝑇2 ⟧, 𝑗 ∈ ⟦1, 𝑛⟧, 𝑛 is the number of variables dependent on 𝑌. Lasso correlation is to find the set {𝛽𝑗 }𝑗∈⟦1,𝑛⟧ that minimize the function: min
1
𝛽0 ,𝛽1 ,… ,𝛽𝑛 2
2
𝑝 𝑛 2 ∑𝑇𝑡= 𝑇1 (𝑦𝑡 − 𝛽0 − ∑𝑗=1 β𝑗 𝑋𝑗,𝑡 ) + 𝜆 ∑𝑗=0|𝛽𝑗 |
parameters fixed at their nominal values. These two limitations
Then the importance of each variable 𝑋𝑗 in 𝑌 are order by
jointly motivate the exploration of alternative techniques for
importance of 𝛽𝑗 . This method is particularly interesting because
model assemblies.
it requires only one simulation. However, Lasso regression is adapted to identify linear dependency between variables.
3.3.Statistical approaches The study of the influence of models within a model assembly is few addressed in the literature. However, Statistical
Explainability in neural network. Recent research in the field of explainability in artificial intelligence offers promising approaches to addressing this issue.
Meyes et al introduced the concept of ablation studies in
Table 1 synthesizes the assessment of each family of approaches
Artificial Neural Networks [17]. The idea of the methodology is
against these five criteria. The evaluation is based on the analysis
to identify the part of a neural network that are the best involve in
conducted in the previous sections and reflects both the theoretical
a decision. The concept is to randomly remove neurons in neural
properties of the methods and their practical constraints when
networks to keep the part of neural network which has the greatest
applied to simulation model assemblies.
influence on the decision variable. This method seems adapted to answer our problematics, because one goal is to identify the subset
4.2.Analysis
of models which are the most influential in simulation architecture.
Several observations emerge from this comparison. First,
But it is not directly applicable, because simulation models cannot
there is a clear trade-off between objectivity and cost. The most
directly be removed from a simulation architecture.
rigorous quantitative methods — variance-based global sensitivity
Limitations in the context of model assembly. Statistical
analysis (Sobol indices) — provide the highest degree of
model methods are suitable for simple model, however, linear
objectivity, with a mathematically grounded decomposition of
correlation between different variables of a models can be
output variance, but at a computational cost that scales
undetectable by granger causality or lasso correlation. Ablation
unfavorably with the number of parameters. Conversely,
study seems to be a suitable approach. However, this method can
statistical correlation
not be adapted to physical model, because in a models assembly,
regression) require only a single simulation run, making them the
the outputs of a models are linked with input of a downstream
most efficient in terms of computational budget, but their
model. If a model would be removed, then the inputs of the
restriction to linear dependencies limits the reliability of their
downstream models will not be connected to the outputs of the
conclusions for nonlinear physical models.
upstream models.
methods (Granger
causality,
Lasso
The scope of applicability reveals a fundamental gap. All credibility assessment frameworks are designed for standalone models and provide no compositional mechanism for assemblies.
4.METHODS COMPARISON
Among the impact evaluation methods, sensitivity analysis
4.1.Comparison The preceding sections have reviewed two broad classes of
operates at the parameter level, not at the submodel level, which
simulation
creates a conceptual mismatch with the problem of assessing
architectures: credibility assessment frameworks applied to
submodel influence. Ablation studies from the AI domain address
individual models (Section 2), and impact evaluation methods —
the right conceptual question — identifying the most influential
including sensitivity analysis and statistical techniques — that
components — but cannot be directly transferred to physical
characterize the influence of submodels within an assembly
model assemblies due to interface connectivity constraints.
approaches
for
assessing
the
credibility
of
(Section 3). Each approach addresses a different facet of the
Maturity does not correlate with applicability to the
problem, and none, taken in isolation, provides a complete answer.
assembly problem. The most mature methods (credibility
This section proposes a structured comparison of the seven
frameworks, local sensitivity analysis, Sobol indices) are well-
identified families of methods across four evaluation criteria:
established for individual models but were not designed for the
objectivity and rigor, cost and execution time, scope and
compositional question. The most conceptually relevant approach
applicability, maturity and standardization.
(ablation studies) is the least mature in the simulation context,
Objectivity and Rigor evaluate whether the method produces quantitative, reproducible, and verifiable results, as opposed to
existing only at a conceptual stage without established methodology or tooling.
qualitative or declarative assessments. Cost and Execution Time
Ease of usage which is not a criterion of this table varies
capture the computational and organizational resources required,
significantly but is generally inversely related to methodological
including the number of model evaluations and the complexity of
depth. Local sensitivity analysis and statistical correlation
the workflow. Scope and Applicability consider the range of
methods are straightforward to implement for any black-box
models and architectures to which the method can be applied, and
model, including FMI-compliant components, while metamodel-
in particular whether it extends from individual models to model
based approaches and variance-based methods require specialized
assemblies. Maturity and Standardization reflect the level of
expertise in experimental design, surrogate modeling, or statistical
theoretical development, community adoption, and availability of
analysis that may not be available in all industrial teams.
standards or tooling.
Table 1 Comparison of Approaches with 4 criteria Family of Approaches
Objectivity & Rigor
Cost & Execution Time
Scope & Applicability
Model credibility Frameworks (NASA-STD-7009B, PCMM, Prostep ivip, IRT SystemX)
Primarily declarative: credibility relies on self-reported evidence.
Requires systematic documentation and expert review; cost scales with model criticality.
PCMM is mostly limited to 3D physicsbased models. Prostep and IRT SystemX are more domain-agnostic.
Well-established due to standardization, industrial use or regulations
Local Sensitivity Analysis (Finite Differences, Adjoint Methods)
Provides quantitative, reproducible derivatives. However, conclusions are strictly valid only locally and cannot be generalized across the parameter space.
Requires p+1 to 2p model evaluations for p parameters. Highly efficient, if analytical or automatic differentiation is available.
Valid only in the neighborhood of the nominal operating point. Cannot detect nonlinearities, threshold effects, or parameter interactions.
Well-established mathematical foundations. Widely implemented in all simulation environments.
Global SA — Screening (Morris, OAT, Sequential Bifurcation)
Provides a qualitative ranking of parameter influence. Morris statistics (mean and standard deviation of elementary effects) offer reproducible indicators, but results are ordinal rather than quantitative.
Morris requires r(p+1) evaluations. OAT is simpler but explores a vanishing fraction of the parameter space. Sequential bifurcation is efficient but needs prior knowledge of effect signs.
Effective as a preliminary filter for high-dimensional problems. Does not quantify interactions precisely. OAT is unsuitable for high-dimensional systems. Sequential bifurcation assumes near-linearity.
Morris (1991) and OAT are standard tools in sensitivity analysis. Sequential bifurcation is less common but well-documented in the literature.
Provides a rigorous, quantitative decomposition of output variance into first order and total-order contributions.
Sobol indices demand large sample sizes. FAST/eFAST reduce cost via spectral methods but remain expensive for a lot of parameters.
Applicable to any black-box model with independent parameters. Not directly suited for time-series outputs; requires reduction to scalar quantities of interest.
Sobol, FAST, eFAST is widely used and has been implemented in many tools
Objectivity depends on surrogate fidelity. Gaussian processes provide prediction uncertainty, enabling quantification of approximation error. Results are indirect
Initial cost for training data generation, but once calibrated, sensitivity indices can be computed analytically or at negligible marginal cost. Suitable when the original model is computationally expensive.
Applicable to any black-box model. Particularly advantageous for expensive simulators. Accuracy degrades in highdimensional parameter spaces (curse of dimensionality) unless active learning strategies are used.
Kriging is well-established in geostatistics and computer experiments. Its use for sensitivity analysis is supported by a mature body of literature and robust software implementations.
Statistical Correlation (Granger Causality, Lasso Regression)
Provides quantitative, reproducible statistical indicators (F-test for Granger, regularization path for Lasso). However, both methods assume linear relationships, limiting their objectivity for nonlinear models.
Requires only a single simulation run to generate the time-series data. Post-processing is computationally inexpensive. The most efficient family of methods in terms of simulation budget.
Effective for detecting linear dependencies between interface variables. Fails to capture nonlinear couplings, which are common in physical simulation assemblies.
Granger causality (1969) and Lasso (1996) are well-established in econometrics and statistics. Their application to simulation model assemblies is recent and not yet standardized.
Explainability in AI (Ablation Studies in Neural Networks)
The ablation protocol is reproducible, but the analogy between neurons and simulation submodels is imperfect. Removing a neuron is well-defined; removing a physical submodel from an assembly disrupts interface connectivity.
Depends on the number of ablation experiments required. In neural networks, cost is manageable; for simulation assemblies, each ablation scenario requires architectural reconfiguration and re-execution.
The idea of identifying the most influential components is directly relevant to model assemblies. However, direct application is not feasible: removing a submodel breaks interface connections, making the assembly undefined.
Ablation studies are standard practice in deep learning (since ~2019). Transfer to simulation model assemblies is at a conceptual stage; no established methodology or tooling exists for this context.
Global SA — Variance-based (Sobol Indices, FAST, eFAST)
Global SA — Metamodelbased (Kriging / Gaussian Process Regression)
Maturity & Standardization
5.DISCUSSION AND CONCLUSION This paper addressed the problem of assessing the credibility of a simulation architecture based on its constituent submodels. The comparison reveals that no single existing method adequately addresses this issue. Credibility frameworks provide a necessary but insufficient foundation: they ensure that individual models meet quality standards but cannot characterize how errors and uncertainties propagate through interfaces and accumulate at the system level. Sensitivity analysis methods offer tools for understanding parameter-level influence, but the step from parameter influence on submodel influence remains an open scientific question. Among the approaches reviewed, the conceptual framework of ablation studies appears the most aligned with the problem structure: it directly targets the identification of the most influential components within a composed system. However, its transfer from neural networks to simulation architectures is not straightforward, as removing a physical submodel disrupts interface connectivity and renders the assembly undefined. A viable adaptation could consist in substitution strategies, where a submodel is replaced by a simplified or degraded version rather than removed entirely, thereby preserving interface integrity while enabling assessment of the submodel's contribution to the architecture's output. Ultimately, the credibility of a simulation architecture cannot be reduced to the aggregation of individual model credibility. It requires a dedicated methodological framework that integrates compositional uncertainty propagation, interface characterization, and decision-relevant impact assessment — a framework that, as this review has shown, remains to be fully developed. 6.ACKNOWLEDGMENT This work has been supported by the French government under the "France 2030” program, as part of the SystemX Technological Research Institute within the AFS project. 7.REFERENCES [1] Romain Barbedienne, Julien Silande, Henri Sohier, Anthony Levillain, Cédric Leclerc, and Maxime Hayet, “Simulation-based multi-organization engineering : Simulation specification,” presented at the SIA simulation numérique, guyancourt, Apr. 2025. [2] Standard for models and simulation, NASA STD 7009B, 2024. [Online]. Available: https://standards.nasa.gov/sites/default/files/standards/NAS A/B/1/NASA-STD-7009B-Final-3-5-2024.pdf [3] W. Oberkampf, T. Trucano, and M. Pilch, “Predictive Capability Maturity Model for computational modeling and
simulation.,” SAND2007-5948, 976951, Oct. 2007. doi: 10.2172/976951. [4] Prostep ivip, “Guard Rails for ‘Simulation Credibility Standards and Recommendation,’” white paper Version 1, Mar. 2024. Accessed: Apr. 05, 2026. [Online]. Available: https://www.prostep.org/fileadmin/proddownload/TechnicalPaper_SimulationCredibility_2024_V8.1_v2.pdf [5] Romain Barbedienne et al., “Simulation-based multiorganization engineering,” IRT SystemX, White paper, Apr. 2025. [6] European commission, laying down rules for the application of Regulation (EU) 2019/2144 of the European Parliament and of the Council as regards uniform procedures and technical specifications for the typeapproval of the automated driving system (ADS) of fully automated vehicles, vol. L 221/1. 2022. Accessed: Apr. 05, 2026. [Online]. Available: https://eur-lex.europa.eu/legalcontent/EN/TXT/PDF/?uri=CELEX:32022R1426 [7] M. Paruggia, “Sensitivity Analysis in Practice: A Guide to Assessing Scientific Models,” J. Am. Stat. Assoc., vol. 101, no. 473, pp. 398–399, 2006, doi: 10.1198/jasa.2006.s80. [8] D. Cacuci, Sensitivity and Uncertainty Analysis, Volume I: Theory, vol. 1. 2003. doi: 10.1201/9780203498798. [9] M. D. Morris, “Factorial Sampling Plans for Preliminary Computational Experiments,” Technometrics, vol. 33, no. 2, pp. 161–174, 1991, doi: 10.1080/00401706.1991.10484804. [10] H. Sohier, H. Piet-Lahanier, and J.-L. Farges, “Analysis and optimization of an air-launch-to-orbit separation,” Acta Astronaut., vol. 108, pp. 18–29, 2015, doi: https://doi.org/10.1016/j.actaastro.2014.11.043. [11] A. Saltelli and P. Annoni, “How to avoid a perfunctory sensitivity analysis,” Environ. Model. Softw., vol. 25, no. 12, pp. 1508–1517, 2010, doi: https://doi.org/10.1016/j.envsoft.2010.04.012. [12] B. Bettonvil and J. P. C. Kleijnen, “Searching for important factors in simulation models with many factors: Sequential bifurcation,” Eur. J. Oper. Res., vol. 96, no. 1, pp. 180–194, 1997, doi: https://doi.org/10.1016/S0377-2217(96)00156-7. [13] Andrea Saltelli, Stefano Tarantola, Francesca Campolongo, and Marco Ratto, “Methods Based on Decomposing the Variance of the Output,” in Sensitivity Analysis in Practice, John Wiley & Sons, Ltd, 2002, pp. 109–149. doi: https://doi.org/10.1002/0470870958.ch5. [14] J. P. C. Kleijnen, “KRIGING METAMODELING IN SIMULATION: A REVIEW”. [15] A. Shojaie and E. B. Fox, “Granger causality: A review and recent advances,” Annu. Rev. Stat. Its Appl., vol. 9, pp. 289–319, 2022. [16] R. Tibshirani, “Regression Shrinkage and Selection Via the Lasso,” J. R. Stat. Soc. Ser. B Stat. Methodol., vol. 58, no. 1, pp. 267–288, Jan. 1996, doi: 10.1111/j.25176161.1996.tb02080.x. [17] R. Meyes, M. Lu, C. W. de Puiseau, and T. Meisen, “Ablation Studies in Artificial Neural Networks,” Feb. 18, 2019, arXiv: arXiv:1901.08644. doi: 10.48550/arXiv.1901.08644.