ConceptioArchivearXiv CS
arXiv CSopen access

Why Model Credibility Isn't Enough: -Rethinking Trust in Simulation Architectures

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
softwarearchitecturesoftwareengineeringtesting
software engineering, software architecture, testing

Why Model Credibility Isn’t Enough: - Rethinking Trust in Simulation Architectures Romain Barbedienne 1) Adeline Lanugue 1) Rim Kaddah 1) Julien Silande 1,2) Anthony Levillain 3) Cedric Leclerc 4) Maxime Hayet 5) Boussaad Soualmi 1) Cristian Maxim 1) 1) IRT SystemX, 2 Bd Thomas Gobert, 91120 Palaiseau, France (E-mail: [email protected]) 2) Keysight Technology, La janais 3 rue Pierre et marie curie, 35131 Chartres-de-Bretagne, France 3) OPmobility Alphatech, ZAC du Bois de Plaisance, 214 Av. de la Mare Gessart, 60280 Venette, France 4)Renault Technocentre, 1 avenue du Golf, 78280 Guyancourt, France 5)Stellantis Green-Campus, 43 Rue Jean Pierre Timbaud, 78300 Poissy, France ABSTRACT: Credibility of a simulation model is an important topic. Several approaches try to quantify the credibility of simulation. However, models are mostly assembled within a simulation architecture. Can the credibility of a simulation architecture be assessed based on the credibility of the models that comprise it? This paper aims to address this issue by providing an overview of the current state of the art in the field of assembly credibility. It will compare sensitivity analysis techniques, qualitative analysis by experts, explainability in AI, and networks. Finally, an assessment of the proposed approaches, based on criteria such as rigor, generalization, and resource requirements, will reveal the strengths and weaknesses of each approach. KEY WORDS: Credibility, Model Architecture, Sensitivity Analysis 1.INTRODUCTION

These submodels can sometimes even be delivered by suppliers

Designing a system requires ensuring that it operates

[1]. In this context, is it possible to evaluate the credibility of a

correctly while guaranteeing safety appropriate for its intended

simulation architecture based on the credibility of the models that

use. In recent years, computer simulation has become widely used

comprise it?

in the industrial sector to validate these functional and safety

This paper is a analyze different solution that aimed at

requirements. Indeed, for complex systems, simulation is less

evaluating scientific approaches to addressing this issue. The first

expensive than conducting real-life tests. Thereby, simulation

part of this paper is focus on the methods that allows to evaluate

leads to anticipate and test a system as soon as possible to enhance

credibility of a simulation model. We will discuss the direct

it. Moreover, it allows to realize a lot of tests in a short time and

application of these methods to the evaluation of simulation

quickly improve the system. However, to be pertinent and prevent

architectures. The second part will explore the methods that can

errors in decision making, these tests should be done in a model

be used to evaluate the impact of a model in a simulation

which is enough credible and robust.

architecture. The third part is an analysis and a comparison of

In the M&S domain, credibility could be defined as the

different methods.

degree of confidence that a stakeholder places in a product for a specific purpose based on established independent evidence. To

2.Evaluation of simulation model credibility

ensure the credibility of a model, it is necessary to conduct efforts

The credibility of simulation model is characterized by the

to ensure that the system is properly designed and that its

degree of confidence that a stakeholder places in a product for a

verification and validation have been carried out correctly. The

specific purpose based on established independent evidence. Thus,

credibility of a simulation model is essential to ensure the right

assessing the credibility of a model is subjective; it cannot be

decision. Improving the credibility of a simulation model is an

measured using precise metrics. However, several studies have

essential step, but it can be very time-consuming. It is therefore

focused on rationalizing credibility.

necessary to strike the right balance between improving the model's credibility and the time spent on this task. In practice, for decisions, it is necessary to combine

2.1.Method for ensuring the credibility of models

models from different teams. Engineers build model architectures

NASA has introduced Standard 7009B [2]. This standard

composed of several submodels. These models are connected to

defines uniform technical and engineering requirements for the

one another via interfaces. Each interface contains one or more

processes, procedures, practices, and methods recognized as

variables that are exchanged between the submodels. The

standards in modeling and simulation. One of the key elements in

submodels constituting this architecture have their own credibility.

assessing modeling and simulation capabilities is the set of

recommendations covering the following topics: data traceability,

concern the credibility assessment when using virtual chains tools

verification, validation, technical review, and product/process

to validate vehicle. Some rules are linked with the experiences of

management. Each requirement must be met throughout a model’s

the team, others about proofs, and the traceability in a

lifecycle, which adds complexity to the design process since the

documentation. To respect these rules, as soon as possible, in the

requirements must be reviewed with every modification. This

V-Model enhance the credibility of the model and the system.

enhances the model’s credibility but is particularly suited for high-

2.4. Comparison and discussion

criticality simulation models.

The five approaches have the common objective of rationalizing credibility assessment but differ in scope and rigor.

2.2.Frameworks for credibility assessment

PCMM and NASA-STD-7009B are the most scientifically

Oberkampf et al. [3] introduced the Predictive Capability

grounded, addressing respectively the technical foundations of a

Maturity Model (PCMM). This framework evaluates the maturity

model. The Prostep ivip and IRT SystemX frameworks adopt a

of a simulation model across several dimensions: representation

more pragmatic stance, offering structured tools that facilitate

and geometric fidelity, physics and material model fidelity, code

communication between stakeholders in industrial contexts.

verification,

solution

uncertainty

quantification

verification,

model

validation,

and

However, these approaches remain largely declarative: credibility

sensitivity

analysis.

Each

relies on self-reported information, with limited mechanisms to

dimension is defined for different maturity levels. PCMM allows

independently verify the claims made — a significant weakness

for a quick assessment of a model’s strengths and weaknesses.

when models are delivered by external suppliers. The European

One limitation of this approach is that it is primarily applicable to

Commission regulation provides a useful normative baseline but

3D models.

stays general and qualitative, oriented toward certification

and

Prostep ivip publishes a white paper [4] where they explain the

compliance rather than scientific rigor.

emphasizes the need to standardize the assessment of simulation

A critical limitation shared by all five frameworks is that they

credibility by incorporating criteria related to quality, accuracy,

are designed to assess standalone simulation models, and their

and criticality, to strengthen stakeholders’ confidence in the

extension to a simulation architecture raises unresolved scientific

results obtained. This document introduces a framework for

challenges.

assessing the credibility of processes. Within this framework, the

experimental reference data, which is often unavailable at the

Wheel of V&V and the Credibility Spider Chart help highlight the

architecture level — particularly in early design phases or in

strengths and weaknesses of a simulation model, enabling

OEM–supplier relationships where test data cannot be shared.

stakeholders to quickly assess the model’s credibility. This highly

Furthermore, none of these frameworks provides a systematic

promising work needs to be extended to other areas of verification

method to address compositional aspects: even if each submodel

and validation; indeed, the credibility of a model also depends on

is individually credible, the propagation of uncertainties across

its design and its specifications.

interfaces and emergent behaviors arising from model coupling

Most

validation

methods

require

access

to

Finally, IRT SystemX publishes a white paper [5] to improve

remain largely uncharacterized. Addressing these issues requires

the interaction between OEM and supplier in the context of

developing methods capable of aggregating submodel credibility

simulation. They provide a questionnaire designed to assess the

and reducing reliance on declarative assessments in favor of more

credibility of simulation models across five key areas: robustness

objective, verifiable evidence. In order to assess the credibility of

and sensitivity, model uncertainty and margin, verification,

a model assembly based on its submodels, the remainder of this

validation, and use. This questionnaire is summarized in a memo

paper will focus on methods for evaluating the impact of a

that provides an overview of the level of rigor applied to each

simulation model within a model assembly.

model. A minimum score is required for each axis, based on the criticality of the simulation, the system design cycle, and whether or not the model needs to be integrated into a simulation architecture.

3.Evaluation of impact of a model in an Assembly 3.1.Problem statement A simulation architecture is composed of submodels connected through interfaces, each submodel carrying its own credibility

2.3. Regulations

level Figure 1. Even when each submodel is found to be

The European commission has established a regulation aimed

sufficiently credible individually, this does not imply that the

at establishing guidelines for the certification of vehicles capable

whole assembly is credible, as allowable errors from submodels

of operating without human intervention [6]. The four parts

may accumulate to unacceptable levels at the system level. Rather

than attempting a direct compositional aggregation of model

3.2.Sensitivity analysis

credibility an alternative strategy consists in evaluating the impact

Sensitivity analysis is a broad family of methods aimed at

of each submodel on the architecture's output variables of interest.

determining how variations in model inputs or parameters

This impact-based approach reframes the problem: instead of

propagate to model outputs. Depending on whether the analysis is

asking whether a submodel is credible in isolation, it asks how

conducted at a single operating point or across the entire parameter

much a submodel's uncertainty or error propagates to the decision-

space, these methods are generally classified as either local or

relevant outputs of the architecture.

global. The following subsections review both approaches, discussing their respective theoretical foundations, practical applicability, and limitations in the context of simulation model assemblies. 3.2.1.

Local sensitivity analysis Local sensitivity analysis characterizes how small

perturbations in model parameters affect model outputs in the vicinity of a nominal operating point [7]. Formally, given a model Figure 1 Credibility of simulation architecture This analysis is inherently context dependent. The validity of a model is defined over a specific domain of model form, inputs, parameters, and responses, which effectively limits its use to the particular application for which it was validated. Thus, the

F:ℝ𝑝 → ℝ𝑞 mapping a parameter vector Λ ∈ ℝ𝑝 to an output vector 𝑌 ∈ ℝ𝑞 , the local sensitivity is defined as the partial derivative of model output 𝑦𝑖 related to parameter 𝜆𝑗 evaluated at a nominal point Λ1 (where 𝑖 ∈ ⟦1, 𝑞⟧ and 𝑗 ∈ ⟦1, 𝑝⟧): 𝜕𝑦𝑖 | Sij = 𝜕𝜆𝑗 Λ=Λ1

contribution of a given submodel to the architecture's output

In practice, when analytical derivatives are unavailable, this

depends on the scenarios and parameter sets used for decision-

quantity is approximated numerically using finite difference

making. Two submodels may have very different impacts

schemes [8].

depending on the operating conditions simulated. This implies that

From a computational standpoint, local sensitivity analysis is

the impact analysis must be conducted specifically for the

highly efficient. For a model with 𝑝 parameters, a full sensitivity

scenarios and parameter ranges used for the decision making.

matrix requires only 𝑝 + 1 model evaluations using a forward

Furthermore, a model may be valid for one set of experimental

difference scheme, or at most 2p evaluations with a centered

conditions and invalid in another; therefore, during the impact

difference scheme. This makes the approach tractable even for

assessment, input variables and parameters of each submodel must

models with a large number of parameters, provided each

remain strictly within their respective validation domains to

evaluation is not prohibitively expensive. When derivatives can be

ensure that the conclusions drawn are meaningful. Finally, one last

computed analytically or via automatic differentiation, the cost

consideration to keep in mind is that simulation models are often

reduces further, requiring only a single forward pass through the

compiled in the Functional Mock-up Interface (FMI) format. This facilitates co-simulation, but also allows models to be exchanged

model alongside its linearization. However, local sensitivity analysis rests on a fundamental

between suppliers and OEMs without the risk of sharing

assumption: that the model behavior is approximately linear in the

proprietary know-how.

neighborhood of the nominal point Λ1 . This assumption carries

To systematically evaluate the contribution of each submodel

several important consequences. First, the method is strictly valid

to the architecture's outputs, two complementary families of

only when parameter variations remain small relative to the scale

methods will be explored in this paper. Sensitivity analysis

over which the model's behavior changes. Second, and more

methods are useful for testing the robustness of results, identifying

critically, local sensitivity indices are entirely dependent on the

model inputs that cause significant impact in the output. Moreover,

choice of nominal point, meaning that the conclusions drawn may

statistical methods offer a complementary perspective by

not be representative of the model's behavior across its full

characterizing the probabilistic distribution of outputs and

parameter range. In the presence of nonlinearities, the sensitivity

quantifying uncertainty propagation across the architecture. These two approaches will be explored in the following sections.

structure can vary dramatically across the parameter space, rendering point-wise estimates potentially misleading. Another example of limitation of local sensitivity is the failure to distinguish between a local minimum and an inflection point. This

difference is, however, crucial for determining the robustness of a model Figure 2.

The Morris method is frequently used as a first step to identify the most influential parameters before proceeding to a more resource-intensive analysis. At this stage, it is critical not to underestimate the influence of any parameter, as a false negative (incorrectly classifying an influential parameter as negligible) would propagate through all subsequent analyses. To address this risk, Sohier et al. [10] proposed an adaptive variant of the Morris method in which the number of elementary effect samples is increased selectively for parameters with low estimated influence,

𝑑𝐹

Figure 2 : Both cases yield 𝑑𝜆𝑖 (𝜆1 ) ≈ 0 so local analysis alone cannot distinguish them

thereby reducing the probability of misclassification while controlling the total number of model evaluations. Two further screening approaches are worth mentioning. One-

3.2.2.

Global sensitivity analysis

At-a-Time (OAT) analysis [11] varies each parameter individually

Unlike local sensitivity analysis, which characterizes model

across its full range while holding all others fixed, and plots the

behavior at a single nominal point, global sensitivity analysis

resulting output response to reveal the form of each parameter's

evaluates the influence of parameters over their entire admissible

influence (including nonlinearities and threshold effects). It is

range. By exploring the full parameter space rather than a local

simpler than the Morris method but shares its inability to detect

neighborhood, this family of methods provides a more robust and

interactions, and the fraction of parameter space explored shrinks

representative characterization of model behavior, in particular

super-exponentially with the number of parameters, making it

when nonlinearities or parameter interactions are present. Global

unsuitable for high-dimensional problems. Sequential bifurcation

sensitivity analysis methods are generally organized into three

[12] takes a different approach: parameters are grouped and tested

main families: screening-based approaches, variance-based

in aggregate, with groups progressively subdivided until

approaches, and metamodel-based approaches

individual influential parameters are isolated. It is both effective and efficient, requiring relatively few simulations run, but

Screening-based approaches. The objective of screening

assumes that the model response can be approximated by a first-

methods is to rank parameters by order of influence at low

order polynomial and that the signs of parameter effects are known

computational cost, typically as a preliminary step before a more

a priori, which limits its applicability to FMI models.

expensive analysis. The most widely used method in this family is the Morris method [9]. The Morris method proceeds as follows.

Variance-based

approaches.

Variance-based

methods

Starting from a randomly sampled initial configuration in the

provide a quantitative decomposition of output variance into

parameter space, each parameter is perturbed individually by a

contributions attributable to each parameter and to their

fixed step, and the corresponding elementary effect (defined as the

interactions. The most rigorous framework in this family is that of

ratio of the output variation to the parameter perturbation) is

Sobol indices [13]. This framework decomposes the variance of

recorded. This procedure is repeated for r independently sampled

the output of the model into fractions that can be attributed to

configurations, yielding a distribution of r elementary effects per

parameters or set of parameters.

parameter. Two statistics of this distribution are then used as

Sobol indices yield rigorous, interpretable results but require

sensitivity measures: the mean of the absolute values estimates the

two conditions: parameters must be mutually independent, and the

overall influence of the parameter on the output, while the

problem must admit a scalar output. When the model produces

standard deviation serves as an indicator of nonlinear behavior and

time-series outputs, it remains possible to apply Sobol indices to

interactions with other parameters. The method requires only

scalar quantities of interest derived from the trajectory, such as the

r(p+1) model evaluations in total, which makes it tractable even

maximum, minimum, or time-averaged value. The principal

for models with a large number of parameters. It is also

practical limitation is computational cost: estimating Sobol

straightforward to implement for any black-box model. Its

indices requires a structured experimental design, typically Monte

principal limitation is that it does not quantify parameter

Carlo sampling or Latin Hypercube Sampling, and the required

interactions precisely, and its conclusions are qualitative rather

number of model evaluations scales as 𝑂(𝑟(𝑝 + 2)) for 𝑝

than quantitative.

parameters, making the analysis potentially expensive for highdimensional problems. Alternative variance-based methods, such

as the Fourier Amplitude Sensitivity Test (FAST) [13] and its

analysis between different variables can be an applied to evaluate

extended variant eFAST [13], offer reduced computational cost by

the influence between different models in an assembly.

estimating sensitivity indices through spectral analysis of the model output.

Statistical correlation approaches. Granger introduced a method in the context of economy [15]. The idea is to evaluate if two time series (𝑋𝑡 , 𝑌𝑡 ) variable are causal from a statistical point

Metamodel-based approaches. When the computational cost

of view (where 𝑡 ∈ [𝑇0 , 𝑇1 ]). This method combined a regression

of direct sampling is prohibitive, metamodel-based methods

analysis, and a statistical test; the F-test. The regression analysis

construct a surrogate that approximates the original simulator at a

consists of Find 𝐴𝑖,𝑘 (where 𝑖 ∈ {1,2}, 𝑘 ∈ ⟦1, 𝑑⟧) such as:

fraction of the cost. Sensitivity analysis is then performed on the

𝑌𝑡+1 = ∑𝑑𝑘=1 𝐴1,𝑘 𝑌𝑡−𝑘 + 𝐴2,𝑘 𝑋𝑡−𝑘

surrogate rather than on the original model. Gaussian process

Where 𝑑 is the depth. Then, If 𝐴2,𝑘 is predominant in front

regression, also known as Kriging [14], is among the most widely

of 𝐴1,𝑘 it mean that 𝑌𝑡+1 depend mainly on 𝑋𝑡−𝑘 . A F-test is

used surrogates for this purpose: it provides not only a point

performed to ensure that the variables are related. Granger

prediction but also a variance estimate that quantifies prediction

causality can be applied in the context of model assembly. Each

uncertainty. Once the surrogate is calibrated on a limited number

model on the simulation assembly, Granger-type statistical

of model evaluations, sensitivity indices can be estimated

correlation can be assessed between each input and output of the

analytically or at negligible additional cost.

model (Error! Reference source not found.).

Limitations in the context of model assembly. Applying global sensitivity analysis in a model assembly context raises two fundamental challenges that go beyond computational cost. First, most sensitivity analysis methods are formulated for scalar or static outputs and are not directly applicable to time-series models; when dealing with dynamic systems, the reduction to scalar metrics necessarily entails a loss of information about the temporal structure of the response. Second, and more fundamentally, all methods described above quantify the influence of individual parameters on an output variable. This raises a conceptual question: if every parameter of a given sub-model has some influence on the decision variable, does that imply that the sub-model itself is influential? Conversely, what if only a small subset of its parameters drives the output, does the remaining sub-model contribute meaningfully? Standard sensitivity analysis provides no direct answer to these questions, because it operates at the parameter level rather than at the model level. Addressing the contribution of a sub-model as a whole requires a different conceptual framework, one in which the object of analysis is the sub-model rather than its individual parameters. Furthermore, standard sensitivity analysis requires varying parameters across their admissible range, whereas the approach sought here aims to assess sub-model influence while keeping all

Figure 3 Example of statistical correlation study of submodel3 in a model assembly The main advantage of this method is than it requires only one simulation to get data, then the granger correlation is applied. Thus, simulation time is really limited. However, the granger causality is adapted for variables linearly correlated. which is not always the case for simulation models. In the context of regression modeling, Tibshirani introduced lasso (Least Absolute Shrinkage and Selection Operator) regression [16]. This method is a linear regression technique for time series. Let 𝑌𝑡 the variable of interest and 𝑋𝑗,𝑡 the variables dependent on 𝑌 where 𝑡 ∈ ⟦𝑇1 , 𝑇2 ⟧, 𝑗 ∈ ⟦1, 𝑛⟧, 𝑛 is the number of variables dependent on 𝑌. Lasso correlation is to find the set {𝛽𝑗 }𝑗∈⟦1,𝑛⟧ that minimize the function: min

1

𝛽0 ,𝛽1 ,… ,𝛽𝑛 2

2

𝑝 𝑛 2 ∑𝑇𝑡= 𝑇1 (𝑦𝑡 − 𝛽0 − ∑𝑗=1 β𝑗 𝑋𝑗,𝑡 ) + 𝜆 ∑𝑗=0|𝛽𝑗 |

parameters fixed at their nominal values. These two limitations

Then the importance of each variable 𝑋𝑗 in 𝑌 are order by

jointly motivate the exploration of alternative techniques for

importance of 𝛽𝑗 . This method is particularly interesting because

model assemblies.

it requires only one simulation. However, Lasso regression is adapted to identify linear dependency between variables.

3.3.Statistical approaches The study of the influence of models within a model assembly is few addressed in the literature. However, Statistical

Explainability in neural network. Recent research in the field of explainability in artificial intelligence offers promising approaches to addressing this issue.

Meyes et al introduced the concept of ablation studies in

Table 1 synthesizes the assessment of each family of approaches

Artificial Neural Networks [17]. The idea of the methodology is

against these five criteria. The evaluation is based on the analysis

to identify the part of a neural network that are the best involve in

conducted in the previous sections and reflects both the theoretical

a decision. The concept is to randomly remove neurons in neural

properties of the methods and their practical constraints when

networks to keep the part of neural network which has the greatest

applied to simulation model assemblies.

influence on the decision variable. This method seems adapted to answer our problematics, because one goal is to identify the subset

4.2.Analysis

of models which are the most influential in simulation architecture.

Several observations emerge from this comparison. First,

But it is not directly applicable, because simulation models cannot

there is a clear trade-off between objectivity and cost. The most

directly be removed from a simulation architecture.

rigorous quantitative methods — variance-based global sensitivity

Limitations in the context of model assembly. Statistical

analysis (Sobol indices) — provide the highest degree of

model methods are suitable for simple model, however, linear

objectivity, with a mathematically grounded decomposition of

correlation between different variables of a models can be

output variance, but at a computational cost that scales

undetectable by granger causality or lasso correlation. Ablation

unfavorably with the number of parameters. Conversely,

study seems to be a suitable approach. However, this method can

statistical correlation

not be adapted to physical model, because in a models assembly,

regression) require only a single simulation run, making them the

the outputs of a models are linked with input of a downstream

most efficient in terms of computational budget, but their

model. If a model would be removed, then the inputs of the

restriction to linear dependencies limits the reliability of their

downstream models will not be connected to the outputs of the

conclusions for nonlinear physical models.

upstream models.

methods (Granger

causality,

Lasso

The scope of applicability reveals a fundamental gap. All credibility assessment frameworks are designed for standalone models and provide no compositional mechanism for assemblies.

4.METHODS COMPARISON

Among the impact evaluation methods, sensitivity analysis

4.1.Comparison The preceding sections have reviewed two broad classes of

operates at the parameter level, not at the submodel level, which

simulation

creates a conceptual mismatch with the problem of assessing

architectures: credibility assessment frameworks applied to

submodel influence. Ablation studies from the AI domain address

individual models (Section 2), and impact evaluation methods —

the right conceptual question — identifying the most influential

including sensitivity analysis and statistical techniques — that

components — but cannot be directly transferred to physical

characterize the influence of submodels within an assembly

model assemblies due to interface connectivity constraints.

approaches

for

assessing

the

credibility

of

(Section 3). Each approach addresses a different facet of the

Maturity does not correlate with applicability to the

problem, and none, taken in isolation, provides a complete answer.

assembly problem. The most mature methods (credibility

This section proposes a structured comparison of the seven

frameworks, local sensitivity analysis, Sobol indices) are well-

identified families of methods across four evaluation criteria:

established for individual models but were not designed for the

objectivity and rigor, cost and execution time, scope and

compositional question. The most conceptually relevant approach

applicability, maturity and standardization.

(ablation studies) is the least mature in the simulation context,

Objectivity and Rigor evaluate whether the method produces quantitative, reproducible, and verifiable results, as opposed to

existing only at a conceptual stage without established methodology or tooling.

qualitative or declarative assessments. Cost and Execution Time

Ease of usage which is not a criterion of this table varies

capture the computational and organizational resources required,

significantly but is generally inversely related to methodological

including the number of model evaluations and the complexity of

depth. Local sensitivity analysis and statistical correlation

the workflow. Scope and Applicability consider the range of

methods are straightforward to implement for any black-box

models and architectures to which the method can be applied, and

model, including FMI-compliant components, while metamodel-

in particular whether it extends from individual models to model

based approaches and variance-based methods require specialized

assemblies. Maturity and Standardization reflect the level of

expertise in experimental design, surrogate modeling, or statistical

theoretical development, community adoption, and availability of

analysis that may not be available in all industrial teams.

standards or tooling.

Table 1 Comparison of Approaches with 4 criteria Family of Approaches

Objectivity & Rigor

Cost & Execution Time

Scope & Applicability

Model credibility Frameworks (NASA-STD-7009B, PCMM, Prostep ivip, IRT SystemX)

Primarily declarative: credibility relies on self-reported evidence.

Requires systematic documentation and expert review; cost scales with model criticality.

PCMM is mostly limited to 3D physicsbased models. Prostep and IRT SystemX are more domain-agnostic.

Well-established due to standardization, industrial use or regulations

Local Sensitivity Analysis (Finite Differences, Adjoint Methods)

Provides quantitative, reproducible derivatives. However, conclusions are strictly valid only locally and cannot be generalized across the parameter space.

Requires p+1 to 2p model evaluations for p parameters. Highly efficient, if analytical or automatic differentiation is available.

Valid only in the neighborhood of the nominal operating point. Cannot detect nonlinearities, threshold effects, or parameter interactions.

Well-established mathematical foundations. Widely implemented in all simulation environments.

Global SA — Screening (Morris, OAT, Sequential Bifurcation)

Provides a qualitative ranking of parameter influence. Morris statistics (mean and standard deviation of elementary effects) offer reproducible indicators, but results are ordinal rather than quantitative.

Morris requires r(p+1) evaluations. OAT is simpler but explores a vanishing fraction of the parameter space. Sequential bifurcation is efficient but needs prior knowledge of effect signs.

Effective as a preliminary filter for high-dimensional problems. Does not quantify interactions precisely. OAT is unsuitable for high-dimensional systems. Sequential bifurcation assumes near-linearity.

Morris (1991) and OAT are standard tools in sensitivity analysis. Sequential bifurcation is less common but well-documented in the literature.

Provides a rigorous, quantitative decomposition of output variance into first order and total-order contributions.

Sobol indices demand large sample sizes. FAST/eFAST reduce cost via spectral methods but remain expensive for a lot of parameters.

Applicable to any black-box model with independent parameters. Not directly suited for time-series outputs; requires reduction to scalar quantities of interest.

Sobol, FAST, eFAST is widely used and has been implemented in many tools

Objectivity depends on surrogate fidelity. Gaussian processes provide prediction uncertainty, enabling quantification of approximation error. Results are indirect

Initial cost for training data generation, but once calibrated, sensitivity indices can be computed analytically or at negligible marginal cost. Suitable when the original model is computationally expensive.

Applicable to any black-box model. Particularly advantageous for expensive simulators. Accuracy degrades in highdimensional parameter spaces (curse of dimensionality) unless active learning strategies are used.

Kriging is well-established in geostatistics and computer experiments. Its use for sensitivity analysis is supported by a mature body of literature and robust software implementations.

Statistical Correlation (Granger Causality, Lasso Regression)

Provides quantitative, reproducible statistical indicators (F-test for Granger, regularization path for Lasso). However, both methods assume linear relationships, limiting their objectivity for nonlinear models.

Requires only a single simulation run to generate the time-series data. Post-processing is computationally inexpensive. The most efficient family of methods in terms of simulation budget.

Effective for detecting linear dependencies between interface variables. Fails to capture nonlinear couplings, which are common in physical simulation assemblies.

Granger causality (1969) and Lasso (1996) are well-established in econometrics and statistics. Their application to simulation model assemblies is recent and not yet standardized.

Explainability in AI (Ablation Studies in Neural Networks)

The ablation protocol is reproducible, but the analogy between neurons and simulation submodels is imperfect. Removing a neuron is well-defined; removing a physical submodel from an assembly disrupts interface connectivity.

Depends on the number of ablation experiments required. In neural networks, cost is manageable; for simulation assemblies, each ablation scenario requires architectural reconfiguration and re-execution.

The idea of identifying the most influential components is directly relevant to model assemblies. However, direct application is not feasible: removing a submodel breaks interface connections, making the assembly undefined.

Ablation studies are standard practice in deep learning (since ~2019). Transfer to simulation model assemblies is at a conceptual stage; no established methodology or tooling exists for this context.

Global SA — Variance-based (Sobol Indices, FAST, eFAST)

Global SA — Metamodelbased (Kriging / Gaussian Process Regression)

Maturity & Standardization

5.DISCUSSION AND CONCLUSION This paper addressed the problem of assessing the credibility of a simulation architecture based on its constituent submodels. The comparison reveals that no single existing method adequately addresses this issue. Credibility frameworks provide a necessary but insufficient foundation: they ensure that individual models meet quality standards but cannot characterize how errors and uncertainties propagate through interfaces and accumulate at the system level. Sensitivity analysis methods offer tools for understanding parameter-level influence, but the step from parameter influence on submodel influence remains an open scientific question. Among the approaches reviewed, the conceptual framework of ablation studies appears the most aligned with the problem structure: it directly targets the identification of the most influential components within a composed system. However, its transfer from neural networks to simulation architectures is not straightforward, as removing a physical submodel disrupts interface connectivity and renders the assembly undefined. A viable adaptation could consist in substitution strategies, where a submodel is replaced by a simplified or degraded version rather than removed entirely, thereby preserving interface integrity while enabling assessment of the submodel's contribution to the architecture's output. Ultimately, the credibility of a simulation architecture cannot be reduced to the aggregation of individual model credibility. It requires a dedicated methodological framework that integrates compositional uncertainty propagation, interface characterization, and decision-relevant impact assessment — a framework that, as this review has shown, remains to be fully developed. 6.ACKNOWLEDGMENT This work has been supported by the French government under the "France 2030” program, as part of the SystemX Technological Research Institute within the AFS project. 7.REFERENCES [1] Romain Barbedienne, Julien Silande, Henri Sohier, Anthony Levillain, Cédric Leclerc, and Maxime Hayet, “Simulation-based multi-organization engineering : Simulation specification,” presented at the SIA simulation numérique, guyancourt, Apr. 2025. [2] Standard for models and simulation, NASA STD 7009B, 2024. [Online]. Available: https://standards.nasa.gov/sites/default/files/standards/NAS A/B/1/NASA-STD-7009B-Final-3-5-2024.pdf [3] W. Oberkampf, T. Trucano, and M. Pilch, “Predictive Capability Maturity Model for computational modeling and

simulation.,” SAND2007-5948, 976951, Oct. 2007. doi: 10.2172/976951. [4] Prostep ivip, “Guard Rails for ‘Simulation Credibility Standards and Recommendation,’” white paper Version 1, Mar. 2024. Accessed: Apr. 05, 2026. [Online]. Available: https://www.prostep.org/fileadmin/proddownload/TechnicalPaper_SimulationCredibility_2024_V8.1_v2.pdf [5] Romain Barbedienne et al., “Simulation-based multiorganization engineering,” IRT SystemX, White paper, Apr. 2025. [6] European commission, laying down rules for the application of Regulation (EU) 2019/2144 of the European Parliament and of the Council as regards uniform procedures and technical specifications for the typeapproval of the automated driving system (ADS) of fully automated vehicles, vol. L 221/1. 2022. Accessed: Apr. 05, 2026. [Online]. Available: https://eur-lex.europa.eu/legalcontent/EN/TXT/PDF/?uri=CELEX:32022R1426 [7] M. Paruggia, “Sensitivity Analysis in Practice: A Guide to Assessing Scientific Models,” J. Am. Stat. Assoc., vol. 101, no. 473, pp. 398–399, 2006, doi: 10.1198/jasa.2006.s80. [8] D. Cacuci, Sensitivity and Uncertainty Analysis, Volume I: Theory, vol. 1. 2003. doi: 10.1201/9780203498798. [9] M. D. Morris, “Factorial Sampling Plans for Preliminary Computational Experiments,” Technometrics, vol. 33, no. 2, pp. 161–174, 1991, doi: 10.1080/00401706.1991.10484804. [10] H. Sohier, H. Piet-Lahanier, and J.-L. Farges, “Analysis and optimization of an air-launch-to-orbit separation,” Acta Astronaut., vol. 108, pp. 18–29, 2015, doi: https://doi.org/10.1016/j.actaastro.2014.11.043. [11] A. Saltelli and P. Annoni, “How to avoid a perfunctory sensitivity analysis,” Environ. Model. Softw., vol. 25, no. 12, pp. 1508–1517, 2010, doi: https://doi.org/10.1016/j.envsoft.2010.04.012. [12] B. Bettonvil and J. P. C. Kleijnen, “Searching for important factors in simulation models with many factors: Sequential bifurcation,” Eur. J. Oper. Res., vol. 96, no. 1, pp. 180–194, 1997, doi: https://doi.org/10.1016/S0377-2217(96)00156-7. [13] Andrea Saltelli, Stefano Tarantola, Francesca Campolongo, and Marco Ratto, “Methods Based on Decomposing the Variance of the Output,” in Sensitivity Analysis in Practice, John Wiley & Sons, Ltd, 2002, pp. 109–149. doi: https://doi.org/10.1002/0470870958.ch5. [14] J. P. C. Kleijnen, “KRIGING METAMODELING IN SIMULATION: A REVIEW”. [15] A. Shojaie and E. B. Fox, “Granger causality: A review and recent advances,” Annu. Rev. Stat. Its Appl., vol. 9, pp. 289–319, 2022. [16] R. Tibshirani, “Regression Shrinkage and Selection Via the Lasso,” J. R. Stat. Soc. Ser. B Stat. Methodol., vol. 58, no. 1, pp. 267–288, Jan. 1996, doi: 10.1111/j.25176161.1996.tb02080.x. [17] R. Meyes, M. Lu, C. W. de Puiseau, and T. Meisen, “Ablation Studies in Artificial Neural Networks,” Feb. 18, 2019, arXiv: arXiv:1901.08644. doi: 10.48550/arXiv.1901.08644.

Related documents

Record · ID 282853 · SHA-256 d52d33619457f449
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.