Fast Bayesian equipment condition monitoring via simulation based inference: applications to heat exchanger health Peter Collett,1 Alexander Johannes Stasik, ,2, 3, ∗ Simone Casolo ,1, † and Signe Riemer-Sørensen 3 1
arXiv:2604.20735v1 [cs.LG] 22 Apr 2026
2
Cognite AS, Oslo, Norway Department of Data Science, Norwegian University of Life Sciences, Ås, Norway 3 Department for Mathematics and Cybernetics, SINTEF Digital, Oslo, Norway (Dated: April 23, 2026)
Accurate condition monitoring of industrial equipment requires inferring latent degradation parameters from indirect sensor measurements under uncertainty. While traditional Bayesian methods like Markov Chain Monte Carlo (MCMC) provide rigorous uncertainty quantification, their heavy computational bottlenecks render them impractical for real-time process control. To overcome this limitation, we propose an AI-driven framework utilizing Simulation-Based Inference (SBI) powered by amortized neural posterior estimation to diagnose complex failure modes in heat exchangers. By training neural density estimators on a simulated dataset, our approach learns a direct, likelihoodfree mapping from thermal-fluid observations to the full posterior distribution of degradation parameters. We benchmark this framework against an MCMC baseline across various synthetic fouling and leakage scenarios, including challenging low-probability, sparse-event failures. The results show that SBI achieves comparable diagnostic accuracy and reliable uncertainty quantification, while accelerating inference time by a factor of 82× compared to traditional sampling. The amortized nature of the neural network enables near-instantaneous inference, establishing SBI as a highly scalable, real-time alternative for probabilistic fault diagnosis and digital twin realization in complex engineering systems.
Keywords: condition monitoring, Bayesian statistics, heat exchanger, simulation based inference, predictive maintenance, prognostics and health management, fault diagnosis.
I.
INTRODUCTION
The operational integrity and thermal efficiency of complex industrial systems are paramount to the economic performance and safety of modern process plants, driving a paradigm shift toward advanced predictive maintenance (PdM) and condition monitoring [1, 2]. A foundational challenge in industrial asset management is that critical process or health parameters such as efficiency or degradation coefficients, internal component wear, or leak rates often cannot be measured directly with standard instrumentation. Instead, these latent variables can be inferred from observable sensor streams by leveraging mathematical models that simulate the physical behavior of the equipment, often by a costly manual trial-and-error process. While this model-based approach is broadly applicable to a wide range of industrial equipment relying on first-principles or empirical simulations, it is particularly vital for systems where internal states are inaccessible during operation [3, 4]. Within this context, direct Bayesian inference provides a rigorous framework for estimating these latent parameters from sensor data, enabling the uncertainty-aware
∗ [email protected] † [email protected]
calibration of physical simulators and digital twins to accurately reflect the true equipment or process state. Heat exchangers serve as a prototypical example of industrial equipment where health parameters must be indirectly determined. Machine learning and deep learning methods have been extensively studied in the last decade for condition monitoring of heat exchangers failures [5–9], with recent momentum heavily favoring time-series networks like LSTMs, explainable gradient boosting (XGBoost), neural networks [10] and hybrid ML ensembles [11–16]. For these units, essential variables such as the fouling resistance (Rf ), the effective heat transfer area (A), and internal mass-loss fractions are unobservable directly. Instead, these parameters are estimated from observable data, such as inlet and outlet temperatures or fluid mass flow rates, utilizing thermal-fluid simulation tools to bridge the gap between sensor readings and the underlying equipment state [17, 18]. Bayesian inference provides a rigorous mathematical framework for resolving this inverse problem by treating these unobservable parameters as random variables [19]. This allows for the quantification of uncertainty through a posterior probability distribution, which is essential for riskaware decision-making and predicting Remaining Useful Life (RUL) [20–22]. However, the practical application of traditional Bayesian tools, such as Markov Chain Monte Carlo (MCMC) [23, 24] or Sequential Monte Carlo (Particle Filtering) [25], is severely constrained by computational overhead when scaled to complex systems. To ensure convergence, MCMC samplers typically require thousands of iterative evaluations of the underlying physical simulation model for every single inference call. This computational bottleneck renders MCMC impractical for online monitoring scenarios where rapid, high-
2 frequency diagnostics are required for real-time process control. To address these limitations, modern Artificial Intelligence (AI) paradigms, including normalizing flows and neural-networked variational inference [26–28], and specifically Simulation-Based Inference (SBI) [29, 30], have emerged as scalable alternatives to bypass highdimensional Bayesian inverse bottlenecks. The primary advantage of SBI is its ability to perform likelihood-free inference by leveraging the forward process of the simulator itself. By generating a comprehensive dataset pairing parameter inputs and simulation outputs in a one-time offline phase, SBI employs neural density estimators to learn the direct mapping from observed data to the full distribution of the posterior, likelihood or likelihood ratio [29]. Once the neural network is trained, the computational burden is effectively amortized; subsequent inference calls for new sensor measurements require only seconds, allowing the method to scale across complex industrial systems and multiple assets simultaneously [31]. Simultaneously, the neural network can be easily updated with additional simulations if observations drift beyond the range covered by the originally sampled pairs of parameter input and simulation output. While SBI is established as a scientific method [29], the literature lacks examples of industrial applications and performance on such systems. In this study, we present a probabilistic framework for the automated diagnosis of failure modes in a shell-and-tube heat exchanger, focusing specifically on fouling and leakage scenarios. We establish stochastic models for two primary failure mechanisms, progressive fouling and internal leakage, governed by latent parameters to be determined in a Bayesian fashion. A systematic comparison is conducted between traditional MCMC sampling and amortized SBI, specifically utilizing sequential Neural Posterior Estimation (NPE, [32]) to identify the onset and evolution of these failures. Our results demonstrate that the SBI approach is orders of magnitude faster than MCMC, providing near-instantaneous posterior characterization without sacrificing diagnostic depth. By highlighting this significant gain in computational efficiency, this work establishes a scalable workflow for deploying high-fidelity Bayesian condition monitoring in real-time thermal-fluid applications. Furthermore, we remark how this approach is transferrable to the condition monitoring of other multi-parameter industrial processes and equipment making it exceptionally well-suited for legacy systems or "black-box" simulators where the underlying governing equations are inaccessible [33]. Alternative methods include learning fast approximations for the simulator [34] to speed up classical MCMC implementation, replacing legacy simulators with their differentiable counterparts [35], or leveraging Physics-Informed Neural Networks (PINNs) to embed governing physical constraints directly into the learning architecture [36, 37].
Shell-and-Tube Heat Exchanger ṁhot , cp,hot Thot,in
Thot,out UA
Q Tcold,in
Tcold,out
ṁcold , cp,cold
FIG. 1: Schematic of a counterflow heat exchanger with key variables. The hot fluid (red) enters at Thot,in with mass flow rate ṁhot and specific heat cp,hot , and exits at Thot,out . The cold fluid (blue) enters at Tcold,in with ṁcold and cp,cold , and exits at Tcold,out . Heat, Q, is transferred across the exchanger wall with overall conductance U A.
II.
METHODS
As physical basis for our diagnostic model, we study a deterministic heat exchanger model (Sec. II A) with stochastic failures (Sec. II B).
A.
Deterministic Heat Exchanger Model
For the deterministic model, we consider a steady-state counterflow heat exchanger as illustrated in Fig. 1. The heat capacity rates for the hot and cold fluid streams are defined by their respective mass flow rates ṁ and specific heat capacities cp : Chot = ṁhot cp,hot ,
Ccold = ṁcold cp,cold .
(1)
For a counterflow heat exchanger, the energy (heat) balances are given by Q = ṁhot cp,hot (Thot,in − Thot,out ) , Q = ṁcold cp,cold (Tcold,out − Tcold,in ) ,
(2) (3)
as well as the heat transfer equation, Q = U A · ∆TLM ,
(4)
where ∆TLM denotes the log-mean temperature difference (LMTD) between the two streams, defined as ∆TLM =
(Thot,in − Tcold,out ) − (Thot,out − Tcold,in ) . (5) T −Tcold,out ln Thot,in hot,out −Tcold,in
While the heat transfer rate Q is traditionally solved using the LMTD method, the nonlinear nature of the LMTD requires iterative root-finding for unknown outlet temperatures. To facilitate efficient Bayesian inference, we utilize the effectiveness-NTU (ϵ-NTU) method [38], which provides an explicit analytical solution. The heat transfer is expressed as: Q = ϵCmin (Thot,in − Tcold,in ) ,
(6)
3 where Cmin = min(Chot , Ccold ). The heat exchanger effectiveness, ϵ, represents the ratio of actual heat transfer to the maximum thermodynamic limit. Here, ϵ is determined by the Number of Transfer Units (NTU) and the capacity rate ratio r: Cmin , Cmax
(7)
1 − exp[−NTU(1 − r)] . 1 − r exp[−NTU(1 − r)]
(8)
NTU =
ϵ=
UA , Cmin
r=
The product of the overall heat transfer coefficient U and the exchange area A is denoted as U A, which characterizes the heat transfer capability of the equipment. The outlet temperatures are then directly recovered from the energy balance: Q Chot Q . Tcold,out = Tcold,in + Ccold Thot,out = Thot,in −
B.
Stochastic Modeling of Failure Mechanisms
For each of the failure mechanisms object of this study, degradation is modeled by introducing a stochastic time-dependency into the parameters of the deterministic model in Sec. II A, aligning with recent reliability frameworks for stochastic degrading devices [39]. These stochastic formulations are selected not because they necessarily provide a more accurate physical representation of degradation than existing deterministic models, but rather as tunable, multi-parameter frameworks for failure evolution. By adjusting the event frequency and the severity scales of the failure mechanisms, we can generate a wide spectrum of failure scenarios. This flexibility allows for the systematic creation of datasets that vary in their level of difficulty for anomaly detection and parameter identification. Consequently, the model serves as a rigorous testbed for evaluating the robustness of the Bayesian inference framework across scenarios ranging from subtle, continuous degradation to sporadic, highimpact failure events. In general, we assume failures initiate at an unknown changepoint τ . To maintain differentiability for gradient-based inference, we utilize a logistic sigmoid transition S(t) to model the induction period: S(t) =
1 , 1 + exp(−k(t − τ ))
(9)
where k is the transition sharpness. Two primary failure mechanisms are considered: 1. Tube Fouling. In this work, grounded in recent analyses of physical deposition mechanisms [40], fouling is modeled as a reduction in the overall conductance U A via a non-negative, dimensionless fouling factor R(t): U A(t) =
U Aclean . 1 + R(t)
(10)
FIG. 2: Fouling and Leakage evolution in time with the effect of the failure parameters (λ, βf , βl ) over the trajectories at a fixed failure onset time τ .
To capture the sporadic, burst-like nature of industrial scaling, R(t) is modelled as a discretized and relaxed Compound Poisson Process [41]. The total fouling at time t is the accumulation of stochastic increments: R(t) =
t X i=1
∆Ri ,
∆Ri = S(i) · Ii · Ji · gf ,
(11)
where gf ∈ {0, 1} is a binary fouling indicator determined by the sampled failure mode z. Specifically, gf = 1 when the selected mode includes fouling and gf = 0 otherwise. The term S(i) represents a steep sigmoid function centered at the random failure time τ , acting as a temporal gate that activates the process for i ≥ τ . The
4 stochastic occurrence of a jump is governed by the relaxed gating variable Ii . At each time step, a jump probability is calculated from the Poisson arrival rate λ as Pjump = 1 − exp(−λ). The operator is defined by passing a latent uniform sample ui ∼ Uniform(0, 1) through a steep sigmoid function: Ii =
1 . 1 + exp(−k(Pjump − ui ))
(12)
The magnitude of the potential jump, Ji , is drawn from an exponential distribution. To ensure a mean jump size of βf , the distribution is parameterized by the rate 1/βf : Ji ∼ Exp
1 βf
growth process. Following the changepoint τ , the cumulative degradation is driven by exponential increments: " #! t X L(t) = Lmax 1 − exp − S(i) · gl · ∆Li , (17) i=1
where Lmax = 0.95 represents the physical saturation of the leak and gl ∈ {0, 1} is a binary leakage indicator determined by the sampled failure mode z. The increments ∆Li are sampled from an exponential distribution ∆Li ∼ exp (1/βl ), where the scale parameter βl determines the growth rate of the failure.
,
where E[Ji ] = βf .
(13)
Overall, the fouling process depends on the arrival rate λ and the strength parameter βf . As λ increases, the frequency of stochastic events grows, while a higher βf increases the average magnitude of each individual jump. The expected growth of the fouling factor can be analyzed via the linearity of expectation. For the postinduction phase (i > τ ), where the temporal gate S(i) ≈ 1, the expected increment is the product of the expected jump probability and the expected jump magnitude: E[∆Ri ] = E[Ii ] · E[Ji ] · gf = (1 − e−λ ) · βf · gf .
(14)
For t > τ , the total accumulated fouling is the sum of these independent increments over the active duration of the failure. Thus, by Wald identity, the expectation of R(t) is linear in time: E[R(t)] ≈ (1 − e−λ ) βf gf (t − τ ) . (15) compatibly with the well known Ebert-Panchal model [42]. This derivation shows that the process is linear in time relative to the random failure point τ . This formulation is consistent with the first-order behavior of deterministic fouling laws while retaining the random-event structure necessary to represent the intermittent and unpredictable nature of industrial fouling. 2. Leakage with Fluid Loss. Internal leakage is modeled as a diversion of fluid from the primary hotstream flow path, representing a shell-side bypass or a loss of integrity in the tube-to-header seals. We define a time-dependent leak fraction L(t) ∈ [0, 1) that reduces the effective mass flow rate participating in the heat exchange: ṁhot,out (t) = ṁhot,in (1 − L(t)) .
(16)
This reduction in the hot-side thermal mass decreases the heat capacity rate Chot , altering the capacity rate ratio r. This results in a simultaneous reduction in outlet mass flow and a decrease in Thot,out . In our probabilistic framework, the evolution of the leak fraction L(t) is modeled as a continuous stochastic
C.
Bayesian Inference and Probabilistic Modeling
Bayesian inference offers a rigorous mathematical framework for quantifying uncertainty in model parameters θ given observed data y. In the case of industrial simulations, it solves the inverse problem of determining the distribution p(θ | y) representing the probability that a set of parameters θ produce the observed sensor data y. The cornerstone of the Bayesian approach is the posterior distribution, p(θ | y) =
p(y | θ) p(θ) , p(y)
(18)
where p(θ) R is the prior, p(y | θ) is the likelihood, and p(y) = p(y | θ) p(θ) dθ is the model evidence. In the context of heat exchanger modeling, θ may include parameters such as the effective heat transfer coefficient, fouling factor, mass flow rates, and failure mode indicators. The likelihood is typically constructed based on a well-defined physical model such as the one in Sec. II B, incorporating measurement noise and potential process disturbances: yi = f (θ, xi ) + ϵi ,
ϵi ∼ N (0, σ 2 ),
(19)
where f (·) denotes the deterministic model mapping, xi are covariates or known inputs, and ϵi are independent, normally distributed errors. Posterior inference enables both point and interval estimation, as well as predictive uncertainty. The predictive distribution for new data y∗ given new input x∗ is Z p(y∗ | y) = p(y∗ | θ, x∗ ) p(θ | y) dθ, (20) fully accounting for parameter uncertainty. In most realistic settings, this integral cannot be evaluated analytically. Instead, posterior inference must rely on numerical methods such as Markov Chain Monte Carlo (MCMC, [43]) or variational inference [44]. MCMC constructs a Markov chain whose stationary distribution is the posterior p(θ | y), which enables asymptotically exact sampling-based inference for both continuous and discrete latent variables. In this work, we use Hamiltonian
5 Monte Carlo through the No-U-Turn Sampler (NUTS) for continuous parameters [45]. For the discrete failuremode formulation we combine NUTS with a Gibbs update over the categorical mode variable. MCMC methods provide a strong posterior benchmark, but becomes computationally expensive as it requires a new sampling procedure for every observation. D.
difference Tcold,out − Tcold,in , the hot-side mass-flow loss ṁhot,in − ṁhot,out , and the outlet temperatures Thot,out and Tcold,out . For each of these five signals, the mean, standard deviation, early-to-late change, dynamic range, and a linear trend measure are computed to yield a 25dimensional summary statistic vector. These summaries are designed to capture the temporal signatures most informative for both failure-mode identification and estimation of the associated degradation parameters.
Simulation-Based Inference (SBI)
Simulation-Based Inference (SBI), also known as likelihood-free inference, is a family of techniques for Bayesian inference in models for which the likelihood function p(y | θ) is analytically intractable or unavailable, but simulation of synthetic data ysim ∼ p(y | θ) is feasible [29]. This is particularly relevant for nonlinear physical systems with complex failure dynamics such as the heat exchanger. One classical SBI approach is Approximate Bayesian Computation (ABC) [46], which defines an approximate posterior as Z pϵ (θ | y) ∝ p(θ) I (d(y, ysim ) < ϵ) p(ysim | θ)dysim , (21) where d(·, ·) is a distance metric and ϵ is a user-defined tolerance. More recent methods leverage neural density estimators to approximate either the posterior p(θ | y), the likelihood p(y | θ), or the likelihood-to-evidence ratio, typically via simulated data sets [32]: p̂ϕ (θ | y) ≈ p(θ | y),
(22)
where p̂ϕ denotes the neural network approximation parameterized by ϕ. These methods permit inference even in complex models for which classical Bayesian updating is not tractable, provided that forward simulation is possible. For recent advances and theoretical guarantees, see [29]. We apply an amortized sequential neural posterior estimation (NPE) [47] as implemented in the sbi toolkit software package [48]. Details on the SBI approach can be found in Figure 3 and Section III C. E.
Summary statistics
In Bayesian inference such as MCMC and SBI, the data and models are compared on so-called summary statistics. For high-dimensional data, lower dimensional representations that preserve the relevant information about the parameters of interest enable comparison to simulated data. In particular, they are crucial when the system is complex and the likelihood function is intractable. The lower dimensional representation can be designed from knowledge or learned through an embedding neural network. In this study, we construct a set of summary statistics tailored to the heat exchanger measurements. Features are derived from the hot-side temperature difference Thot,in −Thot,out , the cold-side temperature
F.
Scoring and distance measures
To compare performance of SBI against MCMC, we use metrics that assess both the similarity between inferred posteriors and their quality with respect to the ground-truth parameters. The Wasserstein distance between the SBI and MCMC posteriors quantifies the discrepancy between the amortized SBI posterior and the MCMC reference posterior. The Continuous Ranked Probability Score (CRPS) and credible-interval coverage are used to evaluate the quality of each inferred posterior. These metrics are evaluated separately for each continuous parameter of interest. CRPS is a generalization of the Mean Absolute Error (MAE) for probabilistic predictions [49]. It is calculated by comparing the empirical cumulative distribution function (ECDF) of posterior samples with the step function CDF at the true value. Lower CRPS indicates a posterior that is sharper and better aligned with the ground-truth. The Wasserstein distance is instead a similarity metric between two probability distributions. The convergence rate of empirical Wasserstein estimators typically behaves as n−1/d , where n is the number of samples and d is the dimension, implying that standard empirical Wasserstein distance becomes a poor measure of similarity for high-dimensional data. It is also sensitive to extreme outliers.
III.
EXPERIMENTS
The aim is to evaluate the robustness of the Bayesian diagnostic framework by identifying the failure mode and determine the failure parameters: induction time τ (Eqn. 9), fouling jump scale βf (Eqn. 13), leakage growth βl (Eqn. 17) and fouling intensity λ. In the model setup, failure mode is represented by a discrete categorical variable z ∈ {none, fouling, leakage, both}, rather than a continuous simplex Dirichlet prior over mode probabilities. This avoids the confounding present in the Dirichlet formulation, where the mode weights also scale degradation magnitude. As a result, the categorical formulation leads to clearer failure-mode identification.
6
n-parameters draws from priors θ1 = {λ1 , βf 1 , τ1 , etc.}, . . . , θn = {λn , βf n , τn , etc.}
λ
n simulation runs
n summary statistics
τ
βf
HX Simulator f (λ, τ, βf , etc.)
{Thot,in , Tcold,in , Thot,out , Tcout , mhot,in , mhot,out , ∆Thot , ∆Tcold , mloss }.
Training
λ
Thot,in
βf
τ
{Thot,in , Tcold,in , Thot,out ,
mhot,in
Tcout , mhot,in , mhot,out , ∆Thot , ∆Tcold , mloss }.
observations y
summary statistics
SNPE (NSF)
posteriors’ marginals under observations p̂ϕ (θ | y)
FIG. 3: The applied SBI scheme. Blue: training data (priors) and process, orange: input data, green: inferred output posteriors. In this work n=50,000 simulations.
Scenario Failure τ βf βl λ Weak Fouling Fouling 18 0.005 5.0 Batch SD Fouling 18 0.030 0.5 Boiler FW Fouling 18 0.050 3.0 Mild Leak Leakage 18 - 0.0005 Severe Leak Leakage 18 - 0.0010 -
TABLE I: Failure parameters for synthetic failure scenarios. τ denotes the failure changepoint (induction time); βf and βl represent the scale parameters for fouling jumps and leakage growth, respectively; and λ defines the fouling arrival rate (intensity).
A.
Test scenarios
For evaluation purposes, we define five distinct failure scenarios representing a spectrum of industrial operational conditions. These scenarios are generated using the stochastic model described in Sec. II, with specific parameters detailed in Table I. The scenarios are designed to test the model’s ability to distinguish between quasi-continuous degradation and sporadic, high-impact events: Weak Fouling (Scenario 1): This represents a wellmaintained system where fouling is slow and progresses in a quasi-continuous manner. The high arrival rate (λ = 5.0) combined with a low jump scale (βf = 0.005) simulates frequent, minor deposition events that approximate a linear degradation
trend. This is typical of systems with high fluid velocity or effective surface conditioning, presenting a challenge for early-stage anomaly detection due to the subtle signal-to-noise ratio of the degradation increments. Batch Process Shutdown (Scenario 2): This scenario simulates fouling that occurs primarily during stagnant periods or process shutdowns (Batch SD). The low arrival rate (λ = 0.5) and high jump scale (βf = 0.03) result in rare but severe “shocks” to the heat transfer efficiency. The diagnosis of such low-probability, time-dependent rare events remains a recognized challenge in industrial fault detection [50], particularly under few-shot observability constraints where Bayesian uncertainty calibration is vital [51]. Boiler Feedwater System (Scenario 3): Representing aggressive scaling in systems with high mineral content (Boiler FW), this scenario utilizes both high frequency and high severity scales. The combination of λ = 3.0 and βf = 0.05 creates a rapid loss of heat transfer capability. This scenario tests the model’s ability to track high-gradient failure progression and its reliability in predicting the remaining useful life of the equipment. The leakage scenarios represent progressive structural failures, such as tube pitting or gasket degradation. In both cases, the failure is characterized by a gradual loss
7
0.020
Mild Leak (Scenario 4): A slow-onset failure where the leak fraction grows moderately (βl = 0.0005). This scenario tests the sensitivity of the Bayesian framework to distinguish small mass-flow discrepancies from measurement noise.
τ Prob. Density
of mass flow and a corresponding thermal fingerprint in the outlet temperatures.
B.
Priors
For the operational scenario detection, a categorical prior is required. As heat exchangers operate normally more often than they experience failures, we assign higher prior probability to the no-failure mode: p(failure mode) = [0.4; 0.2; 0.2; 0.2] (none, fouling, leakage, both). We note that the scenario where both failure mechanisms are active, is a very unlikely situation: albeit we still consider it in defining the categorical prior, we do not generate data with both leakage and fouling and consider it only as a confounding scenario. For the changepoint (τ ), we assume a uniform prior over the interval τ ∈ [1, T − 1], excluding boundary cases in which the failure is either already present (τ ≤ 0) or does not occur within the observation window (τ ≥ T ). The positive degradation parameters are assigned log-normal
15
30
45
τ
60
90
75
βf Prob. Density
30 20
βf = 0.005 βf = 0.03 βf = 0.05
10
0 0.000 0.025 0.050 0.075 0.100 0.125 0.150
βf
βl Prob. Density
2500 2000 1500 1000
βl = 0.0005 βl = 0.001
500 00.000
λ Prob. Density
For each scenario, an ensemble of 500 synthetic datasets was generated to ensure statistical significance in the evaluation of the inference methods. For scenario identification, each discrete failure mode is handled by predicting logits, which are transformed via a softmax function to obtain mode probabilities. At inference time, we draw posterior samples, assign each draw the label with the highest softmax probability, and thus obtain discrete samples of the latent realized trajectory, z, analogous to the MCMC output. A scenario is considered correctly identified if the most frequent posterior label matches the true simulated label. In addition to the failure mode status label, we also infer the failure parameters {τ, λ, βf , βl } directly from the data and independently of the failure category.
0.005
40
Finally, we also include a baseline no-failure case: No Failure (Scenario 6): This scenario represents normal heat exchanger operation without fouling or leakage. It serves as a baseline for assessing whether the inference framework can correctly identify the absence of failure and avoid falsepositive diagnoses.
0.010
0.000 0
Severe Leak (Scenario 5): A rapid loss of system integrity (βl = 0.0010). This scenario creates a significant and fast-evolving divergence between inlet and outlet mass flow, requiring the model to maintain stability while the heat capacity rate ratio r shifts rapidly.
τ = 18
0.015
0.001
βl
0.002
0.003
0.4 0.3 0.2
λ=5 λ = 0.5 λ=3
0.1 0.0 0.0
2.5
5.0
7.5
λ
10.0
12.5
15.0
FIG. 4: Prior probability densities for the changepoint time τ , fouling strength βf , leak rate βl , and fouling-event arrival rate λ used in the heat-exchanger model. Vertical dashed lines mark true parameter values for the study scenarios reported in Table I.
priors: βf ∼ LogNormal(log 0.015, 1.0), βl ∼ LogNormal(log 0.0004, 0.4), λ ∼ LogNormal(log 2.0, 0.5). These priors enforce positivity while remaining broad enough to cover the range of fouling and leakage behaviors considered in the study as shown in Fig. 4.
0.09
40
0.06
20
0.03
βf
τ 00 1.5e-03
20
40
60 0
0.03
0 0.09 3
0.06
2 7.5e-04 1
βl
λ
00 0.00075 0.0015 0 1 2 30 MCMC posterior median MCMC posterior median
SBI posterior median
Process data were generated by modelling the heat exchanger in a two-stage procedure. In the first stage, we specify a parameter vector x that characterizes the stationary properties of the system, such as the failure mode, failure strength, time constant τ , and related quantities. Given these parameters, we generate a trajectory z from a stochastic process. This trajectory represents the latent (unobserved) timeseries for fouling and leakage dynamics. In the second stage, the latent trajectory z is used as input to a simulator of the heat exchanger to generate the observable time series y, such as temperatures, flows, and other measured signals. In its most vanilla form, simulation-based inference (SBI) is trained by providing it with query access to the simulator, yielding pairs (x, y). At inference time, SBI produces an approximation to the posterior distribution p(x | y). Since the latent trajectory z is never observed and the inference procedure is not informed about the internal latent structure of the simulator SBI has no opportunity to directly infer the realized trajectory z. This is in contrast to MCMC methods, which have access to the full probabilistic model. A practical SBI approach is therefore to infer x given y, and then sample trajectories z from the simulator conditioned on the inferred parameters. However, if the latent variable is stochastic this procedure can never recover the actual realized trajectory, only samples from its conditional distribution as implied by the simulator. Alternatively, one could modify the inference setup and train SBI on pairs (x, z) mapped to observations y. At inference time, conditioning on y would then yield a joint posterior over both x and the trajectory z. In principle, this allows direct estimation of latent trajectories rather than only their distributional family. In practice, however, this substantially increases the complexity of the inference problem, as SBI must explore a very highdimensional latent space corresponding to the full time series z. We benchmark the amortized SBI approach against a MCMC baseline. For MCMC, the stochastic degradation model is implemented in NumPyro [52, 53], and the No-U-Turn Sampler [45] draws joint samples of the discrete failure mode z and the continuous degradation parameters. This requires extensive simulator evaluations for every new observation. We found sufficient a setup made of 4 chains, each of 2, 000 warm up steps and 3, 000 samples, for a total of 20, 000 MCMC evaluations per inference task. For SBI, we employ the scheme depicted in Figure 3 and in Ref.[32]. The neural density estimator is trained offline on a conservative 50, 000 forward simulations, learning a direct mapping from the 25-dimensional summary statistics of the observed parameters to an approximate posterior distribution for the failure mode and failure parameters. For the neural density estimator within the Sequential Neural Posterior Estimation (SNPE) framework, we employed a Neural Spline Flow (NSF) architecture [47, 54]. Unlike simpler affine autore-
60
SBI posterior median
Experimental setup
SBI posterior median
C.
SBI posterior median
8
FIG. 5: Scatterplot comparing posterior median parameter estimates inferred by MCMC and SBI across the 2,500 failure-case observations (500 realizations for each of Scenarios 1–5). Colors denote scenarios: blue (Weak Fouling), orange (Batch Shutdown), green (Boiler Feedwater), yellow (Mild Leakage), and lilac (Severe Leakage). The dashed identity line indicates perfect agreement between the two inference methods.
gressive flows, NSFs utilize monotonic rational-quadratic splines, providing the high representational capacity required to map the 25-dimensional summary statistics to the complex, potentially multimodal posterior distributions of the heat exchanger’s latent degradation parameters. The normalizing flow was configured with a standard multivariate normal base distribution and comprised 5 sequential transforms. The underlying conditioner network was implemented as a multi-layer perceptron (MLP) featuring 2 hidden layers with 50 hidden units each, utilizing Rectified Linear Unit (ReLU) activations. The neural network was trained using the Adam optimizer with an initial learning rate of 5 × 10−4 and a batch size of 256. To prevent overfitting during the offline amortization phase, training was subject to an early stopping criterion, halting if the validation logprobability failed to improve over 20 consecutive epochs. All inference architectures were built and trained utilizing the PyTorch-based sbi toolkit [48]. To evaluate robustness, we use the six benchmark scenarios described in Sec. II. For each scenario, we evaluate 500 independent observation records generated with the same ground-truth parameters but differing measurement noise realizations.
9
FIG. 6: The marginalized distributions of the posterior medians inferred via MCMC (orange) and SBI (blue) across 500 noisy observations per scenario. The red dashed lines indicate the ground-truth parameter values. Each row represents a single continuous parameter, while each column corresponds to a specific failure scenario.
IV.
RESULTS
This section compares SBI and MCMC over the six benchmark scenarios using 500 noisy realizations per scenario. We first assess failure-mode identification, then compare continuous-parameter inference quality and posterior shape agreement, and finally discuss Scenario 2 as a representative stress test under weak observability. Table II shows that both MCMC and SBI identify fouling and leakage scenarios with near-perfect reliability, assigning most posterior mass to the correct failure mode across the 500 datasets per scenario. The healthy state (Scenario 6) is slightly more challenging, as transient sensor noise can emulate weak degradation and occasionally shift posterior mass toward a fault mode. Even in this regime, SBI remains statistically equivalent to the MCMC reference for failure-mode classification. Figure 5 provides a pointwise comparison of posterior medians obtained with MCMC (horizontal axis) and SBI (vertical axis) for all failure realizations. Each marker corresponds to one inference task, so the concentration of points around the identity line indicates sample-by-
Scenario
Failure MCMC SBI mode accuracy accuracy 1: Weak Fouling Fouling 100% 100% 2: Batch SD Fouling 100% 100% 3: Boiler FW Fouling 100% 100% 4: Mild Leak Leakage 99.8% 100% 5: Severe Leak Leakage 99.6% 100% 6: No Failure None 98.2% 98.6%
TABLE II: Categorical failure-mode identification accuracy for MCMC and SBI across the six scenarios. Examples of time evolution curves for the scenarios are shown in Fig 2.
sample agreement between the amortized and samplingbased estimators. Across scenarios, the cloud is tightly aligned with the diagonal for mode-discriminative parameters, confirming that SBI preserves the same central diagnostic conclusions as MCMC. The largest spread is observed for the changepoint τ , particularly in sparse-event fouling regimes, where delayed and weakly informative
τ
1.0
1.00
βf
0.50
0.5
τ
βf
20
0.02
10
0.01
0 k k ak Ftch SDler FWld Lea re Lea e W Ba Boi Mi eve S
0.00 ak F tch SD ler FW We Ba Boi
CRPS
0.75
0.25
0.0 k k ak F SD FW ea ea We BatchBoiler Mild Levere L S
0.00 SD FW ak F We Batch Boiler 1.5
βl
λ
Wd (MCMC-SBI)
30
0.2
1.0 0.1
0.5
0.0
d
Mil
k Lea
re
e Sev
0.0 SD FW ak F We Batch Boiler
k Lea
FIG. 7: Normalized, 1-dimensional Wasserstein distance between MCMC and amortized SBI posterior samples for the same failure parameter, scaled by the absolute value of the true parameter. Color coding follows the one in Figure 5
λ
βl
0.0003 CRPS
Wd (MCMC-SBI)
10
2
0.0002 1
0.0001 0.0000
Mi
eak ld L
k
Sev
e
ea re L
0 ak F tch SD ler FW We Ba Boi
FIG. 8: Continuous Ranked Probability Score (CRPS) distributions comparing the sharpness and accuracy of MCMC and SBI posterior predictions for each parameter.
.
transients reduce practical identifiability. An additional pattern in Fig. 5 is the comparatively narrow SBI posterior for λ relative to MCMC. This behavior is consistent with an amortization-induced shrinkage effect in weakly identifiable settings: when the observation carries limited information about jumparrival intensity, SNPE tends to regularize toward dominant simulator-supported regions learned during training, yielding sharper posteriors. In contrast, MCMC explores a flatter likelihood landscape more explicitly and therefore retains wider uncertainty for λ [55, 56]. The key point is that this difference is primarily in posterior width, not in failure-mode decision; both methods still produce consistent classification outcomes. To systematically compare MCMC and SBI, we analyze the inference results over the 500 independent noisy realizations per scenario. Fig. 6 presents the marginalized distributions of posterior medians for each parameter and scenario. These distributions represent variation of point estimates across datasets, not the uncertainty width of individual posteriors. For the induction time τ , medians are centered close to the ground truth in all scenarios, with MCMC yielding slightly narrower between-dataset spread. For βf , βl , and λ, SBI and MCMC remain broadly consistent in central tendency, while both methods show reduced accuracy for λ in fouling scenarios. This behavior reflects identifiability lim-
its of the stochastic degradation model rather than a method-specific failure. Consistent with Fig. 5, SBI posteriors for λ are often narrower than MCMC in weakly identifiable regimes, whereas MCMC preserves broader uncertainty by exploring near-flat likelihood regions. Importantly, despite these limits in continuous-parameter recovery, failure-mode classification remains correct in virtually all 2,500 failure cases. Similar qualitative conclusions are obtained when using posterior maximumlikelihood point summaries. To assess similarity of full posterior shapes, we evaluate the 1D Wasserstein distance between MCMC and SBI posteriors for each parameter (Fig. 7). The distributions are generally concentrated at low values, indicating that the amortized SBI posterior remains close to the MCMC reference over most inference tasks. Larger distances appear primarily in weakly identifiable regimes, where multiple parameter combinations explain the observations nearly equally well. This pattern is consistent with the scenario-level analysis above: discrepancies are driven more by structural identifiability limits than by systematic bias of the SBI estimator. Finally, Fig. 8 reports the Continuous Ranked Probability Score (CRPS), a proper scoring rule that jointly evaluates probabilistic sharpness and statistical accuracy. The CRPS distributions for SBI and MCMC are of similar magnitude across parameters, indicating that SBI preserves most of the probabilistic predictive quality of the MCMC baseline. For the induction time τ and degra-
11
FIG. 9: Posterior predictive checks for a Batch Process Shutdown (Scenario 2). Top panel: generated sensor data for Thot,out (full line), average over 500 realizations (dashed line) and 90% confidence interval. Vertical dashed line show failure onset time (τ ). Mid panel: Fouling resistance computed for the observed realization (black) together with an example trajectory from parameters sampled from the posteriors inferred by MCMC (orange) and SBI (blue). The average (dashed lines) and 95% confidence intervals (shaded) are shown for trajectories generated from 500 posterior samples. Bottom panel: The inferred posterior probabilities per process parameter from MCMC (orange) and SBI (blue). The ground truth is shown as red lines in the diagonal density plots and crosses in the covariance plots.
dation parameters βf , βl , both metrics demonstrate acceptable predictive accuracy levels in most scenarios. The exception is the determination of λ in the Batch Process Shutdown (Scenario 2), where the posterior location is consistently overestimated due to the extreme sparsity of the fouling jumps. In this regime, the un-
derlying degradation is driven by a very low event arrival rate (λ = 0.5) combined with a large jump scale (βf = 0.03). Because these stochastic shocks are exceedingly rare, a given finite observation window may capture very few—or even zero—fouling events. Consequently, the observed thermal-fluid signals carry minimal information regarding the true arrival rate, resulting in a nearly flat likelihood surface for λ. In the absence of informative data, the posterior distribution naturally shrinks toward the prior. Since the assigned log-normal prior for λ is centered at a median of 2.0 —significantly higher than the true scenario value—both MCMC and SBI predictably overestimate this parameter. This behavior highlights a fundamental limit of structural identifiability in sparseevent regimes, rather than a systematic failure of the inference algorithms. The challenge is compounded by a fundamental identifiability trade-off between λ and βf . Infrequent large jumps and frequent small jumps can produce near-identical cumulative fouling trajectories, making it difficult to distinguish the true regime from the dominant prior. MCMC is particularly susceptible because it must simultaneously infer λ alongside 100 hidden per-timestep jump indicators, entangling the global arrival rate with local decisions about whether each individual timestep carried a jump. SBI avoids this by absorbing the individual timestep variables into the simulation process and instead relies on summary statistics. This compresses what little signal exists about the rate of fouling accumulation into a lower-dimensional representation. Even so, neither method can fully overcome the prior dominance when the observation window contains so few events. This highlights a fundamental limit of structural identifiability in sparse-event regimes, rather than a systematic failure of the inference algorithms. Detailed model calibration analyses are provided in the Supplementary Material. Scenario 2 (Batch Process Shutdown) provides therefore a representative stress test for probabilistic diagnosis under weak observability. In this regime, fouling evolves through sparse but potentially large stochastic jumps (low event rate, λ = 0.5, and relatively large jump scale, βf = 0.03), so the thermal response can resemble measurement noise over extended time windows. This setting is therefore informative for assessing whether the inference framework can separate persistent degradation dynamics from transient fluctuations. Both MCMC and SBI assign dominant posterior probability to the fouling mode (Table II) and recover a consistent posterior for the changepoint τ . As shown in Fig. 9, marginal posteriors are also in close agreement across methods, indicating that amortized SBI preserves the main uncertainty structure of the MCMC reference posterior. Although both approaches show limited identifiability for λ and, to a lesser extent, βf in this sparse-event regime, posterior predictive trajectories remain physically plausible and reproduce the observed trend in fouling resistance. From a reliability perspective, this is the key requirement for downstream condition monitoring tasks, such as
predicting the Remaining Useful Life of the equipment. Even when individual parameters are weakly identifiable, the inferred posterior ensemble still supports robust fault recognition and uncertainty-aware forecasting of degradation progression. Accurately recovering these parameters and their associated uncertainties is a critical first step for downstream
A.
Resources consumption
The primary motivation for adopting SBI in industrial asset management is its exceptional computational scalability at deployment time. As summarized in Table III, MCMC must restart its iterative sampling process from scratch for every new diagnostic request (inference), so its total simulation cost grows linearly with the number of inference tasks. In contrast, SBI training is a one-time offline phase whose upfront cost is rapidly amortized over subsequent diagnoses. We therefore pushed both approaches to a compact setup where the simulator calls are as few as possible while retaining acceptable accuracy on the parameters posteriors, in order to test the computational cost of inference in the two cases. MCMC was run on 4 chains with 150 burn-in and 75 samples (900 transitions/call) while SBI was trained on 5,000 simulations. As illustrated in Fig. 10, the two methods reach cost parity after fewer than six inference calls, beyond which SBI is 82 times faster per diagnosis while maintaining identical classification accuracy and τ convergence. In this study, the simulator is a lightweight JAX-JITcompiled effectiveness-NTU model of a single heat exchanger, chosen as a toy model for testing. Each simulator call is estimated to require 0.03 s of computing on a 4-core CPU Apple M4 Pro with JAX-JIT pre-warmed. Therefore, the absolute time saved per call is modest. However, the 82× acceleration is a property of the inference algorithms rather than of the particular simulator, and it scales with simulator cost. For networks of interacting heat exchangers or plant-wide process simulations where a single forward evaluation can take minutes to hours, the same factor translates to savings of tens of minutes to many hours per diagnosis. Across a modern process plant where condition monitoring must be performed simultaneously on dozens of assets at high frequency, the near instantaneous posterior evaluation offered by SBI provides a pathway to real-time, risk-aware decision-making that is entirely intractable with classical Method Sim. calls Acc. Infer. time Train cost MCMC 900 / call 100 % 2.4 s / call — SBI 5,000 (once) 100 % 0.029 s / call 19 s
TABLE III: Computational cost of inference at 100% accuracy for MCMC and SBI on an Apple M4 Pro laptop.
Total Simulation Runs
12
10000 7500 5000 2500 00
2
4
6
8
Inference Calls
10
12
FIG. 10: Resource comparison at the minimal configuration achieving both ≥95% classification accuracy on Scenario 2 (fouling, τ = 18). MCMC used 4 chains with 150 burn-in + 75 samples (900 transitions/call); SBI used 5,000 training simulations. SBI is 82× faster per call after a break-even of ∼6 inference calls.
Bayesian sampling.
V.
DISCUSSION
This study demonstrates that simulation-based inference (SBI) provides a practical and scalable alternative to classical Bayesian sampling for condition monitoring of heat exchangers. Across all evaluated scenarios, SBI achieves diagnostic performance comparable to MCMC in both failure-mode classification and parameter estimation, while reducing inference time by several orders of magnitude. The close agreement in posterior summaries and distributional metrics indicates that SBI with amortized neural posterior estimation can efficiently approximate high-fidelity MCMC baselines, even in nonlinear thermo-fluid systems with stochastic degradation dynamics. A key implication is the shift in computational burden enabled by SBI. By front-loading simulation cost into an offline training phase, the framework enables nearinstantaneous inference during deployment. This is particularly relevant for industrial settings where continuous monitoring must be performed across multiple assets under strict latency constraints. The results suggest that, beyond a small number of inference queries, SBI becomes more computationally efficient than MCMC, making it suitable for real-time diagnostics and large-scale deployment. From a reliability engineering perspective, the framework supports probabilistic fault diagnosis by providing full posterior distributions over failure modes and degradation parameters. This enables uncertainty-aware decision-making, which is critical for risk assessment, maintenance planning, and integrating uncertainty information directly into robust RUL predictions [57]. Notably, the results indicate that failure onset and mode
13 identification is highly robust, even in regimes where individual parameters are weakly identifiable. This distinction highlights that reliable anomaly detection may be achievable even when precise quantification of degradation dynamics remains challenging. Several limitations should be considered. First, parameter identifiability depends strongly on the observability of the underlying process; for example, the fouling intensity and changepoint parameters are difficult to estimate in scenarios with sparse or inactive degradation. These challenges arise from both measurement noise and structural properties of the stochastic model. Second, the study relies on synthetic data generated from a simplified degradation model, which does not fully capture the complexity of real industrial systems. In particular, assumptions such as stationary noise and simplified fouling dynamics may limit direct transferability. Additionally, the use of engineered summary statistics, while computationally efficient, may discard information present in the full time series. Despite these limitations, the approach is broadly applicable due to its model-agnostic nature. SBI does not require an explicit likelihood function and can be applied to black-box simulators, making it suitable for legacy systems where governing equations are inaccessible. Compared to purely data-driven methods, the framework retains physical interpretability through latent parameters while providing calibrated uncertainty estimates. Future work should focus on validation with real operational data, where additional sources of variability such as sensor drift and unmodeled disturbances are present. However, for real-life industrial systems of increased complexity, the MCMC baseline will be computational infeasible and hence the verification will most likely be limited to few hand-labeled failures. Extensions to adaptive or online training could improve robustness under distributional shift, while joint inference of latent trajectories and parameters may further enhance diagnostic resolution. More generally, integrating SBI-based inference into digital twin frameworks offers a pathway toward scalable, probabilistic condition monitoring in complex industrial systems.
VI.
and latent degradation parameters. The method remains robust across a range of scenarios, including sparseevent cases with weak parameter identifiability. From a reliability perspective, the framework provides the uncertainty-aware diagnostics essential for risk-informed decision-making and predictive maintenance. Crucially, its amortized and model-agnostic nature allows for the continuous inference of process and health parameters in realistic, plant-wide industrial settings where traditional MCMC sampling is entirely computationally unfeasible. Future work will focus on validation with real operational data and improving robustness under model mismatch.
VII.
DATA AND CODE AVAILABILITY
The code and experiments used in this paper are available on GitHub at: https://github.com/ petercollett-cognite/sbi_mcmc_heat_exchanger. git.
VIII.
AUTHORSHIP CONTRIBUTION STATEMENT
Peter Collet: Writing – review & editing, formal analysis. Simone Casolo: Writing – review & editing, formal analysis, conceptualization. Alexander J. Stasik: Review & editing, supervision, methodology, conceptualization. Signe Riemer-Sørensen: Writing – review & editing, supervision, methodology, conceptualization.
IX.
DECLARATION OF COMPETING INTEREST
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
CONCLUSIONS
This work introduces a Bayesian framework for heat exchanger condition monitoring using simulation-based inference (SBI) powered by amortized neural posterior estimation. The results show that SBI matches MCMC in diagnostic accuracy while reducing inference time by a factor of 82. This massive computational acceleration enables near real-time estimation of failure modes
[1] A. Ahmed Murtaza, A. Saher, M. Hamza Zafar, S. Kumayl Raza Moosavi, M. Faisal Aftab, and F. Sanfilippo,
X.
ACKNOWLEDGEMENTS
This publication has been funded by the SFI NorwAI, (Centre for Research-based Innovation, 309834). The authors gratefully acknowledge the financial support from the Research Council of Norway and the partners of the SFI NorwAI.
Results in Engineering 24, 102935 (2024).
14 [2] L. Villa and C. Zanini Brusamarello, The Canadian Journal of Chemical Engineering 103, 1786 (2025). [3] R. Qi, J. Zhang, and K. Spencer, Algorithms 16, 9 (2022). [4] A. Sahu, S. K. Palei, and A. Mishra, Expert Systems 41, 10.1111/exsy.13360 (2023). [5] J. Zou, T. Hirokawa, J. An, L. Huang, and J. Camm, Frontiers in Energy Research 11, 10.3389/fenrg.2023.1294531 (2023). [6] S. Sundar, M. C. Rajagopal, H. Zhao, G. Kuntumalla, Y. Meng, H. C. Chang, C. Shao, P. Ferreira, N. Miljkovic, S. Sinha, and S. Salapaka, Intl. J. of Heat and Mass Transf. 159, 120112 (2020). [7] Z. Wu, B. Zhang, H. Yu, J. Ren, M. Pan, C. He, and Q. Chen, Chem. Eng. Sci. 282, 119285 (2023). [8] R. Jradi, C. Marvillet, and M. R. Jeday, Sci. Rep 12, 20437 (2022). [9] M. Markowski and P. Trzcinski, Sci. Rep 15, 42078 (2025). [10] A. A. Bash, H. Karah, and A. A. Khail, Engineering Applications of Artificial Intelligence 151, 110723 (2025). [11] G. Hou, D. Zhang, Q. Yan, S. Wang, L. Ma, and M. Jiang, International Communications in Heat and Mass Transfer 164, 108809 (2025). [12] G. Hou, Z. An, D. Zhang, X. Du, Y. Ding, and J. Fu, International Communications in Heat and Mass Transfer 175, 111129 (2026). [13] H. Xu, S. Yan, X. Qin, W. An, and J. Wang, Processes 13, 10.3390/pr13010219 (2025). [14] S. Hosseini, A. Khandakar, M. E. Chowdhury, M. A. Ayari, T. Rahman, M. H. Chowdhury, and B. Vaferi, Energy Reports 8, 8767 (2022). [15] J. Berce, M. Bucci, M. Zupančič, M. Može, and I. Golobič, Applied Thermal Engineering 277, 126954 (2025). [16] A. Aboul Khail and A. Karah Bash, Thermal Science and Engineering Progress 71, 104565 (2026). [17] R. Ben-Mansour, S. El-Ferik, M. Al-Naser, B. A. Qureshi, M. A. M. Eltoum, A. Abuelyamen, F. Al-Sunni, and R. Ben Mansour, Energies 16, 10.3390/en16062812 (2023). [18] A. Etminan, K. Pope, and K. Mashayekh, AI Thermal Fluids 4, 100022 (2025). [19] A. Gelman, J. B. Carlin, H. S. Stern, and D. B. Rubin, Bayesian data analysis (Chapman and Hall/CRC, 1995). [20] R. Moradi, S. Cofre-Martel, E. Lopez Droguett, M. Modarres, and K. M. Groth, Reliability Engineering & System Safety 222, 108433 (2022). [21] M. E. Cholette, H. Yu, P. Borghesani, L. Ma, and G. Kent, Reliability Engineering & System Safety 183, 184 (2019). [22] S. Bardeeniz, C. Panjapornpon, P. Chomchai, and M. A. Hussain, Reliability Engineering & System Safety 262, 111250 (2025). [23] G. L. Jones and Q. Qin, Annual Review of Statistics and Its Application 9, 557 (2021). [24] S. Xiao and W. Nowak, Reliability Engineering & System Safety 266, 111718 (2026). [25] E. Zio and G. Peloni, Reliability Engineering & System Safety 96, 403 (2011). [26] N. Nazemzadeh, J. Loyola-Fuentes, G. Righetti, E. DiazBejarano, S. Mancin, and F. Coletti, AI Thermal Fluids 5, 100030 (2026). [27] J. Mo and W.-J. Yan, Reliability Engineering & System Safety 264, 111337 (2025).
[28] A. Dasgupta and E. A. Johnson, Reliability Engineering & System Safety 242, 109729 (2024). [29] K. Cranmer, J. Brehmer, and G. Louppe, Proceedings of the National Academy of Sciences 117, 30055 (2020). [30] M. Deistler, J. Boelts, P. Steinbach, G. Moss, T. Moreau, M. Gloeckler, P. L. C. Rodrigues, J. Linhart, J. K. Lappalainen, B. K. Miller, P. J. Gonçalves, J.-M. Lueckmann, C. Schröder, and J. H. Macke, Simulation-based inference: A practical guide (2025), arXiv:2508.12939 [stat.ML]. [31] Álvaro Tejero-Cantero, J. Boelts, M. Deistler, J.-M. Lueckmann, C. Durkan, P. J. Gonçalves, D. Greenberg, and J. H. Macke, arXiv 10.48550/arxiv.2007.09114 (2020). [32] G. Papamakarios and I. Murray, Fast ϵ-free inference of simulation models with bayesian conditional density estimation (2018), arXiv:1605.06376 [stat.ML]. [33] G. Papamakarios, D. Sterratt, and I. Murray, in Proceedings of the Twenty-Second International Conference on Artificial Intelligence and Statistics, Proceedings of Machine Learning Research, Vol. 89, edited by K. Chaudhuri and M. Sugiyama (PMLR, 2019) pp. 837–848. [34] S. He, Y. Ye, M. Wang, J. Zhang, W. Tian, S. Qiu, and G. Su, Nuclear Engineering and Design 432, 113759 (2025). [35] J. Degrave, M. Hermans, J. Dambre, and F. Wyffels, A differentiable physics engine for deep learning in robotics (2018), arXiv:1611.01652 [cs.NE]. [36] R. Majumdar, V. Jadhav, A. Deodhar, S. Karande, L. Vig, and V. Runkana, Numerical Heat Transfer, Part B: Fundamentals 86, 1910 (2025). [37] M. Shi, L. Sun, Z. Bi, and R. He, Computers & Chemical Engineering 201, 109259 (2025). [38] T. L. Bergman, A. S. Lavine, F. P. Incropera, and D. P. DeWitt, Fundamentals of heat and mass transfer (John Wiley & Sons, 2011). [39] J.-X. Zhang, J.-L. Zhang, Z.-X. Zhang, T.-M. Li, and X.-S. Si, Reliability Engineering & System Safety 250, 110223 (2024). [40] P. Wang, Q. Liu, D. Xu, G. Jiang, W. Peng, F. Liu, and J. Hou, International Journal of Heat and Mass Transfer 256, 128183 (2026). [41] P. Tankov and R. Cont, Financial Modelling with Jump Processes (Chapman and Hall/CRC, Boca Raton, 2003). [42] W. Ebert and C. B. Panchal, in Fouling mitigation in industrial heat exchanger equipment, edited by C. B. Panchal, T. R. Bott, E. F. C. Somerscales, and S. Toyama (Begell House, 1997) pp. 451–460. [43] D. van Ravenzwaaij, P. Cassey, and S. D. Brown, Psychonomic Bulletin & Review 25, 143 (2018). [44] D. M. Blei, A. Kucukelbir, and J. D. McAuliffe, Journal of the American Statistical Association 112, 859–877 (2017). [45] M. D. Hoffman and A. Gelman, J. Mach. Learn. Res. 15, 1593–1623 (2014). [46] M. Sunnåker, A. G. Busetto, E. Numminen, J. Corander, M. Foll, and C. Dessimoz, PLoS Computational Biology 9, e1002803 (2013). [47] D. S. Greenberg, M. Nonnenmacher, and J. H. Macke, in International Conference on Machine Learning (2019). [48] J. Boelts, M. Deistler, M. Gloeckler, Álvaro TejeroCantero, J.-M. Lueckmann, G. Moss, P. Steinbach, T. Moreau, F. Muratore, J. Linhart, C. Durkan, J. Vetter, B. K. Miller, M. Herold, A. Ziaeemehr, M. Pals,
15 T. Gruner, S. Bischoff, N. Krouglova, R. Gao, J. K. Lappalainen, B. Mucsányi, F. Pei, A. Schulz, Z. Stefanidi, P. Rodrigues, C. Schröder, F. A. Zaid, J. Beck, J. Kapoor, D. S. Greenberg, P. J. Gonçalves, and J. H. Macke, Journal of Open Source Software 10, 7754 (2025). [49] J. E. Matheson and R. L. Winkler, Management Science 22, 1087 (1976). [50] C. Yang, M. Ai, Z. Tang, and Y. Xie, Engineering Applications of Artificial Intelligence 162, 112236 (2025). [51] L. Chang and Y.-H. Lin, Engineering Applications of Artificial Intelligence 142, 109980 (2025). [52] D. Phan, N. Pradhan, and M. Jankowiak, arXiv preprint arXiv:1912.11554 (2019).
[53] E. Bingham, J. P. Chen, M. Jankowiak, F. Obermeyer, N. Pradhan, T. Karaletsos, R. Singh, P. A. Szerlip, P. Horsfall, and N. D. Goodman, J. Mach. Learn. Res. 20, 28:1 (2019). [54] C. Durkan, A. Ivanov, A. Maddison, and G. Papamakarios, in Advances in Neural Information Processing Systems, Vol. 32 (2019) pp. 7511–7522. [55] A. Delaunoy, B. K. Miller, P. Forré, C. Weniger, and G. Louppe, Preprint (2023), arXiv:2304.10978 [stat.ML]. [56] M. Falkiewicz, N. Takeishi, I. Shekhzadeh, A. Wehenkel, A. Delaunoy, G. Louppe, and A. Kalousis, Preprint (2023), arXiv:2310.13402 [stat.ML]. [57] X. Xu, J. Zhou, X. Weng, Z. Zhang, H. He, F. Steyskal, and G. Brunauer, Reliability Engineering & System Safety 250, 110250 (2024).