Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice Chem Rev . 2026 Mar 4;126(7):4189–4236. doi: 10.1021/acs.chemrev.5c00931 Search in PMC Search in PubMed View in NLM Catalog Add to search Uncertainty Quantification for In Silico Chemistry Tom Frömbgen Tom Frömbgen † Mulliken Center for Theoretical Chemistry, Clausius-Institute for Physical and Theoretical Chemistry, University of Bonn, Beringstraße 4, 53115 Bonn, Germany Find articles by Tom Frömbgen † , Elizaveta Surzhikova Elizaveta Surzhikova ‡ Institute of Physical and Theoretical Chemistry, Technische Universität Braunschweig, Gaußstr. 17, 38106 Braunschweig, Germany Find articles by Elizaveta Surzhikova ‡ , Jürgen Dölz Jürgen Dölz ¶ Institute for Numerical Simulation, University of Bonn, Friedrich-Hirzebruch-Allee 7, 53115 Bonn, Germany Find articles by Jürgen Dölz ¶, * , Jonny Proppe Jonny Proppe ‡ Institute of Physical and Theoretical Chemistry, Technische Universität Braunschweig, Gaußstr. 17, 38106 Braunschweig, Germany Find articles by Jonny Proppe ‡, * , Barbara Kirchner Barbara Kirchner † Mulliken Center for Theoretical Chemistry, Clausius-Institute for Physical and Theoretical Chemistry, University of Bonn, Beringstraße 4, 53115 Bonn, Germany Find articles by Barbara Kirchner †, * , Christoph R Jacob Christoph R Jacob ‡ Institute of Physical and Theoretical Chemistry, Technische Universität Braunschweig, Gaußstr. 17, 38106 Braunschweig, Germany Find articles by Christoph R Jacob ‡, * Author information Article notes Copyright and License information † Mulliken Center for Theoretical Chemistry, Clausius-Institute for Physical and Theoretical Chemistry, University of Bonn, Beringstraße 4, 53115 Bonn, Germany ‡ Institute of Physical and Theoretical Chemistry, Technische Universität Braunschweig, Gaußstr. 17, 38106 Braunschweig, Germany ¶ Institute for Numerical Simulation, University of Bonn, Friedrich-Hirzebruch-Allee 7, 53115 Bonn, Germany * Email: [email protected] . * Email: [email protected] . * Email: [email protected] . * Email: [email protected] . Received 2025 Oct 28; Accepted 2026 Feb 13; Revised 2026 Feb 4; Collection date 2026 Apr 8. © 2026 The Authors. Published by American Chemical Society This article is licensed under CC-BY 4.0 PMC Copyright notice PMCID: PMC13067284 PMID: 41779034 Abstract The rapid growth of worldwide computing power has transformed in silico chemistry into a discipline that is integrated into the daily work of many chemists. Nowadays, researchers find it increasingly straightforward to predict a wide range of molecular properties and chemical processes at reasonable computational cost. The resulting abundance of data generated by quantum chemistry, molecular dynamics simulations, and chemical machine learning naturally raises questions about accuracy, precision, and reliability as well as the systematic treatment of errors and uncertainties. Addressing these questions through rigorous mathematical frameworks is at the heart of uncertainty quantification. In the past years, the incorporation of uncertainty quantification into in silico chemistry has gained attention, motivated by its ability to provide deeper insights into chemical phenomena. In this review, we establish a common language for uncertainty quantification with respect to in silico chemistry, introduce the key mathematical formalisms, and survey the growing body of work that applies uncertainty quantification across different areas of in silico chemistry. 1. Introduction A preprint of this work was previously published at https://chemrxiv.org/ . In physics one is constantly faced with the task of having to draw conclusions from imperfect information. Computational chemistry and the associated software have reached a level of maturity where even novices or nonexperts in the field can perform calculations, conduct computer-based experiments, and generate data. At the same time, scientific targets are becoming increasingly complex, often demanding significant computational resources to uncover new insights into systems or phenomena that are not yet (fully) understood. Examples posing such challenges include modeling large ensembles of molecules or interfaces and performing high-accuracy quantum-chemical calculations, which require substantial sampling and/or large basis sets. Furthermore, the desire to tackle more complex chemical systems is frequently hindered by a lack of sufficient information or input data, casting doubt on the reliability of highly accurate calculations. Therefore, the resulting data are often far from being perfect and carry uncertainties that modern in silico chemistry should take into account. In that sense, the expressive motto from the review article of von Toussaint on Bayesian methods in physics seamlessly translates to in silico chemistry. The aforementioned aspects are illustrated by the transformation the daily routine of in silico chemistry has taken. For example, in ionic liquid research, − a couple of decades ago, research focused on studying structures or component specification with sophisticated force fields in molecular dynamics simulations or intermolecular interactions in quantum-chemical calculations between one cation and one anion. − Nowadays, more involved multicomponent mixtures or transport properties like diffusion coefficients or ionic conductivities are considered. − This requires longer simulation times, larger system sizes, and also more refined force fields. − To deal with future requirements and applied topics, e.g., deep eutectic solvents, − gas absorption, or energy storage materials ,,, additional phenomena such as electric fields, solid–liquid interfaces, and growth phenomena need to be accounted for, adding a layer of complexity and dependence on possibly inexact information. A similar transformation can be observed in the field of reaction mechanism exploration. In the past, computational studies were often limited to isolated reaction steps, treating only reactants, transition states, and products within a single elementary reaction. With increasing computational power and more sophisticated algorithms, researchers now construct complex reaction networks where each node represents not just a single structure but an ensemble of conformations. Automated approaches have been developed to construct complex reaction networks and explore reaction mechanisms for numerous reactant molecules. − The complexity of these networks can be further increased by introducing environment effects such as microsolvation. , Beyond reaction networks, similar advancements have transformed other domains of computational chemistry. The study of catalytic systems, for example, has shifted from static models of active sites to dynamic simulations incorporating explicit solvent, surface reconstructions, and even operando conditions. , Machine learning methods now complement traditional quantum-chemical calculations by predicting activation energies, optimizing reaction pathways, or accelerating molecular dynamics simulations. − A particularly pressing challenge for the future is the accurate and efficient description of excited-state processes, which are critical in fields ranging from photochemistry and spectroscopy to optoelectronic materials. While current methods provide valuable insights, they are often computationally prohibitive when applied to large or complex systems. The ability to construct and analyze potential energy surfaces that span multiple electronic states would mark a major milestone in theoretical chemistry, further pushing the boundaries of what computational approaches can achieve. Recent developments in computational methods have enabled the efficient calculation of excited states, allowing reaction networks to represent multiple potential energy surfaces. , However, all such applications of in silico chemistry to ever more complex systems rely on imperfect models and the resulting data is prone to uncertainties. When dealing with imperfect data, it is key to acknowledge the fact that imperfect data require appropriate statistical treatment to quantify the level of “imperfectness”. By taking a step back from the illusion of perfection, and by applying techniques from uncertainty quantification (UQ), even more beneficial conclusions can be drawn. As a valuable starting point for uncertainty quantification, the following definition of Molinero and co-workers can be considered: Uncertainty quantification (UQ) describes the process of quantifying uncertainties and how they are propagated in a model system of interest, such as a molecular simulation. Following this process and drawing appropriate conclusions should be one of the key components of any scientific workflow to assess the quality of the scientific results. The assessment of uncertainties in experimental sciences is well-established. The most basic and general work is provided by Evaluation of Measurement DataGuide to the Expression of Uncertainty in Measurement (GUM). It is an international standard for quantification of uncertainties independent of the methods applied. An easier-to-digest work is provided in the Guidelines for Evaluating and Expressing the Uncertainty of NIST Measurement Results (NIST Technical Note 1297). This work is comparable to GUM, but less general and mainly focused on experiments. For chemistry, there is a IUPAC guide for quantifying uncertainties in experimental measurements, and basic requirements for reporting experimental uncertainties are provided, e.g., in guidelines of scientific journals such as The Journal of Physical Chemistry . While many principles from the experimental sciences can also be applied in the computational sciences, the latter often require different strategies for UQ. The mathematical foundations of uncertainty quantification in computer simulations have been well investigated in the last decades, see for example the several mathematics-centered books − and reviews − that are available. These mathematical tools have become an indispensable part of different scientific disciplines, such as climate modeling and engineering. − In the field of in silico chemistry, the assessment of uncertainties is less well-established, and has only been emerging rather recently. For overview articles highlighting the importance of uncertainty quantification in computational chemistry, we refer the reader to refs − . Here we review applications of UQ in different areas of in silico chemistry, covering quantum chemistry, molecular dynamics simulations, and chemical machine learning. We will address UQ in materials modeling only as far as it directly relates to atomistic simulations. Also, UQ for quantum computing with applications to in silico chemistry is beyond the scope of this review, and the reader is kindly referred to other works and reviews. − The above quote by Molinero and co-workers highlights that uncertainty quantification is a multistep workflow. As a minimal or initial step, UQ involves identifying the stages at which uncertainties can enter a computation, such as through uncertain parameters. At a more advanced level, it requires assessing how these uncertainties propagate through the computation, ultimately affecting the quantity of interest (QoI). Additional insights into QoIs can be gained by comparing them to highly accurate reference values (referred to as benchmarks) to quantify the uncertainties. We stress here that the latter step is optional, and while this requires benchmark values to compare with, the aim of UQ is different from those of benchmark calculations. UQ helps to understand how actionable conclusions can be drawn from current state-of-the-art chemical computationsan endeavor that may still lead to an added value, despite possible shortcomings of the calculations. This in turn leads to another step of UQ, where UQ can guide researchers to improved approaches, for example with inverse UQ or multifidelity methods. Thus, UQ clearly holds the potential for novel method developments. Moreover, besides basic research interest, there is also the potential for the development of novel approaches that UQ permits, as will be shown in this review. UQ can also serve as a control mechanism to counteract faulty data production and misconduct behavior. , In view of the ever-increasing amount of data produced as well as the growing capabilities of artificial intelligence (AI) and machine learning (ML) , in almost all scientific research fields, UQ constitutes a cornerstone of control mechanisms to maintain scientific quality and reproducibility. , Similarly, as outlined in this review, UQ is not at all restricted to machine learning albeit being an important and necessary part of itand vice versawhich is why this review should not be subsumed under many excellent reviews about machine learning in chemistry, ,− but contain some of the work that strongly focus on UQ aspects. In this review, we establish a common language for uncertainty quantification with respect to in silico chemistry, introduce the key mathematical formalisms, and survey the growing body of work that applies uncertainty quantification across different areas of in silico chemistry. We first establish a common language for uncertainty quantification with respect to in silico chemistry in section by presenting key terms and definitions in section , discussing the most important sources of uncertainty in section , and demonstrating the UQ workflow in section . This is followed by a rigorous introduction of basic mathematical principles on a level that should be understandable to in silico chemists in section : we address the modeling of uncertainties in computations ( section ), problem classes of UQ ( section ), sensitivity analysis ( section ), sampling methods ( section ), surrogate models ( section ), uncertainty validation ( section ), dimension reduction ( section ) and provide a brief overview on UQ software with potential applications to in silico chemistry ( section ). Section deals with the applications of UQ to in silico chemistry, focusing on uncertainties in energies in section , in trajectories and other simulation-based properties in section and in static properties and electronic structure applications in section . We close this review with recommendations and conclusions in section . 2. Taxonomy of Uncertainty Quantification 2.1. Key Terms and Definitions This section introduces core concepts and definitions for clear and easier understanding of meaningful uncertainty quantification for in silicio chemistry. • Quantity of interest is the particular quantity that is going to be calculated and that is affected by the uncertainty. This is usually an output quantity of a calculation, rather than an input quantity. • Reference data or reference values represent the best available estimate of the true values and are usually themselves uncertain. They may be derived from high-quality experimental data or from more accurate or established computational methods. • Error refers to the difference between a quantity of interest (outcome) and a suitable reference value. • Precision refers to the consistency or repeatability of outcomes under the same conditions. Precision is the inverse of noise , also denoted random error (see Figure ). 1. Open in a new tab Illustration of accuracy and precision by black dots on gray/red targets. • Accuracy describes how close an outcome is to its reference value. Accuracy is the inverse of bias , also denoted systematic error (see Figure ). • Uncertainty denotes the expected deviation of the outcome from a suitable reference value, prior to any comparison. While an error is the deviation of one individual calculated quantity of interest from the reference value, uncertainty refers to the expected distribution of the (unknown) error of an individual calculated quantity. This distribution can be made tangible by various quantities, such as, e.g., its standard deviation or a confidence interval. • Epistemic uncertainty arises from a lack of knowledge about the chemical system or the computational model. It can, in principle, be reduced through improved models or more precise parametrizations. • Aleatoric uncertainty is the part of overall uncertainty that cannot be reduced by more extensive sampling or by collecting more data. In computational chemistry, aleatoric uncertainty typically reflects variability in experimental reference data or stochasticity in simulations. • Uncertainty quantification (UQ) refers to the systematic assessment and documentation of uncertainty in computational results. It encompasses the identification of uncertainty sources (see next section), their quantification, and their impact on the outcome. • Benchmarking refers to the comparison of computational methods against reference data or other methods using predefined metrics. Unlike UQ, benchmarking assesses global performance rather than the reliability of individual predictions. ,− • Errors are deterministic if they do not change when a calculation is repeated with the exact same input data. • Errors are stochastic if they change randomly upon repeating a calculation with the exact same input data. These definitions serve as the conceptual basis for discussing uncertainties with regard to in silico chemistry. 2.2. Sources of Uncertainty in In Silico Chemistry Generally speaking, computational chemistry targets some quantity of interest (QoI). The reference value f (ξ) is a function of the chemical system ξ and can (at least in principle) be measured in an experiment. This can, for instance, be a reaction energy, a bulk property such as the density or the diffusion coefficient, or an infrared (IR) spectrum. One then employs the computer implementation of some suitable model f comp (ξ chem , η, h ), which will depend on parameters describing (a model of) the chemical system ξ chem as well as additional parameters η and h of the model and its numerical approximations, respectively, to calculate this QoI. The total error of this model is then given by e tot = f comp ( ξ chem , η , h ) − f ( ξ ) 1 Uncertainty quantification aims at the systematic assessment and documentation of this error, i.e., the deviation of the predictions for a specific chemical system from the (usually unknown) chemical reality or (as proxy) an experimental or computational benchmark. It encompasses the identification of uncertainty sources (see next section), their quantification, and their impact on the outcome. The uncertainty is introduced by many different sources. Their relative importance, and whether they are classified as aleatoric or epistemic, will depend both on the specific problem at hand (i.e., the quantity of interest in the computations and the benchmark that is targeted) and on the computational methods that are used. As different sources of uncertainty will generally require different approaches for their quantification, it will be important to specify which sources of uncertainty are included in a specific method for uncertainty quantification in computational chemistry and to acknowledge sources of uncertainty that are not addressed. Neglecting some sources of uncertainty will generally correspond to redefining the benchmark that is used for comparison. While many sources of uncertainty in computational chemistry have been discussed in the literature, there is currently no generally accepted classification scheme. Here, we will build on the categorization employed by Lejaeghere, who distinguishes representation uncertainties , level of theory uncertainties , and numerics uncertainties ( Figure ). In all three classes, uncertainties can be both deterministic and stochastic. 2. Open in a new tab Illustration of the models (in colored boxes) and errors (denoted by arrows connecting the models) typically introduced during in silico chemistry workflows (also refer to Tables and ). Representation Uncertainties Lejaeghere defines these as follows: “This broadest class of errors represents all deviations between theory and reality because of deliberate approximations or inadvertent assumptions [that are not directly linked to the level of theory]. They correspond to a mismatch between what is calculated and what is experimentally measured, i.e., an incomplete atomistic representation of the conditions of the macroscopic material under study.” First, this includes simplifications that are necessary to tackle a chemical problem computationally, which will lead to inadvertent mismatch between experiment and computation, i.e., real chemical systems are much more complicated than any computational model. In calculations for solids, this might include the neglect of impurities, disorder or microstructure. Similarly, when modeling solution chemistry side products or other impurities might be neglected. Boundary conditions usually cannot be represented faithfully in computational models, and one instead uses, e.g., periodic boundary conditions instead of finite-size bulk system, or an isolated molecule in vacuum instead of a gas-phase system at finite pressure. Second, there are deliberate simplifications and approximations. Often, complex chemical systems are computationally not directly accessible, but simplified molecular models (e.g., using simplified ligands in a transition metal complex) have to be used instead. Similarly, it is common to either neglect the chemical environment of the chosen molecular model or to treat it using approximate models (e.g., a continuum solvation model). This can both be a deliberate approximation or a necessary simplification because of a lack of knowledge on the atomistic structure of the environment. Altogether, such approximations define a chemical model ξ chem , and we can define the representation error as e repr = f ( ξ ) − f ( ξ chem ) 2 Level of Theory Uncertainties These are related to the approximations that are introduced in the computational chemistry methods that are used, which give rise to a well-defined mathematical model. This mathematical model f math (ξ chem , η) will (explicitly or implicitly) depend on additional parameters η. For quantum-chemistry, sources of such level of theory uncertainties include: the use of the Born–Oppenheimer approximation and a neglect of adiabatic corrections, the use of the nonrelativistic Schrödinger equation (i.e., neglect of relativistic and QED effects), the use of a simplified ansatz for the electronic wavefunction or the use of an approximate exchange–correlation functional in density-functional theory, the use of pseudopotentials for the core electrons, and the truncation of the one-electron basis set. For theoretical spectroscopy, further approximations in the respective methodology for calculating spectra (such as, e.g., the use of the harmonic approximation in the calculation of vibrational spectra) will introduce uncertainties. For molecular dynamics simulations, , sources of such uncertainties include the neglect of nuclear quantum effects in the Born–Oppenheimer approximation, the model that is assumed for the potential energy surface (e.g., the force field that is used), and the time scales of the simulation. The approximations within the mathematical model lead to the level of theory error, e lot = f ( ξ chem ) − f math ( ξ chem , η ) 3 Numerics Uncertainties The computer implementation of the mathematical model provides the final output of a computation, f comp (ξ chem , η, h ) and will depend on additional parameters h . Lejaeghere defines numerics uncertainties as follows: “Numerical errors [...] originate in the technical parts of a [DFT] calculation and vary from purely hardware- or algorithm-related effects, such as floating-point precision, to more physical approximations to reduce the computational load, such as series truncations or the use of a particular pseudopotential.” For the case of solid-state DFT calculations using periodic boundary conditions, Herbst et al. list discretization error (such as Brillouin zone sampling, finite basis set of plane waves), algorithmic error (e.g., self-consistent Kohn–Sham equations are solved by a self-consistent field (SCF) algorithm, in which at each step of the SCF algorithm, a linear eigenvalue equation is solved with an iterative eigensolver), and arithmetic error (due to finite floating-point precision). They also point out that “to these must be added programming errors, not neglected in codebases consisting of millions of lines of code, and hardware errors, which are expected to become significant for exascale architectures”. For molecular dynamics simulations, the choice of a finite time step and cutoffs in the evaluation of the force field as well as the integration of the equations of motion itself contribute to numerics uncertainties. The numerics error can be expressed as e num = f math ( ξ chem , η ) − f comp ( ξ chem , η , h ) 4 Discussion The different models defined by Lejaeghere as discussed above are presented in Table , which also includes examples of corresponding approximations in computational chemistry. The errors introduced in the different steps are summarized in Table . Using the introduced notation and the triangle inequality, the magnitude of the total error can be decomposed into the error made by each step as follows: | f ( ξ ) − f comp ( ξ chem , η , h ) | ︸ total error ≤ | f ( ξ ) − f ( ξ chem ) | ︸ representation error + | f ( ξ chem ) − f math ( ξ chem , η ) | ︸ level of theory error + | f math ( ξ chem , η ) − f comp ( ξ chem , η , h ) | ︸ numerics error Often, the errors of the different steps can be investigated separately. 1. Functions Describing the Sources of Uncertainty Introduced in a Typical Computational Chemistry Workflow, as Discussed in Section , Including Descriptions and Examples from Quantum Chemistry (QC) and Molecular Dynamics Simulations (MD) Perspectives . function of model description example f (ξ) The QoI as a function of the underlying chemical system ξ. QC: electronic energy, vibrational spectrum, reaction barrier; MD: density, diffusion coefficient, spectrum, (thermodynamic and transport properties), free energy. f (ξ chem ) Chemical model. ξ chem denotes the description of molecules (atomic resolution, in- or excluding electrons, coarse graining), structure ( xyz coordinates of molecules), amount (single molecule, cluster, ensemble), and dynamical aspects (conformation, propagation in time). Isolated molecule for electronic energy, molecule dissolved in liquid for explicit solvation, transition state of reaction for determining barrier, many molecules of a liquid simulated for several ns to determine diffusion coefficient, many molecules in crystal or surface, interface between many liquid molecules and solid molecules. f math (ξ chem , η) Mathematical model . Everything that can be expressed in a mathematical notation and that is related to the computational chemistry method; may depend on additional parameters η. Electronic-structure method (HF, post-HF, DFT), including LCAO ansatz, force field (atomistic, polarizable, coarse grained), equation of motion (velocity Verlet, leapfrog), equations for thermostats or barostats. f comp (ξ chem , η, h ) Result of the computer implementation of f math (ξ chem , η), including parametrizations, mesh width, tolerances of solvers, etc., encoded in h . Electronic energy at a chosen basis set, trajectory (includes coordinates and velocities at certain time steps for a certain duration of time), postprocession of the trajectory, free energy. Open in a new tab a For abbreviations used, refer to textbooks on QC or MD. − . 2. Compilation of Errors Occurring with In Silico Chemistry Workflows and Their Relation to Different Models . representation error e repr = f (ξ) – f (ξ chem ) level of theory error e lot = f (ξ chem ) – f math (ξ chem , η) numerics error e num = f math (ξ chem , η) – f comp (ξ chem , η, h ) total error e tot = f (ξ) – f comp (ξ chem , η, h ) Open in a new tab a Also refer to Table and Figure . However, the classification of different types of uncertainty is often not straightforward in computational chemistry. For instance, while the truncation of the one-electron basis set in molecular quantum chemistry (using atom-centered basis functions) is best classified as level-of-theory uncertainty, the truncation of a plane-wave basis set in condensed phase DFT calculations is governed by a single cutoff and can thus be classified as a source of numerics uncertainty. Moreover, a core challenge in UQ for chemical computations is the lack of a clear hierarchy among sources of error. In the above classification, it is implicitly assumed that a computational model is built in a sequential manner: First, a structural model is chosen (e.g., positions of point-like nuclei), then a mathematical model (e.g., a specific level of approximation), and finally, numerical parameters are set. However, this linear view can be misleading. Theoretical assumptions often constrain which structural models are even meaningful or valid. Point-like nuclei, for example, are not merely a convenience but a precondition for the analytical integration of Coulomb operators. This feedback loop blurs the boundaries between different sources of error and complicates efforts to assign uncertainties in a strictly hierarchical fashion. To navigate this complexity, it is often useful to distinguish between the modeling workflow and the classification of error sources. While the scheme proposed by Lejaeghererepresentation, level of theory, numericsmaps well onto the typical structure of computational models and was therefore used here to introduce key uncertainty concepts, it does not fully resolve the hierarchy problem. An alternative taxonomy, put forward by Pernot, offers a more technically appropriate framework for analyzing uncertainty propagation and assessing the potential for uncertainty reduction. He distinguishes model uncertainty , parameter uncertainty , and numerical uncertainty . Using the terminology introduced above, model uncertainty refers to uncertainties due to both the chemical model ξ chem and the functional form of the mathematical model f math . Uncertainties due to the chemical or mathematical model are almost always present in in silico chemistry. Their estimation is usually infeasible due to a lack of understanding of the property of interest f (ξ) and/or it requires a significant amount of domain specific knowledge. The best available option to quantify such uncertainties seems to be to compare models between each other or to consider a presumably highly precise or almost accurate model as the ground truth. Model uncertainties are systematic. Uncertainties in the parameters of the mathematical model η, for example due to measurement errors and limited data availability, propagate through the model and affect its outcome. Since the precise error is unknown or uncertain , it is often modeled as random. Measurement errors are thus usually statistical errors and as such, can be quantified by statistical methods. Numerical uncertainty comprises all uncertainties caused by algorithms for implementing the mathematical model of the target system and during runtime of the calculations. This includes round-off errors during linear algebra computations and tolerances h in approximate solvers. Computational errors are mostly systematic errors when computations are done on a single processor. Computations performed on parallel architectures are well-known for implying a stochastic aspect in the outcomes, particularly when many processors are involved. The two discussed schemes for the classification of sources of uncertainty by Lejaeghere and Pernot share the distinction of numerical uncertainties . For the remaining sources of uncertainty, they employ complementary classifications. Pernot’s approach assumes a well-defined mathematical model that depends on uncertain parameters to distinguish model uncertainty and parameter uncertainty . While this might be an appropriate picture in some cases (such as regression models or more general ML-models with a well-defined reference), general modeling problems in in silico chemistry are often more complex and do not allow for a clear distinction of the computational model and its parameters. The scheme by Lejaeghere attempts to tame this complexity by distinguishing a chemical model ( representation uncertainty ) and a mathematical model ( level of theory uncertainty ). Both schemes are complementary, but neither seems to be perfectly suited to discuss all sources of uncertainty (and methods for their quantification) in in silico chemistry. 2.3. Uncertainty Quantification Workflow Given a target quantity of interest f (ξ) and a computational model f comp (ξ chem , η, h ), an outline of an uncertainty quantification workflow could be sketched as follows: 1. Analyze the sources of uncertainty in the parameters of the computational model as outlined in section . 2. Analyze the sources of uncertainty in the parameters to determine whether they are deterministic or stochastic. 3. Understand and bound deterministic errors. 4. Develop a prior stochastic model for the stochastic errors. 5. If benchmarks (reference values) are available: (a) Repeat steps items 1 to 4 for the benchmark. (b) Use benchmark to update the prior model for the stochastic errors by means of an algorithm. 6. Compute desired statistical quantities of interest by means of an algorithm. While items 1 and 2 can usually be answered by looking into the design of the computational model, the remaining three steps are quite challenging, but for different reasons. Items 3 and 4 are challenging from a theoretical perspective. As outlined in section , the error in the output data of a computational chemistry method can be bounded by the chemical and mathematical modeling errors, the computational error, and the measurement error. Modeling and computational errors are often considered as deterministic errors, but the complexity of the underlying models makes them difficult to understand even for domain experts. Measurement errors are usually modeled as stochastic error and can be quantified by computing various statistical quantities of interest for which a stochastic model is required. Developing stochastic models for measurement errors and errors in benchmarks is a topic on its own, and we refer to the relevant literature. , On the other hand, the challenge of items 5 and 6 is mainly due to computational complexity, as computational methods in uncertainty quantification require many evaluations of computational chemistry methods, which are notoriously compute intensive. Two main problem classes can be distinguished, which we review in the following. 3. Theory of Uncertainty Quantification 3.1. Modeling Uncertainties in Computations 3.1.1. A Crash Course in Probability Theory and Bayes’ Formula The input data subject to stochastic uncertainty can be considered as a random experiment . A random experiment is associated with • the outcomes of the experiment, also referred to as samples or realizations ; • the sample space , that is, a set Ω consisting of all conceptually possible outcomes; • the events , being subsets of Ω for which probability is defined. A canonical example of a random experiment is the toss of a coin, where Ω = {heads, tails}. However, in the framework of computational chemistry, we can also think of the sample space as a set of possible molecule configurations or a set of functions describing a charge density. To each event A ⊂ Ω we can assign a value P ( A ) ∈ [0, 1]. We say that P is a probability function , if P (Ω) = 1 and P ( ∪ i = 1 ∞ A i ) = ∑ i = 1 ∞ P ( A i ) 5 for mutually disjoint sets A i ∈ Ω which do not need to form a partition of Ω. From this definition, the basic rules of calculus for probabilities can be derived. , For example, it is impossible that no event from the sample space occurs, i.e., P (Ø) = 0, the probability of the opposite of A ⊂ Ω is P (Ω \ A ) = 1 – P ( A ), and the sum of probabilities to two not necessarily disjoint events A , B ∈ Σ can be expressed as P ( A ) + P ( B ) = P ( A ∪ B ) + P ( A ∩ B ) 6 and eq also holds for a finite number of sets. Here and in the following we write A ∪ B for the union and A ∩ B for the intersection of the sets A and B (illustrated in Figure ). 3. Open in a new tab Illustration of the intersection (top) and union (bottom) of two sets A and B . Given two events A , B ⊂ Ω with P ( B ) > 0, the conditional probability of A given B can be expressed as P ( A | B ) = P ( A ∩ B ) P ( B ) 7 For P ( A ) = 0 we define P ( A | B ) = P ( A ). Using Bayes’ formula , P ( A | B ) can be expressed through P ( A | B ) = P ( B | A ) P ( A ) P ( B ) 8 In this formula, P ( A | B ) is also termed the posterior , P ( B | A ) is termed the likelihood of A given B , P ( A ) is termed the prior , and P ( B ) is termed the marginal probability. This formula is of particular importance when reference data and benchmarks need to be incorporated in inverse uncertainty quantification. Finally, we say that two events A , B ⊂ Ω with P ( A ), P ( B ) > 0 are (stochastically) independent if P ( A | B ) = P ( A ) , or equivalently , P ( B | A ) = P ( B ) 9 3.1.2. Statistical Models and Random Variables The quantities in computational chemistry which are affected by stochastic uncertainty can be considered as random variables . For the sake of simplicity in the presentation, we assume in the following that Ω ⊂ R d , i.e., the sample space of the input parameters consists of vectors in R d , where d is the dimension. This can be motivated by the fact that most uncertainties in the input data of computational chemistry methods can be parametrized by a finite collection of real numbers. A random variable is a function g (ω), ω ∈ Ω, for which there exists a probability distribution ρ: Ω → [0, ∞ ) such that P ( { g ( ω ) : ω ∈ A } ) = ∫ A ρ ( ω ) d ω 10 It is this probability distribution which encodes most of the properties of the random variable. It follows from the definition that probability densities must satisfy ∫ Ω ρ ( ω ) d ω = 1 11 Similar properties as for probabilities can be derived. For example, we say that two random variables x 1 and x 2 with probability distributions ρ 1 and ρ 2 are (stochastically) independent , if their joint probability distribution is given by ρ(ω 1 , ω 2 ) = ρ 1 (ω 1 )ρ 2 (ω 2 ). Moreover, conditional probability distributions and Bayes’ formula can be derived, which reads ρ ( ω | ω ′ ) = ρ ( ω ′ | ω ) ρ ( ω ) ρ ( ω ′ ) 12 A canonical example of a probability distribution is the normal or Gaussian distribution for real numbers x ∈ R : ρ ( x ) = 1 2 π σ exp [ − ( x − μ ) 2 2 σ 2 ] 13 where μ ∈ R and σ > 0. We write X ∼ N ( μ , σ ) for any random variable with this distribution. However, it is important to note that the concepts we discuss in the following are not restricted to Gaussian distributions, unless explicitly stated. We note that the values of a random variable can be quite general, covering real values, molecule configurations, or functions describing charge distributions, for example. In the setting of computational chemistry, the results of the computational chemistry methods g = f comp (ξ chem , η, h ) or the outcome of a benchmark experiment g = f bench (ξ) are canonical examples for random variables. The collection of a random variable with its probability distribution is also referred to as a statistical model . 3.1.3. Statistical Quantities of Interest The exact behavior of a random variable is often too complicated to be of interest in an uncertainty quantification analysis. Instead, different measures are usually considered. Most importantly, the mean of a random variable g is defined as E [ g ] = ∫ Ω g ( ω ) ρ ( ω ) d ω 14 and the variance and standard deviation are defined as V [ g ] = E [ g 2 ] − E [ g ] 2 , σ [ g ] = V [ g ] 15 Moreover, the correlation and covariance of g and some second random variable h are defined as Cor [ g , h ] = E [ g h ] 16 Cov [ g , h ] = Cor [ g , h ] − E [ g ] E [ h ] 17 We write Cor[ g ] = Cor[ g , g ] and observe that V [ g ] = Cov[ g , g ]. While a common understanding of random variables is that they should be real-valued, we mention for the sake of completeness that these quantities are also well-defined when g and h are function-valued. That is, they are real- or complex-valued functions g ( x , ω), h ( x , ω), depending on random events ω ∈ Ω and some additional parameter x . In this case, E [ g ] , V [ g ] , and σ[ g ] are functions of x as well. The correlation Cor[ g , h ] and covariance Cov[ g , h ] become functions in two variables, defined through Cor [ g , h ] ( x , y ) = ∫ Ω g ( x , ω ) h ( y , ω ) d ω Cov [ g , h ] ( x , y ) = Cor [ g , h ] ( x , y ) − E [ g ] ( x ) E [ h ] ( y ) For the even more general case when the product or integral are not defined, more advanced mathematical techniques can be used. Another measure to quantify the behavior of a random variable are confidence intervals , which express the probability that a random variable is within a certain range, i.e., P ( { a ≤ g ≤ b } ) ≤ ε 18 where 0 < ε < 1 and a and b are such that the comparison with g is reasonable (see Figure for an illustration). A special case thereof is the cumulative distribution function F ( a ) = P ( { a ≤ x } ) 19 which for real-valued random variables is the primitive of the probability density ρ. 4. Open in a new tab Illustration of normal (Gaussian) distributions. The top panel shows three distributions with different means and variances. The bottom panel illustrates the 1σ, 2σ, and 3σ confidence intervals. Finally, sensitivity indices (cf. section ) are measures that quantify how much a stochastic input parameter affects the outcome of a random model. As it turns out, it is quite common that the largest model variations are induced by only a few variables, with the remaining variables of reduced importance or negligible. The gained knowledge can be used for dimension reduction and the construction of surrogate models to obtain approximate but efficient computational models. 3.2. Problem Classes of Uncertainty Quantification Generally speaking, one can distinguish two problem classes of uncertainty quantification. In a forward uncertainty quantification problem (see section ) one considers given or assumed uncertainties of the input data and model parameters, and investigates how these uncertainties propagates through the computational model. This mainly corresponds to items 3 and 4 of the UQ workflow sketched in section . In an inverse uncertainty quantification problem, one takes benchmarks (reference data) into account and infers the uncertainties in the input data and/or model parameters by comparing these benchmarks to the output of the computational model. This mainly corresponds to item 5 of our general UQ workflow. In both problem classes, some preliminary knowledge or assumptions on the behavior of the uncertainty are required. Most prominently, it needs to be decided whether the errors are modeled as deterministic errors or stochastic errors or both. 3.2.1. Forward Uncertainty Quantification: Propagation of Uncertainty Forward uncertainty quantification evaluates how errors in input data lead to errors in output data of a computational chemistry method. This is under the assumption that the input data is modeled as a random variable with known probability distribution . In the simplest case, when only the model parameters η are affected by stochastic uncertainties, we may model their uncertainty as η = η 0 + ν 20 where ν is a random variable with known probability distribution ρ. This allows for the computation of statistical quantities of interest such as the mean, E [ f comp ( ξ chem , η , h ) ] = ∫ f comp ( ξ chem , η , h ) ρ ( η ) d η 21 by replacing the integral by a numerical approximation based on evaluations of f comp (ξ chem , η, h ) for certain values of η. Note that η may lie in a high-dimensional parameter space and, thus, the numerical evaluation of the integral may suffer from the curse of dimensionality and prohibitive computational cost. We comment on appropriate numerical methods below. 3.2.2. Inverse Uncertainty Quantification: Incorporating Benchmarks Inverse uncertainty quantification problems most prominently occur in item 5 of our general UQ workflow (see section ). Our starting point is a benchmark f bench (ξ) (i.e., experimental or computational reference data for a certain input ξ). The benchmark can itself by subject to some uncertainty with respect to the true reference value f (ξ). For simplicity of presentation we assume in the following that the such uncertainties are negligible and refer to refs and for the more general case. Given a benchmark f bench (ξ), we perform a comparison to the prediction of the computational model, f bench ( ξ ) − f comp ( ξ chem , η , h ) = ν 22 and interpret this as a measurement subject additive Gaussian noise which is modeled as a realization of a random variable. Given a prior probability density ρ prior (η) for the model parameters η, our aim is to update this probability distribution by incorporating knowledge from the benchmark. It then follows that ν is an N ( 0 , Σ ) -distributed random variable with probability distribution ρ N . From Bayes’ formula ( eq ), we obtain the a posterior distribution for η: ρ post ( η | f bench ) = ρ N ( f bench − f comp ( ξ chem , η , h ) ) ρ prior ( η ) Z 23 where Z = ∫ ρ N ( f bench − f comp ( ξ chem , η , h ) ) ρ prior ( η ) d η 24 As for the forward problem, the posterior distribution can be used to compute various statistical quantities of interest, such as the mean: E post [ f comp ( ξ chem , η , h ) ] = ∫ f comp ( ξ chem , η , h ) ρ post ( η | f bench ) d η 25 Here we need to point out that the integral of the normalization constant Z of the posterior distribution and for the posterior mean itself may be high-dimensional and thus compute intensive in practice. Further, we also note that the only difference between computing Z and the posterior mean is the presence of an additional factor f comp (ξ, ζ, h ) in the integrand. 3.3. Sensitivity Analysis 3.3.1. Global Sensitivity Analysis: Identifying Influence of Parameters Given a function in several variables, such as f : [ 0 , 1 ] n → R , the idea of sensitivity analysis is to identify the variables which influence the result of f the most. The most popular tool to do so are Sobol’ indices , which explore that f can be decomposed into f ( x 1 , x 2 , ... , x n ) = f 0 + ∑ i 1 f i 1 ( x i 1 ) + ∑ i 1 i 2 f i 1 i 2 ( x i 1 , x i 2 ) + ... 26 with f 0 = ∫ f ( x 1 , ... , x n ) d x 1 ··· d x n 27 f i 1 ( x i 1 ) = ∫ f ( x 1 , ... , x n ) ∏ k ≠ i 1 d x k − f 0 28 f i 1 i 2 ( x i 1 , x i 2 ) = ∫ f ( x 1 , ... , x n ) ∏ k ≠ i 1 , i 2 d x k − f 0 − f i 1 ( x i 1 ) − f i 2 ( x i 2 ) 29 etc. It is intuitively clear that the terms in eq provide some insight on how f depends on its various variables. For this reason, one also refers to eq as an analysis of variance, or similiarly, ANOVA decomposition . One can show that the variances of the entities are given as V [ f ] = ∫ f ( x ) 2 d x − f 0 2 30 V [ f i 1 ··· i s ] = ∫ f i 1 ··· i s ( x i 1 , ... , x i s ) 2 d x i 1 ··· d x i s 31 and that it holds V [ f ] = ∑ s = 1 n ∑ i 1 < · · · < i s V [ f i 1 ··· i s ] 32 The latter expression makes clear that the global sensitivity indices or Sobol’ indices S i 1 ··· i s = V [ f i 1 ··· i s ] V [ f ] 33 allow identification of the contributions to f which contribute the most to V [ f ] . Therein, s is called the order or dimension of the index. One can easily check that the indices are nonnegative and their sum is 1. The latter fact is particularly useful in practice, as one may only compute first order indices to begin with and can then increase the order until the sum of the indices is approximately one. Several other indices to quantify the impact of parameters exist and we refer to the relevant literature for an overview. ,, 3.3.2. Local Sensitivity Analysis: Perturbation Based Uncertainty Quantification While a global sensitivity analysis usually requires approaches requiring multiple or many evaluations of f ( x ), local sensitivity analysis , , also called the perturbation approach , , is well-suited for situations where the values of x are small random fluctuations of bounded magnitude h 0 around a reference value x 0 , i.e., x = x 0 + h , with h random and | h | < h 0 34 The randomness of the model outcomes can then be approximated as a Taylor expansion, i.e., f ( x ) = f ( x 0 + h ) ≈ f ( x 0 ) + h f ′ ( x 0 ) + h 2 2 f ″ ( x 0 ) + ... 35 Assuming some knowledge of the statistics of h , this allows for the approximation of statistical quantities of interest, such as the mean: E [ f ( x ) ] ≈ f ( x 0 ) + E [ h ] f ′ ( x 0 ) + E [ h 2 ] 2 f ″ ( x 0 ) + ... 36 For small perturbations, i.e., small h 0 , these approximations are often more accurate than sampling-based approaches and can save orders of magnitude in compute time. The quantities E [ h ] , E [ h 2 ] , etc. can be computed in closed form or approximated by efficient numerical methods without the need for sampling. ,− While local sensitivity analysis is traditionally only employed for forward problems, expansions for Bayesian inverse problems also exist. 3.4. Sampling Methods 3.4.1. Monte Carlo Methods One of the central tasks when computing statistical quantities of interest (see section ) such as the mean in eq is to evaluate the corresponding integrals. In general, since the functions to integrate are usually not available analytically, these integrals need to be evaluated by numerical methods. The most prominent and simplest approach is the Monte Carlo estimator . ,, That is, given a random variable f ( x ) with probability distribution ρ( x ), the integral is approximated as an average E [ f ] = ∫ Ω f ( ω ) ρ ( ω ) d ω ≈ 1 N ∑ i = 1 N f ( ξ i ) ≕ E N [ f ] 37 where ξ i , i = 1, ..., N , are identically distributed, stochastically independent (i.i.d. in short) random samples drawn from the distribution ρ. It is important to realize that the Monte Carlo estimator is a sum of random variables and, thus, a random variable itself, i.e., its outcome will change with every realization. To this end, it is well-known that the Monte Carlo estimator satisfies the following bound for the root-mean-square error : E [ ( E [ f ] − E N [ f ] ) 2 ] = V [ f ] N 38 and that for any ε > 0 it holds the failure probability P ( { | E [ f ] − E N [ f ] | > ε } ) ≤ V [ f ] N ε 2 39 The first estimate tells us that the root-mean-square error of the Monte Carlo method indeed converges to the true mean, with a convergence rate of O ( N − 1 / 2 ) . Unfortunately, this rate can be painstakingly slow in practice since in practice an evaluation of f ( x i ) will usually involve the run of a computational chemistry method, which can be quite expensive. For example, estimating the mean of f up to an accuracy of 10 –3 in the root-mean-square error requires 10 6 samples due to eq . However, the latter estimate tells us that, no matter how many samples we use in our Monte Carlo estimator, we can never exclude the possibility that the error of a single realization can be arbitrarily wrong, although with a probability which decreases with N . This unsatisfactory behavior of the Monte Carlo estimator has led to the development of a vast amount of methods which promise to overcome its downsides, the most prominent of which we review below. However, let us note that in some cases simple variance reduction techniques allow for small modifications in the sampling process to lower the magnitude of V [ f ] and thus improve the error and the failure probability. This being said, the Monte Carlo estimator is still the simplest method to implement and all of the following methods require additional information on the problem at hand. In contrast, the only assumption made by the Monte Carlo method is that V [ f ] is finite and that one can sample from the probability distribution of f . The latter assumption can be weakened using Markov chain Monte Carlo (MCMC) methods, , which do not require explicit access to the probability distribution. The development of MCMC methods is an active field of research, and over the years various popular developments such as adaptive MCMC, , Hamiltonian Monte Carlo, , and hybrid Monte Carlo have evolved. 3.4.2. Resampling Schemes In practice, it may happen that sampling from probability distribution ρ of a random variable X is not accessible due to theoretical, experimental, or computational constraints, but a finite number of samples X = { X 1 , ... , X M } of X is available. For example, properties derived from quantum-chemical calculations are often both high-dimensional with respect to molecular structure and expensive to obtain, which implies that the underlying distribution ρ remains effectively unknown. The idea of resampling schemes is to draw samples from the set X rather than sampling from ρ directly. This resampling procedure can then be used to compute statistical quantities of interest with the Monte Carlo method, for example. Depending on whether the samples are drawn with or without replacement and whether the set of drawn samples has smaller or the same cardinality, the arising estimators are termed the randomization test, − the jackknife, , or the bootstrap method. − For a comprehensive comparison of the sampling schemes we refer the reader to ref (see also Figure ). 5. Open in a new tab Comparison of resampling schemes. Reproduced with permission from ref . Copyright 1999 Lawrence Erlbaum Associates. The simplicity of resampling methods has led to enormous popularity across the fields. On the other hand, it is intuitively clear that it heavily depends on how representative the samples in X are compared to X . Still, for certain cases one can show that the empirical cumulative distribution function of X converges in a reasonable sense to the one of X as M → ∞ , i.e., if we enlarge the size of our sample set to infinity. For an in-depth discussion and examples we refer to the respective references. 3.4.3. Quadrature Methods: Quasi-Monte Carlo and Sparse Grids Quadrature methods can be considered as a modification of the Monte Carlo estimator where the random samples are replaced by deterministic ones. To this end, one approximates E [ f ] = ∫ Ω f ( ω ) ρ ( ω ) d ω ≈ ∑ i = 1 N w i f ( ξ i ) ≕ Q [ f ] 40 where the ξ i , i = 1, ..., N , are deterministic quadrature nodes and the w i , i = 1, ..., N , are quadrature weights . As a straightforward modification of the Monte Carlo approach, quasi-Monte Carlo methods , choose the quadrature nodes ξ i , i = 1, ..., N , deterministically and the quadrature nodes as w i = 1 N , i = 1, ..., N . The intuition behind this is that the Monte Carlo estimator can potentially become more accurate if the quadrature nodes are of low discrepancy (i.e., evenly distributed) but still chosen at reasonably sophisticated places. Indeed, classical quasi-Monte Carlo methods provide a convergence rate of almost 1, i.e., O ( N − 1 + δ ) for any δ > 0, if first-order mixed regularity of the integrand is available. , If the integrand provides even more regularity, even higher orders of convergence can be achieved when using modern quasi-Monte Carlo methods. − Due to their deterministic choice of the quadrature nodes, quasi-Monte Carlo methods lead to a deterministic approximation of the integral, rather than a statistical estimator as the Monte Carlo Method. Traditional deterministic quadrature rules leverage the freedom of also choosing the quadrature weights to obtain improved deterministic approximations of the integral. For single parameter problems, i.e., Ω ⊂ R , it is textbook knowledge that choosing quadrature nodes and quadrature weights according to Gauss quadrature rules yields the best possible approximation of the integral in eq . However, the situation is different when the domain of integration Ω is higher dimensional. Then, the naive combination of one-dimensional quadrature formulas by means of a tensor-product approach leads to methods suffering from the curse of dimensionality . To overcome the curse of dimensionality, the idea of sparse grid methods , is that a more sophisticated combination of one-dimensional quadrature formulas can break the curse of dimensionality, if the integrand is sufficiently smooth. More precisely, given a sequence of nested one-dimensional quadratures on Ω = I ⊂ R , i.e., Q l [ f ] = ∑ i = 1 N l w l , i f ( ξ l , i ) , l = 0 , 1 , 2 , ... 41 with { ξ l , 1 , ... , ξ l , N l } ⊂ { ξ 1 , l + 1 , ... , ξ N l + 1 , l + 1 } , l = 0, 1, 2, ..., we introduce the differences Δ l [ f ] = Q l [ f ] − Q l − 1 [ f ] , l = 0 , 1 , 2 , ... 42 with the convention that Q –1 [ f ] = 0. We note that Δ l [ f ] is similar to Q l [ f ] in the sense that Δ l [ f ] is a quadrature rule with the same nodes but with different weights. Whereas the tensor product quadrature rule for a d -dimensional domain Ω = I × ··· × I ︸ d times ⊂ R d is given by application of Q l [·] in each variable, i.e., Q l ( d , TP ) [ f ] = ( Q l ⊗ ··· ⊗ Q l ) ︸ d times [ f ] = ∑ i 1 = 1 N l ··· ∑ i d = 1 N l w i 1 ··· w i d f ( ξ i 1 , ... , ξ i d ) 43 we can rearrange terms to rewrite it as Q l ( d , TP ) [ f ] = ∑ ∥ i ∥ ∞ ≤ l ( Δ i 1 ⊗ · · · ⊗ Δ i d ) [ f ] 44 where i = ( i 1 , ... , i d ) ⊂ N 0 d and ∥ i ∥ ∞ = max k =1,..., d | i k |. The sparse quadrature rule is then given as , Q l ( d , SG ) [ f ] = ∑ ∥ i ∥ 1 ≤ l + d − 1 ( Δ i 1 ⊗ ··· ⊗ Δ i d ) [ f ] 45 where i = ( i 1 , ... , i d ) ⊂ N 0 d and ∥ i ∥ 1 = ∑ k =1,..., d | i k |. Using the combination technique, the latter expression can be simplified to , Q l ( d , SG ) [ f ] = ∑ ∥ i ∥ 1 ≤ l + d − 1 ( − 1 ) l + d − ∥ i ∥ 1 − 1 ( d − 1 ∥ i ∥ 1 − l ) · ( Q i 1 ⊗ ··· ⊗ Q i d ) [ f ] 46 One can show that this quadrature rule satisfies an error bound of O ( N − r ( log N ) ( d − 1 ) ( r + 1 ) ) , where N is the number of required evaluations of f if f provides r th-order mixed regularity. This rate of convergence can be improved, if one allows for different quadrature rules in the different variable directions which can be accomplished using a priori information , or adaptively. We refer the interested reader to the relevant literature ,, for more details and remarks and to Figure for an illustration. 6. Open in a new tab Illustration of different quadrature methods for 25 points in Ω = [0, 1] 2 . Top left (black): Random Monte Carlo points. Top right (red): Quasi-Monte Carlo points from a 2,3 Halton series. Bottom left (blue): Gaussian quadrature points approximating a uniform distribution. Bottom right (green): Clenshaw–Curtis quadrature points approximating a uniform distribution. Fortunately, sparse grid techniques are available in most software packages for uncertainty quantification nowadays, see section , and the average user does not need to dive into the implementation details of sparse grids. 3.4.4. Spectral Approaches: Stochastic Collocation and Polynomial Chaos Spectral approaches , belong to the class of surrogate models and are based on polynomial model approximations. They can be expected to work well if the model changes sufficiently smooth in the parameters. More precisely, for a model in d variables represented by a sufficiently smooth function f : [ 0 , 1 ] d → R , one approximates f ( x ) ≈ ∑ i ∈ I c i Ψ i ( x ) 47a Ψ i ( x ) = Ψ i 1 ( x 1 ) ··· Ψ i d ( x d ) 47b where I ⊂ N 0 d is a suitable (multi-index) set of polynomial degrees, i = ( i 1 , ..., i d ) is an element thereof, c i are suitable coefficients, and { Ψ l } l is a family of polynomials in one variable. The main challenge in using this approach is choosing a family of polynomials Ψ l and polynomial degrees (i.e., the index set I ) and then to determine the coefficients c i . The family of polynomials directly affects the way in which the coefficients can be determined, the approximation quality, and other properties. Similar to tensor product quadrature one could use different types of polynomials in each variable, but we keep this simple presentation for the sake of illustration. We typically speak of stochastic collocation methods if we choose to interpolate f at a series of predetermined points in the parameter space. A classical choice to do so is to choose a set of pairwise distinct interpolation points (ξ 0 , ..., ξ N ) ⊂ [0, 1] and consider associated one-dimensional Lagrange polynomials Ψ l ( t ) = ∏ i = 0 i ≠ l N t − ξ i ξ l − ξ i , l = 0 , ... , N 48 Lagrange polynomials (visualized in Figure , top panel) have the convenient property that Ψ l ( ξ k ) = { 1 , l = k 0 , l ≠ k k , l = 0 , 1 , ... , N 49 which allows an approximation of the type shown in eq 47 to be obtained by setting f ( x ) ≈ ∑ ∥ i ∥ ∞ ≤ N f ( ξ i 1 , ... , ξ i d ) Ψ i 1 ( x 1 ) ··· Ψ i d ( x d ) 50 Similar to tensor product quadrature, this naive construction will suffer from the curse of dimensionality in the evaluation points. If the function to approximate is sufficiently smooth, sparse grid techniques can be used in complete analogy to the case of quadrature methods. ,,, 7. Open in a new tab Illustration of three Lagrange polynomials of degree 2 for interpolation points ξ 0 = 0, ξ 1 = 0.5, ξ 2 = 1, denoted by black lines and dots (top panel) and the first six Legendre polynomials (bottom panel). In contrast to stochastic collocation methods, polynomial chaos expansions , − classically choose orthogonal polynomials . Assuming that f ( x ) = f ( x 1 , ..., x d ) has finite variance and depends on i.i.d. random variables x 1 , ..., x d , each of them following the probability distribution ρ 1 , we introduce the joint probability distribution ρ( x 1 , ..., x d ) = ρ 1 ( x 1 )···ρ d ( x d ) and an inner product and norm : ⟨ f , g ⟩ ρ = ∫ Ω f ( x ) g ( x ) ρ ( x ) d x 51 ∥ f ∥ ρ = ⟨ f , f ⟩ ρ = ∫ Ω f ( x ) 2 ρ ( x ) d x 52 Orthogonal polynomials are families of polynomials { Ψ l } l = 0 ∞ which satisfy ⟨ Ψ l , Ψ k ⟩ ρ = { 1 , l = k 0 , l ≠ k k , l = 0 , 1 , 2 , ... 53 and each Ψ l is a polynomial of degree l . Typical examples for orthogonal polynomials are Hermite polynomials (with ρ 1 being the standard normal distribution on R ) and Legendre polynomials (with ρ 1 being the uniform distribution on [−1, 1]). Legendre polynomials are visualized in the bottom panel of Figure . Orthogonal polynomials exist for all probability distributions, but their computation is nontrivial except for very simple cases. Once the orthogonal polynomials are available, an approximation of the form eq 47 is given by f ( x ) ≈ ∑ ∥ i ∥ ∞ ≤ N ⟨ f , Ψ i 1 ⊗ ··· ⊗ Ψ i d ⟩ ρ Ψ i 1 ( x 1 ) ··· Ψ i d ( x d ) 54 We note that evaluating ⟨ f , Ψ i 1 ⊗ ··· ⊗ Ψ i d ⟩ ρ requires the computation of an integral, which can for example be done with quadrature methods (cf. section ) or a least-squares fit. Similar to stochastic collocation, our presentation can be modified to allow for different probability distributions in each random variable and the assumption of stochastic independence can be lifted. Moreover, the polynomial chaos expansions as introduced here will suffer from the curse of dimensionality which can be tacklet with sparse polynomial chaos expansions. , Once the approximation polynomial is available, it can be used to compute statistical quantities of interest such as the mean, variance, or sensitivity measures, for example. For a discussion on the advantages and disadvantages of the approaches, we refer the reader to ref . 3.4.5. Multilevel and Multifidelity Methods The idea of multilevel , and multifidelity methods is to exploit (approximate) multilevel hierarchies to construct a variance reduction technique and shift a large amount of sampling costs from accurate but expensive models to less accurate but cheap models. The origin of these methods is in the uncertainty quantification of partial differential and integral equations, where different levels of model accuracy and computational cost can easily be constructed by nested hierarchies of meshes and/or approximation spaces. To this end, the classical multilevel Monte Carlo method considers a model f which is approximated by a sequence of approximations f 0 , f 1 , ..., f L with increasing accuracy. In this way, the mean of the most accurate approximation E [ f L ] can be approximated by a telescopic sum: E [ f ] ≈ E [ f L ] = E [ f 0 ] + ∑ l = 1 L E [ f l − f l − 1 ] ≈ E N 0 [ f 0 ] + ∑ l = 1 L E N l [ f l − f l − 1 ] = E L ML [ f ] 55 where the N l are the number of samples for the various Monte Carlo estimators. The latter approximation is called the multilevel Monte Carlo estimator , visualized in Figure . Now we write f –1 = 0 and let C l be the computational cost for the evaluation of one sample of f l − f l − 1 . Then the overall cost of the multilevel Monte Carlo estimator is C = ∑ l = 0 L N l C l , and its variance is V [ E L ML [ f ] ] = ∑ l = 0 L V [ f l − f l − 1 ] N l 56 The estimator will be accurate if this variance is small. One can show that for any desired upper bound for this variance, choosing N l ∼ V [ f l − f l − 1 ] / C l minimizes the overall computational cost of the estimator. Thus, if V [ f l − f l − 1 ] decays sufficiently quickly with l , then we obtain small V [ E L ML [ f ] ] (i.e., an accurate estimator) with small sample numbers N l for computationally expensive models (large l ) and large sample numbers N l for computationally cheap models (small l ). Compared to the conventional (single-level) Monte Carlo method (cf. section ), where all samples need to be drawn on the level with the highest accuracy (here: L ) this can yield significant savings in compute time. Considerations with the same outcome can be made when replacing the single-level Monte Carlo estimators in the multilevel estimator by quadrature methods such as quasi-Monte Carlo, sparse grids, stochastic collocation, etc. In fact, the multilevel Monte Carlo method can itself be interpreted as a sparse grid method, giving rise to the family of multi-index Monte Carlo methods. 8. Open in a new tab Illustration of multilevel Monte Carlo methods. The multifidelity Monte Carlo method generalizes the multilevel Monte Carlo method from a hierarchy of approximations to a single model to a hierarchy of different models which are correlated to a high-fidelity model of high accuracy and high computational cost. To this end, given a high-fidelity model f hi and low-fidelity models f 1 , ..., f L , the Monte Carlo estimator given in eq is based on sample numbers N 0 ≤ ··· ≤ N L , i.i.d. samples ξ 1 , ... , ξ N l , and single-level estimators E N 0 [ f hi ] = 1 N 0 ∑ l = 0 N 0 f hi ( ξ l ) 57 E N l [ f l ] = 1 N l ∑ l = 0 N l f l ( ξ l ) 58 E N l − 1 [ f l ] = 1 N l − 1 ∑ l = 0 N l − 1 f l ( ξ l ) 59 and reads E [ f ] ≈ E N 0 [ f hi ] + ∑ l = 1 L α k ( E N l [ f l ] − E N l − 1 [ f l ] ) ≕ E L MF [ f ] 60 The sample numbers N 0 , ..., N L and the control variates α 1 , ..., α L ∈ R are determined from the Pearson correlation coefficients between the models, the variances of the models, and the computational cost of the various model evaluations (see ref ). 3.5. Surrogate Models Approximations to high-fidelity models which are cheap to evaluate are refereed to as surrogate models . These surrogate models play an important role in reducing the computational burden of sampling based methods, particularly when used as low-fidelity models in multifidelity methods. 3.5.1. Surrogate Modeling and Machine Learning From an abstract point of view, surrogate models are a mathematical model, the surrogate method , whose hyper-parameters are fitted to reference data, called the training set . Given suitable surrogate methods and training sets, surrogate methods often generalize well to other data. Clearly, stochastic collocation and polynomial chaos approaches (cf. section ) as well as any machine learning method which infers models from training data can act as surrogate method. In the language of section , we may consider (ξ, f (ξ)) as the hypothetically exact training data in form of input–output pairs. In practice, we need to replace ξ by suitable approximations ξ chem , introducing a representation error, such that the surrogate method is a mapping ξ chem → f comp (ξ chem , η, h ), where the model parameters η depend on the training data. Thus, the error of the surrogate method can be interpreted as a combination of a level of theory error and a numerics error. The error of surrogate methods, and particularly machine learning methods, themselves is an active field of research, with some methods providing exact mathematical results on the error behavior and other methods whose approximation properties are still not fully understood (see refs , , and for comprehensive reviews). 3.5.2. Gaussian Processes A particularly popular instance of surrogate models with connections to Bayesian inverse problems are Gaussian processes , illustrated in Figure . A Gaussian process is a collection of random variables, where all finite subsets thereof have a joint Gaussian distribution. We write f ( x , ω), where ω is a random event and x can be considered as an index to the random variables. It is well-known that a Gaussian process is fully specified by its mean m ( x ) = E [ f ] ( x ) and its covariance function k ( x , x ′) = Cov[ f ]( x , x ′) (cf. section for their definition). Most importantly, the covariance function does not need to be a Gaussian in the sense of eq and its multivariate analog. In fact, it is sufficient to satisfy symmetry, i.e., k ( x , x ′) = k ( x ′, x ), and positive semidefiniteness, i.e., it holds ∑ i , j = 1 n k ( x i , x j ) c i c j ≥ 0 61 for all n ∈ N , x 1 , ..., x n , and c 1 , ..., c n . We write f ( x , ω) ∼ G ( m ( x ) , k ( x , x ′ ) ) or f ( x , ω) ∼ N ( m ( x ) , k ( x , x ′ ) ) and note that f ( x , ω) has a natural representation in terms of a Karhunen–Loève expansion: f ( x , ω ) = m ( x ) + ∑ l = 1 ∞ λ l ψ l ( x ) y l 62 where ( λ l , ψ l ) , l = 1 , 2 , ... , are the eigenpairs of the integral operator ( C g ) ( x ) = ∫ k ( x , x ′ ) g ( x ) d x ′ 63 and y l ( ω ) = 1 λ l ∫ [ f ( x , ω ) − E [ f ] ( x ) ] ψ l ( x ) d x 64 are random variables. We note that the latter ones are usually not available in practice. 9. Open in a new tab Visualization of Gaussian process regression, using a Gaussian kernel function with a mean of zero. The top panel shows samples drawn from the prior distribution. The central panel shows samples drawn from the posterior distribution with training points indicated by black dots. The bottom panel shows the predicted mean, with 95% confidence intervals indicated in gray. Gaussian processes offer a natural capability to incorporate knowledge from measurements by using the Bayesian framework. Assuming as prior knowledge that our random variable f ( x , ω) is a Gaussian process and that we have measurements y i = f ( x i , ω ) + ε i ( ω ) , i = 1 , ... , n 65 at certain locations x 1 , ..., x n and subject to independently distributed Gaussian noise ε i ∼ N ( 0 , σ 2 ) , the posterior distribution of f is explicitly given as a Gaussian process G ( m δ ( x ) , k δ ( x , x ′ ) ) , with m δ ( x ) = k̂ ( x ) T ( K̂ + σ 2 I ) − 1 ( y − m ) 66 k δ ( x , x ′ ) = k ( x , x ′ ) − k̂ ( x ) T ( K̂ + σ 2 I ) − 1 k̂ ( x ) 67 where m = [ m ( x 1 ), ..., m ( x n )] T and k̂ ( x ) = [ k ( x 1 , x ) ⋮ k ( x n , x ) ] 68 K̂ = [ k ( x 1 , x 1 ) ⋯ k ( x 1 , x n ) ⋮ ⋱ ⋮ k ( x n , x 1 ) ⋯ k ( x n , x n ) ] 69 This approach is also known as Gaussian process prediction or interpolation , kriging , , or kernel ridge regression . , See ref for a thorough review in the context of chemistry. 3.6. Uncertainty Validation Assessing the reliability of a prediction’s uncertainty is as important as assessing the reliability of the prediction itself. If a model’s estimated uncertainty is too small, one is at risk to overly trust the model’s predictions, resulting in running fewer (computer) experiments than necessary, thereby missing important examples. If a model’s estimated uncertainty is too large, on the other hand, one is at risk to overly distrust the model’s predictions, resulting in running more experiments than necessary, thereby wasting resources. Both situations are to be avoided by research facilities which have to act in the interest of both knowledge and economics, which stresses the importance of uncertainty validation. The uncertainty of a prediction usually refers to a confidence interval of a probability function for a given confidence level γ ∈ [0, 1]. The confidence level is also referred to as nominal coverage probability as it reflects the expected (nominal) relative frequency (probability) at which the true value lies within (is covered by) the associated confidence interval. The actual coverage probability refers to the observed (actual) relative frequency at which a benchmark lies within the associated confidence interval, which requires an out-of-sample/test set. Nominal and actual coverage probability match if, for an infinitely large test set, lim M → ∞ M − 1 ∑ i = 1 M 1 ( y i ∈ I γ , i ) = γ 70 The indicator function 1 ( y i ∈ I γ , i ) equals 1 if the benchmark value of the i th out-of-sample point, y i , lies within the confidence interval I γ , i ; if it lies outside the interval, the indicator function equals zero: 1 ( y i ∈ I γ , i ) = { 1 , y i ∈ I γ , i 0 , y i ∉ I γ , i 71 It is customary to report uncertainties (i.e., confidence intervals) for a given confidence level. In thermochemistry, for instance, γ = 0.95 is a convention. This corresponds to ca. 2 standard deviations if the underlying error distribution is normal. However, a single confidence level is insufficient in terms of uncertainty validation as the information that the actual coverage probability matches or does not match a given confidence level cannot be transferred to another confidence level. In the forecasting community, one speaks of a well-calibrated model if the actual coverage probability matches any confidence level, lim M → ∞ M − 1 ∑ i = 1 M 1 ( y i ∈ I γ , i ) = γ , ∀ γ ∈ [ 0 , 1 ] 72 For instance, if a normal distribution of the true value with some mean value and variance is assumed, nominal and actual coverage probabilities will only match for any value of γ if the true underlying probability distribution is a normal distribution with the same mean value and the same variance. Here, the term “uncertainty validation” refers specifically to the validation of uncertainty estimates themselves, and should not be confused with model validation as used in verification and validation (V&V) or VVUQ frameworks (cf. section ). 3.7. Dimension Reduction As noted in the sections on quadrature methods ( section ) and spectral approaches ( section ), these approaches suffer from the curse of dimensionality. In other words, the number of required samples to obtain a desired accuracy grows exponentially with the number of parameters. This issue applies to surrogate models as reviewed in section as well. Under certain circumstances, the curse of dimensionality can be mitigated using specialized methods, as for example discussed in sections and . However, quite often it turns out that the behavior of a model is such that it can be (approximately) described by significantly less parameters than given. When a model is (approximately) reparametrized to depend on fewer parameters, this is referred to as dimension reduction . Significant computational gains can be made when the dimensionality of a model is reduced before applying any of the above-mentioned methods. Dimension reduction is an active research area, particularly in machine learning, whose review is beyond the scope of this review. We content ourselves by mentioning two of the most popular approaches are based on principal component analysis and global sensitivity analysis. Both approaches provide us with a linear decomposition where parameters with small contribution can be identified and discarded. For further details on dimension reduction approaches, we refer the reader to refs − . 3.8. Software for Uncertainty Quantification Along with the emergence of UQ in the research fields outlined above and beyond, various software packages have been developed in the context of UQ. Available software presently offers a broad range of tools, covering simple basic statistical analyses as well as advanced approaches to specific applications. Below, we provide a nonexhaustive list of UQ software and libraries that are potentially applicable in the field of in silico chemistry. The tools are listed alphabetically, including a short summary highlighting their main features. ChaosPy : Open-source python library providing specific tools for polynomial chaos expansions and advanced MC sampling. A focus is put on its compatibility with other (python) libraries for UQ. DAKOTA : Large open-source C++ project with broad functionality between optimization (both gradient and nongradient methods), UQ (sampling and reliability as well as stochastic expansion methods), parameter analysis using nonlinear least-squares optimization, and sensitivity analysis (design of experiments and parameter study methods). One of its main advantages is the possible combination of the available methods in a user-defined fashion. EasyVVUQ : , Open-source python library focusing on UQ for MD simulations, providing, for example, polynomial chaos expansions, stochastic collocation, and sensitivity analysis functionality. Part of the SEAVEA toolkit. MUQ Open-source python/C++ library for inverse and forward UQ, featuring various MCMC-based sampling techniques, regression and prior modeling via Gaussian processes, as well as polynomial chaos expansions and nonlinear optimization. NIST Uncertainty Machine : Open-source web-based application that can be operated to assess the measurement uncertainty of an output quantity by means of Gaussian error propagation and MC sampling. The output quantity is required to be a real, explicitly known function of any number of input quantities which are modeled as random variables. OpenTURNS : Open-source python/C++ library for multivariate probabilistic data modeling, including Bayesian calibration and sensitivity analysis. physical_validation : Open-source python library that offers a wide range of tests that can be applied to MD and MC simulations a posterior to validate their physical correctness (e.g, in terms of chosen cutoffs, thermostat, simulation time step etc.) Moreover, functionality for checking code-correctness of MD software is offered, initially developed for GROMACS but designed to be extended to other software. PyApprox : Open-source python/C++ library for high-dimensional approximation and UQ. Special features include tools for building surrogate models, sparse grid interpolation, multifidelity approximations, Bayesian inference methods, and more. SALib : Open-source python library for global sensitivity analysis of general computational models. SEAVEA toolkit : Open-source collection of tools for VVUQ with applications on large-scale computing infrastructures ( www.seaveatk.org ). Includes, for example, EasyVVUQ and is a successor of the VECMA toolkit. SMCPy : Open-source python library developed to enable UQ via parallel, sequential MC sampling. As its key features, it offers an alternative to the MCMC method when treating Bayesian inference problems and an unbiased marginal likelihood estimation in terms of Bayesian model selection. Sparse Grids Matlab Kit Matlab implementation of sparse grids, designed to approximate functions of high order. This tool is applicable for surrogate model-based UQ. Review and comparison to other software packages. UncertainSCI : Open-source python tool for UQ in the field of biomedical simulations, focusing on the propagation of input uncertainties on model outputs by polynomial chaos expansions. UM-Bridge : Open-source python/C++ toolkit designed for combining any kinds of computational models with various statistical/optimization methods of UQ. This is achieved by containerization of software, enabling the coupling of models and UQ methods from different programming languages and computational environments, with the intention to provide a unified interface for UQ methods from all fields of science. UQLab : Open-source Matlab framework for UQ in applied science and engineering, providing software that can be combined in a user-defined and modular fashion. Supported UQ techniques include among others: Bayesian model calibration, sensitivity analysis, support vector machines, polynomial chaos expansions. UQTk : Open-source python/C++ toolkit comprising several libraries for UQ on numerical models, focusing on polynomial chaos expansions, uncertainty propagation, sensitivity analysis and Bayesian inference. Uranie : Open-source C++ software with python support, designed to process and handle large data when performing UQ. Specifically, it supports design of experiments and surrogate model generation, sensitivity analysis, calibration and parametric optimization as well as dedicated visualization tools. WESTPA : Weighted Ensemble Simulation Tool , and tutorial for rare event sampling. 4. Applications of Uncertainty Quantification in In Silico Chemistry 4.1. Energy Uncertainties 4.1.1. Wavefunction-Based Quantum Chemistry Wavefunction-based quantum chemistry provides methods for approximately solving the nonrelativistic many-electron Schrödinger equation within the Born–Oppenheimer approximation (i.e., for a fixed molecular structure with nuclear coordinates R 1 , R 2 , ...): Ĥ el Ψ n ( x 1 , ... , x N ) = E n el Ψ n ( x 1 , ... , x N ) 73 with the electronic Hamiltonian Ĥ el , the electronic wavefunction Ψ n , and the combined spatial and spin coordinates x i = ( r i , s i ) of the electrons. Here, we will mainly consider the ground-state electronic energy E 0 ( R 1 , R 2 , ...) as the quantity of interest. Single-reference wavefunction-based quantum-chemical methods, such as configuration interaction (CI) and coupled-cluster (CC) methods, generally use the Hartree–Fock approximation as their starting point, in which the electronic wavefunction is approximated by a single Slater determinant Ψ SD , which is formed from the occupied Hartree–Fock orbitals ϕ i ( r ). The Hartree–Fock orbitals are in turn approximated by expanding them in suitable atom-centered basis functions χ μ ( r ): ϕ i ( r ) = ∑ μ = 1 N bas c μ ( i ) χ μ ( r ) 74 Besides the occupied orbitals, the solution of the Hartree–Fock equations also provides a set of unoccupied (virtual) orbitals, which can be used to construct additional Slater determinants Φ ij ··· by replacing the occupied orbitals i , j , ... in the Hartree–Fock determinant by virtual orbitals a , b , .... The many-electron wavefunction is then expanded in a suitable basis of these Slater determinants as Ψ el ( x 1 , ... , x N ) = ∑ n C n Φ n ( x 1 , ... , x N ) 75 where the expansion is generally terminated based on the excitation level (i.e., the maximum number of occupied orbitals that are replaced by virtual orbitals). Multireference methods − perform the optimization of the molecular orbitals simultaneously with the determination of the expansion coefficients of the Slater determinants, and usually define their many-electron wavefunction in terms of an active space of orbitals. For further details on specific quantum-chemical methods, we refer the reader to ref . We note that explicitly correlated methods, which extend the ansatz of eq by including terms that depend on the interelectronic distance | r i – r j |, are nowadays widely used. The following discussion also applies to such explicitly correlated methods. When considering errors in wavefunction-based quantum chemistry, one usually considers the exact nonrelativistic ground-state energy of eq as the reference. Therefore, uncertainties in the calculated ground-state energy will be solely due to the approximation introduced by the quantum-chemical methods. Generally, there are two main sources of errors in approximate wavefunction-based quantum-chemical methods: First, the incompleteness of the basis set that is used for expanding the one-electron orbitals (see eq ). Second, the incompleteness of the basis of Slater determinants that are used for expanding the many-electron wavefunction (see eq ). In principle, both of these approximations can be controlled and systematically improved (see Figure ). The exact ground-state energy of the nonrelativistic Schrödinger equation can be obtained with a full-CI calculation in an infinitely large one-electron basis set. Here, it is important to note that neither full-CI calculations in a finite one-electron basis set or calculations using a truncated expansion in Slater determinants in a complete one-electron basis set provide the exact solution, but both limits have to be approached simultaneously. 10. Open in a new tab Errors in quantum-chemical calculations. Adapted from ref . However, this exact solution (or an energy that is within the required numerical accuracy) is hardly achievable in practice for molecules, except for the smallest systems (see, e.g., refs − ). Nevertheless, numerous benchmarking studies have demonstrated the systematic convergence of total energies as well as atomization energies and reaction energies toward experimental reference values. For a particularly systematic example of such a study, we refer the reader to Chapter 15 of ref . Error Estimation by Extrapolation For quantum-chemical calculations using accurate wavefunction-based methods, one can try to systematically approach the exact solution in calculations for individual systems. One common approach to this end is based on the use of extrapolation schemes, in which one performs calculations with increasing basis set sizes and uses these to extrapolate to the basis set limit. This can be further combined with an extrapolation of the correlation contributions by performing calculations that systematically increase the number of included Slater determinants (e.g., by increasing the excitation level in a CI expansion or a CC ansatz). For a review of methods for highly accurate quantum-chemical calculations, including a comprehensive overview of extrapolation schemes and composite methods, we refer the reader to ref . Such extrapolation schemes also provide a basis for estimating the remaining error of the extrapolated energies. For instance, one can use the difference between the extrapolated energies and those of the calculations with the largest basis set and correlation level. Alternatively, the error can be estimated by comparing the results of different extrapolations. The details will depend on the specific composite and/or extrapolation scheme. An example of such an approach can be found in ref , where the basis-set error of CCSD(T) calculations is estimated by comparing the predictions of a basis-set extrapolation scheme to those obtained with the largest basis set. A general scheme using random walks, that can be used to estimate the uncertainty of energies obtained by extrapolating to the complete basis-set limit, has recently been put forward by Lang et al. To also include uncertainties arising from electron correlation, Schuurman et al. performed a focal-point analysis to obtain accurate estimates of the heats of formation for NCO, HNCO, HOCN, HCNO, and HONC. They used the convergence behavior and internal consistency of their results to estimate that the uncertainty in their predictions amounts to less than 0.2 kcal mol –1 . Also using a focal point analysis, Hajgató et al. calculated the singlet–triplet gaps in polyacenes, and estimated the uncertainty in their predictions by comparing different extrapolations. An elegant and detailed protocol for estimating uncertainties within a composite scheme has been put forward by Bakowies, , who developed a sophisticated model for the different relevant sources of errors. A very instructive example showcasing systematic error estimation for highly accurate quantum-chemical calculations was recently presented by Jeziorski and co-workers, who calculated the three-body potential of neon. They carefully estimated the error in each term by considering the difference between the two largest employed basis sets. Another example are highly accurate calculations of the potential energy curve of the chromium dimer by Larsson et al. They combine an estimate of the correlation error with a small basis set with the estimated uncertainty of their basis set extrapolation. While these examples show that for wavefunction-based quantum chemistry, the possibility to systematically approach the exact solution can be used to estimate uncertainties, we are not aware of any recommendations or accepted protocols for reporting estimated errors in such calculations. Multireference Diagnostics We note that the approaches discussed above are mostly restricted to single-reference methods. However, such approaches fail in cases that are dominated by static correlation (i.e., in which the dominating contribution to the exact electronic wavefunction is not from a single determinant, but from several determinants with similar weights). In such cases (e.g., for many open-shell transition metals), multireference methods − are required. While for single-reference methods, a clear hierarchy can be defined based on the excitation level, this is a nontrivial problem when selecting the active space in multireference methods (for attempts to overcome this limitation, see refs − ), which often precludes the application of systematic extrapolation schemes for error estimation. There have been many attempts to devise diagnostic measures that allow one to detect whether single-reference methods are applicable for a certain system (for an overview, see refs and ). Within the Weizmann W4 composite method, , the percentage of the total atomization energy (TAE) due to quadruple and quintuple excitations, % TAE [ T̂ 4 + T̂ 5 ] = TAE [ CCSDTQ 5 ] − TAE [ CCSDT ] TAE [ CCSDTQ 5 ] × 100 % 76 or the %TAE due to perturbative triples excitations, % TAE [ ( T ) ] = TAE [ CCSD ( T ) ] − TAE [ CCSD ] TAE [ CCSD ( T ) ] × 100 % 77 where TAE[X] is the TAE calculated with method X, was used as an indicator of multireference character and considered as an estimate of the uncertainty in the calculated total atomization energies. This is in the same spirit as similar uncertainty estimators for extrapolation schemes and composite methods discussed above. In ref , it was observed that %TAE[ T̂ 4 + T̂ 5 ] and %TAE[(T)] show a very good correlation, and consequently, it has been suggested to use %TAE[(T)] as a diagnostic of multireference character and as an a priori estimate for the contribution of higher excitation levels. Many different measures have been put forward as multireference diagnostics. For coupled-cluster methods, the use of the norm of the singles amplitudes, T 1 , is widespread, and related diagnostics based on the coupled-cluster amplitudes have also been put forward. − For coupled-cluster methods, it was recently proposed that the extent of nonhermiticity of the one-electron density matrix can be used to estimate the error in the correlation energy. Another class of multireference diagnostics is based on occupation numbers from finite-temperature DFT calculations or on natural occupation numbers from MP2 or CASSCF calculations. − Such indicators, which are usually cheeper to calculate than %TAE[(T)], have been discussed as predictors of the latter (i.e., as quantitative uncertainty estimate), or at least as a qualitative indicator whether single-reference methods are applicable. However, the correlation between different multireference diagnostics is generally rather poor. ,, Machine learning approaches, which combine different easily computable multireference diagnostics as well as geometric descriptors, have been shown to be promising in predicting multireference character as measured by %TAE[(T)] for transition metal complexes. − This has been applied to exclude problematic (i.e., high uncertainty) cases in virtual high-throughput screening. Accuracy Predictor for MP2 For Møller–Plesset second-order perturbation (MP2) theory applied to noncovalently bonded complexes, Vuckovic et al. have developed an accuracy predictor by employing the adiabatic connection. The adiabatic connection considers the total energy as a function of a scaled electron–electron interaction, while keeping the electron density constant. The authors point out that MP2 theory approximates this adiabatic connection curve by a straight line, and propose that the deviation of the exact adiabatic connection curve from linear behavior can be used to estimate the accuracy of MP2 interaction energies. They show that for weakly interacting complexes, the exact adiabatic connection curve can be approximated by interpolating between the weakly and strongly interaction limits and by using an approximation for the latter. The resulting predictor can be calculated with no additional effort and correlates well with the errors in the MP2 interaction energies in comparison to high-level reference data. Rigorous Error Bounds Since the electronic Schrödinger equation is a (Hermitian) eigenvalue problem, it is possible to apply error bounds on approximate eigenvalues to quantum-chemical methods for its approximate solution. In particular, the Kato–Temple theorem (see, e.g., ref ) states that for an approximate normalized eigenvector Ψ̃ and the corresponding approximate eigenvalue Ẽ = ⟨Ψ̃| H̃ |Ψ̃⟩, the distance to the closest exact eigenvalue E is bounded by | Ẽ − E | ≤ ∥ r ∥ 2 δ 78 where δ is the distance from E to the remaining eigenvalues (i.e., | E – E j | ≥ δ for all other eigenvalues E j ) and r is the residual, given by r = Ĥ Ψ̃ − Ẽ Ψ̃ 79 Such error bounds have been applied for the analysis of CI as well as CC methods. For a detailed discussion, we refer the reader to ref and other works mentioned therein. While the formal analysis of the errors is valuable by itself, such rigorous error bound for energies from wavefunction based quantum-chemical methods have to the best of our knowledge not been used to calculate uncertainties for individual calculations, which we believe is for two reasons. First, the calculation of the norm of the residual, ∥ r ∥ 2 , is difficult in practice, as it has to be represented outside of the (finite) Hilbert space that is used in the calculations. In particular for the many-electron basis, this presents a insurmountable obstacle. Second, error bounds for the total energy are rarely relevant in practice, as chemistry is concerned with relative energies. In computational chemistry, the calculation of energy differences always profits from systematic error cancelation, even with the most accurate methods. Therefore, uncertainties in relative energies that can be derived from error bounds for total energies are rarely useful. 4.1.2. Density Functional Theory Kohn–Sham density functional theory (KS-DFT) is the workhorse of quantum chemistry, mostly because of its favorable cost–accuracy ratio. − It approximates the total electronic energy as a functional of the KS orbitals { ϕ i ( r )}, yielding the electron density ρ ( r ) = ∑ occ | ϕ i ( r ) | 2 : E [ { ϕ i } ] = T s [ { ϕ i } ] + V nuc [ ρ ] + J [ ρ ] + E xc [ ρ ] + E NN 80 where T s [{ ϕ i }] is the noninteracting kinetic energy, V nuc [ρ] is the electron–nuclear attraction energy, J [ρ] is the classical Coulomb repulsion of the electrons, E NN is the nuclear repulsion energy, and E xc [ρ] is the exchange–correlation (xc) energy. In practice, the latter is treated using approximate functionals of the electron density (and possibly additional ingredients, such as occupied and/or virtual Kohn–Sham orbitals). Minimization of the above energy functional leads to the one-electron KS equations, from which the KS orbitals are determined. When considering the exact solution of the nonrelativistic Schrödinger equation for a fixed molecular structure as benchmark, uncertainties in the energies calculated with DFT arise from the approximate xc functional, the incompleteness of the basis set used for the expansion of the KS orbitals as well as, for instance, numerical integration, convergence thresholds, the use of pseudopotentials to represent core electrons, and Brillouin zone sampling in the case of periodic boundary conditions. The classification of these error sources as level of theory uncertainties and as numerics uncertainties is not always clear and varies in different contexts (see section ) Rigorous Bounds for Numerical Errors in DFT While in principle numerical errors in DFT calculations are well-controllable by systematically investigating the convergence with respect to the relevant parameters, such as the size of the basis set or convergence thresholds, the choice of these parameters is not trivial in practice. A large-scale comparison of periodic boundary condition DFT calculations for solids revealed striking differences between the implementations in different DFT codes (which each come with different numerical approximations and default choices). Even though these differences could be reconciled, this highlights the importance of systematically estimating and controlling numerical errors in DFT calculations A comprehensive discussion of the sources of numerical errors in plane-wave DFT calculations can found in ref . Herbst et al. implemented rigorous error bound for plane-wave DFT calculations, excluding contributions of the SCF iterations, i.e., for the nonself-consistent Kohn–Sham equations. To this end, the authors applied the Kato–Temple bound ( eq ). The calculation of the residual is simplified when using a plane-wave basis set, and an upper bound can be calculated using a larger plane-wave cutoff. This is combined with an lower bound for the energy gap between eigenvalues and with a bound on the numerical error, which is tracked by using interval arithmetics in their implementation. The resulting error bars present a rigorous bound on the error in the orbital energies within the considered model. Further mathematical analyses of convergence and error bounds are available for DFT calculations, , including error bounds for the basis set error with plane-waves, in augmented plane-wave approaches, and for atomic orbital-like basis functions. Also for error sources that are not treated in the implementation of Herbst et al., such as the SCF procedure , and k -point sampling, the mathematical tools for establishing rigorous error bounds for total energies are available. However, as for wavefunction methods, it is unclear whether these translate to useful error bounds on relative energies. Error of Exchange–Correlation Approximations A main source of error in DFT calculations is the approximate exchange–correlation (xc) functional E xc [ρ]. A plethora of approximations are available, − and while numerous studies exist that benchmark their accuracy for carefully devised test sets (see, e.g., refs and ), it is generally unclear how accurate a chosen approximate xc functional will be for a specific molecule. Here, we will focus on approaches that are able to assign an uncertainty to the energy calculated for a specific molecule . One simple measure that often gives an indication of the accuracy of a DFT calculation is the variation with the percentage of exact exchange λ that is included in the xc functional. This can be captured by measures such as the B 1 diagnostic: B 1 = TAE [ BLYP ] − TAE [ B 1 LYP ] n bonds 81 where TAE[BLYP] is the total atomization energy calculated with the nonhybrid functional BLYP, TAE[B1LYP] is the total atomization energy calculated with the hybrid functional B1LYP, which contains 25% of exact exchange, and n bonds is the number of covalent bonds in the molecule. The A λ diagnostic uses a different normalization. When, for instance, applied to the PBE0 functional (with λ = 0.25), it is defined as A 25 [ PBE ] = 1.00 0.25 TAE [ PBE ] − TAE [ PBE 0 ] TAE [ PBE ] 82 While such measures can help to identify molecules for which a large uncertainty can be expected for DFT calculations in general, they do not provide an accuracy that is related to a specific choice of xc approximation. Therefore, they cannot be used to guide the selection of a specific xc approximation, or help identify a suitable amount of exact exchange. Many approximations to the xc functional contain empirical parameters a that are fitted such that the root-mean-square error (or a similar cost function C ( a )) for a training set is minimized. The framework for Bayesian error estimation functionals (BEEFs) in DFT originally proposed by Mortensen et al. extends the best-fit set of parameters a 0 by a probability distribution p a for the empirical parameters in the xc approximation, i.e., instead of a single (best-fit) approximate functional, one uses an ensemble of functionals. The BEEF framework assumes that for parameters that enter the xc functional linearly, p a ∝ exp ( − C ( a ) / T ) 83 where the spread of the distribution is determined by the ensemble temperature T = 2 C ( a 0 ). For details on the determination of this probability distribution, we refer the reader to refs and and . By sampling from this probability distribution and performing DFT calculations with different sets of parameters, it becomes possible to obtain a probability distribution for the resulting energies, from which a mean and a variance can be assigned to each result. An extension of this theoretical framework was put forward by Aldegunde et al., who apply a relevance vector machine learning model and obtain probability distributions for the model parameters using Bayesian inference (see section ). Within the BEEF framework, Wellendorff et al. developed the BEEF-vdW GGA functional as well as the mBEEF meta-GGA functional. Both restrict the parameter space to linear parameters in the GGA enhancement factor of the exchange functional. The sampling of the parameter space is only performed post-SCF, i.e., the electron density is not reoptimized for each set of parameters, which allows for an efficient evaluation of the uncertainty. By sampling the parameter distribution, the BEEF framework makes it possible to not only obtain uncertainties for individual energies, but also for energy differences and other properties derived from the energy. This has been exploited to assess the uncertainties in harmonic vibrational frequencies (see section ), finite-temperature thermodynamic properties of solids, as well as for various applications of DFT in heterogeneous catalysis. − While these BEEF functionals provide a possibility for estimating uncertainties due to the xc approximation, it is unclear whether the parametric uncertainty in the BEEF model is able to correctly predict the error distribution of these functional. ,, Therefore, Reiher and co-workers developed a system-specific long-range corrected PBE functional within the BEEF framework. , They reoptimize five (linear and nonlinear) parameters in this functional for a specific class of systems of interest, but only use the one linear parameter for estimating the uncertainty. They stress that a system-specific parametrization is crucial not only for improving the accuracy of the xc functional for a problem at hand, but also for the reliable estimation of uncertainties. Walker et al. used a latent variable model to predict uncertainties in DFT-calculated energies. For the energies of the species of interest y , they performed DFT calculations using four different approximate xc functionals and fitted a model of the form y = W z + μ + e 84 where z is a normally distributed latent variable, μ are the mean values predicted over the considered functionals, and e is a zero-mean Gaussian distributed noise with specific variances Ψ . The unknown parameters W and Ψ are obtained by maximizing a log-likelihood function such that the model is consistent with the considered ensemble of approximate xc functionals. The uncertainties provided by this model can then be propagated to assign uncertainties to other quantities of interest, such as turnover frequencies, apparent activation barriers, and reaction orders. This was applied to model the various experimental kinetic data and to identify active sites of the water–gas shift reaction over Pt-based catalysts. A similar approach has been applied to investigate ethane dehydrogenation and hydrogenolysis on Pt surfaces. An alternative approach to quantifying the uncertainties due to the xc approximation is the parametrization of Δ-machine learning models that predict the difference between approximate DFT calculations and accurate high-level quantum-chemical calculations. The simplest case is a linear regression model, X = β T + ϵ 85 which predicts the experimental values X from the calculated values T , and where β is a linear regression coefficient and ϵ is a random error with zero mean. Lejaeghere et al. applied such a model to the prediction of cohesive energies, bulk moduli, and elastic constants of elemental crystals, and assign uncertainties to the predictions of solid-state DFT calculations with approximate xc functionals. Similarly, Pernot et al. employed linear calibration models to access the prediction uncertainty of properties from solid-state DFT calculations. Furthermore, uncertainties have been assigned − to linear correction models that provide corrections to DFT energies (e.g., by interpolating between DFT and DFT-U) , in solid-state chemistry. Simm and Reiher applied Gaussian process regression to learn the difference between PBE and extrapolated MP2 atomization energies, and apply this for the exploration of chemical reaction networks (see also section ). To this end, their model uses geometry-dependent descriptors. While the Δ-machine learning model provides a prediction of the error in the DFT calculation, they actually use the energies corrected by this model (i.e., approximations to the high-level MP2 results) as their quantity of interest, and employ the uncertainty provided by the Gaussian process (see section ) to assess the accuracy of their model. If this uncertainty becomes too large, additional high-level calculations are performed to improve upon the model in a rolling reparametrization (active learning). For an extensive test set of binary and ternary oxides, Yuk et al. trained a random forest regression model to predict the error with respect to experimental reference data for four approximate xc functionals. As features, their model uses descriptors from materials informatics, specifically the fractional ionicity, charge per valence electron, valence electrons per atomic number, Pauling electrostatic strength of the metal–oxygen bond, mass density, oxygen fractional occupation, and element specific DFT ionization errors. Similarly, but without reference to experimental or computational reference data, Alfonso-Ramos et al. trained a machine learning model using structural fingerprints as features to predict the variance in the activation energies of pericyclic reactions calculated by 20 approximate xc functionals. For transition metal complexes, Kulik and co-workers developed a Δ-machine learning model that predicts the difference between the spin-state energies calculated with DFT using an approximate xc functional and with DLPNO-CCSD(T), i.e., the error due to the xc functional. Their model relies on descriptors that depend on the electron densities calculated for the high-spin and low-spin state with a specific xc functional (B3LYP). By parametrizing such models for 48 different approximate xc functionals, they are able to recommend the best-performing xc functional for a specific transition metal complex. Errors of Dispersion Corrections Since semilocal xc approximations are known to miss dispersion interactions, such xc functionals are mostly combined with semiclassical dispersion corrections. , These corrections model the error in the dispersion energy via an atomic-pairwise potential in terms of C 6 (and possibly C 8 , C 10 , ...) coefficients, in combination with an empirical damping function in the short and medium range. This damping function contains empirical parameters that are fitted to experimental or high-level computational reference data. Semiclassical dispersion interactions provide a well-defined model that can be used as a test case for uncertainty quantification in computational chemistry. For first-generation (DFT-D) dispersion interactions, , Hanke assessed the sensitivity of the dispersion energy with respect to their different parameters. Using ad hoc estimates of uncertainties in these parameters, they investigated the resulting uncertainties in calculated binding energies. Conversely, by comparing to reference data, they estimated the uncertainty in a global parameter of the dispersion correction. For the D3 dispersion correction, Weymuth et al. performed an in-depth statistical analysis of the sensitivity with respect to its empirical parameters. They performed a bootstrapping analysis ( section ) to assign uncertainties both to these parameters and to the predicted dispersion energies. Moreover, they performed a jackknife analysis to assess the quality of the reference dataset that was used to parametrize D3. Proppe et al. trained a Gaussian process model for the dispersion energy using the interaction energy difference between DFT and coupled-cluster calculations as training data, using only structural information that is also available to the D3 model as input features. This Gaussian process provides not only a prediction of the dispersion energy, but also for the associated uncertainty. This active-learning strategy , can, in turn, be used to refine the model in regions where this uncertainty becomes too large. 4.1.3. Semiempirical Quantum Mechanics Semiempirical quantum-chemical (SQM) methods − employ a Hamiltonian akin to the one appearing in the HF or DFT formalism, usually in a minimal basis set, and approximate the parameters appearing in it using simplified models. These parameters are generally obtained by a fit to suitable reference data, which makes then amenable to uncertainty quantification. One attempt at quantifying uncertainties in the PM7 method was undertaken by Oreluk et al. Using the Bound-to-Bound Data Collaboration framework, − they assigned feasible intervals to each parameter that are consistent with the training data, and propagated these to error intervals for the heats of formation of linear alkanes. They found that the accuracy of the training data was preserved for the predicted heats of formation within bounds of chemical accuracy if predictions were made for the molecules of comparable size, but that the error grows linearly with the relative size of the molecules. Finally, we mention that Δ-machine learning models based on Gaussian processes as discussed in section for DFT are also applicable in combination with SQM methods. 4.1.4. Classical Force Fields Classical force fields model the ground-state potential energy surface using classical interaction terms. In their most common form, ,, E FF ( R 1 , ... , R N ) = 1 2 ∑ i bonds k i ( r i − r i 0 ) 2 + 1 2 ∑ j angles k j θ ( θ j − θ j 0 ) 2 + ∑ n torsions cos ( n n ω n − γ n ) + ∑ I ∑ J > I q I q J r I J + ∑ I ∑ J > I 4 ϵ I J [ ( σ I J r I J ) 12 − ( σ I J r I J ) 6 ] 86 they include bonding terms, which depend on bond lengths r i , bond angles θ j , and torsional angles ω n , as well as nonbonding electrostatic and van der Waals interactions, which depend on interatomic distances r IJ = | R I – R J |. Such force fields form the basis of molecular dynamics simulations that can be used to determine finite-temperature properties as well as free energies. The empirical parameters in the force field, namely, the bonding parameters k i , r i , θ j , θ j , γ n , n n , the partial charges q I , and the Lennard-Jones parameters ϵ IJ and σ IJ , are determined using a wide range of strategies. These parameters will therefore be subject to uncertainties that will incur uncertainties in the sampling-based properties determined in molecular dynamics simulations using such classical force fields. Brunken and Reiher proposed a system-specific parametrization of classical force fields of this form with respect to quantum-chemical reference data. For their parametrizations, they obtained uncertainty estimates by using the standard deviations from a k -fold cross-validation and from a separate Δ-ML model. However, in contrast to machine learning interatomic potentials (see section ), classical force fields are usually not parametrized to reproduce the potential energy surface from quantum-chemical calculations. Therefore, the uncertainty of force fields parameters cannot be determined directly, but has to be inferred from their impact on sampling-based properties. Such approaches will be discussed in section , in particular in section . 4.1.5. Machine Learning Interatomic Potentials Machine learning interatomic potentials (MLIPs) aim to model the potential energy surface E 0 ( R 1 , R 2 , ...) of large systems at low computational cost with the accuracy of high-level ab initio calculations. , In reality, these are often targeted to reproduce either static DFT or ab initio molecular dynamics simulations. , Thus, MLIPs will inherit the flaws and errors of the reference data and will never be more accurate than them. Therefore, UQ in the context of MLIPs always aims at quantifying the uncertainty with respect to their reference method. Various forms and architectures of MLIPs have been proposed in recent years, and the most relevant ones in terms of this review and uncertainty quantification are briefly recapitulated in the following. Neural network potentials (NNPs)as the name suggestsuse neural networks (NNs) to represent the potential energy of a system as an arbitrary analytical function of the atomic positions. Initiated by the work of Behler and Parrinello, various types of NNPs have since been developed, as comprehensively reviewed by Behler. An integral part of the training process of NNPs and many other MLIPs is active learning (AL), which refers to the iterative selection of the most informative data points (e.g., atomic configurations) to include in the training set, guided by uncertainty or model disagreement. This approach reduces the need for large, uniformly sampled datasets by focusing expensive reference (e.g., DFT) calculations on regions of configurational space where the model is least confident. Several strategies exist to guide this data selection process based on uncertainty quantification: Query-by-committee (QBC) uses an ensemble (committee) of modelstypically neural networksand quantifies uncertainty via their disagreement on a given input, e.g., by the variance or maximum deviation in predicted forces. If the ensemble variance is large for a given configuration, a reference calculation is performed and the new data is added to the training set. A key advantage of ensemble-based AL methods is the computational efficiency of their uncertainty quantification. While some authors have found that the predicted uncertainty and the actual error are not always linearly related, techniques to overcome this issue have been proposed. An alternative UQ approach is Gaussian process regression (GPR), , where uncertainty is estimated analytically from a covariance (kernel) function that measures the similarity between unseen and previously observed data points. Another class of machine learning interatomic potentials is based on the atomic cluster expansion (ACE) − introduced by Drautz and co-workers, a framework that expresses the potential energy as a linear combination of invariant basis functions constructed from atomic clusters. In the ACE approach, uncertainties can be estimated, for instance, using Bayesian linear regression, where the posterior distribution over model parameters naturally yields predictive variances. Fur further discussions of UQ for MLIPs, we refer to other reviews and perspectives. ,,,,, QBC approaches have for example been calibrated and applied to Raman spectroscopy, solvation of organic molecules as well as the radial distribution function and thermodynamic state functions of liquids by Ceriotti and co-workers. Their ensemble calibration is based on GPR and discussed in section . Isayev, Dral, and co-workers applied QBC to estimate uncertainties in the prediction of thermochemical properties with MLIPs. Recently, Kellner and Ceriotti introduced direct propagation of shallow ensembles, where the NN models of an ensemble share all but the last layer weights. Using datasets of DFT-computed properties for liquid water, lithium thiophosphate (Li 3 PS 4 ), barium titanate (BaTiO 3 ), and QM9 molecules, the authors demonstrate that their approach balances computational efficiency with reliable uncertainty estimation. Upon comparison against deep ensembles and conformal prediction methods, the approach improves predictive uncertainty without significant computational overhead. In this regard, conformal prediction can be used to estimate error bars for energy or force predictions, with formal statistical guarantees. Deep ensembles refer to ensembles of NNs, where each NN is augmented to also predict an uncertainty (e.g., the variance) for each input. These were investigated and compared to traditional committees by Carrete et al. They developed different MLIPs for the ionic liquid ethylammonium nitrate using both approaches and find that when using homogeneous training data (i.e., all data points generated at the same quality and/or method), deep ensembles do not show superiority compared to committees in terms of their uncertainty metrics. However, when training data from different sources was used, the deep ensemble could clearly discriminate between those sources. A different approach was chosen by Lin et al., who focus on generating neural network (NN) based reactive potentials, specifically applied to H 2 dissociation on Ag(111) and Ag(100) surfaces. They train an ensemble of NNs, then choose the two models with the lowest training errors, and from a newly defined uncertainty metric (the negative of squared difference surface) of these two models, determine points on the PES that should be subject to new ab initio reference data calculations. For their examples of hydrogen on silver surfaces, that approach worked reasonably well. The dropout method was developed to overcome the computational costs of training a large committee of models required for good statistics. Dropout NNs are created from fully connected NNs by randomly removing the outgoing connections of some nodes in each layer, which is promising, because it has been shown that when using a variational inference approach, a dropout NN approximates a Bayesian NN thus enabling Bayesian UQ. Dropout-based and related techniques are best understood as pragmatic tools for probing predictive variability in ML models, rather than as comprehensive uncertainty quantification methods in the traditional UQ framework. Such approaches provide heuristic and approximate uncertainty estimates and should therefore be interpreted with appropriate caution, particularly with respect to overconfidence. The reliability of uncertainty estimates obtained from such methods should therefore be assessed using dedicated uncertainty validation techniques (cf. section ) rather than inferred from the uncertainty model alone. Wen and Tadmor used dropout NNPs in predicting mechanical properties of graphene. With a similar motivation, Zaverkin et al. use biased MD simulations to study the capability of MLIPs to sample the complex conformational space of alanine dipeptide and a flexible MOF. The authors use gradient-based uncertainty estimates, which they had developed in a previous work: that is, the gradient of the output with respect to the model parameters indicates how much the prediction changes when the parameters are perturbed. If the model’s prediction is highly sensitive to small changes in the weights, it implies that the prediction is likely to be uncertain. Then, the uncertainty is used to bias the MD potential to drive the simulation to previously unexplored (i.e., high-uncertainty) regions. It is found that this approach greatly accelerates the exploration of the conformational space and that the accuracy of the developed MLIPs is on par with selected ensemble-based methods, while at a notably reduced computational cost. In the context of ACE-based MLIPs, Lysogorskiy et al. compared uncertainty estimation strategies based on ensemble learning and the D-optimality criterion. The latter selects the most informative atomic configurations by maximizing the determinant of the information matrix (constructed from the ACE basis functions), which corresponds to minimizing the overall uncertainty in the fitted ACE model parameters. The authors test the two strategies on structural properties of copper, MD simulations of water and structures of Li 4 clusters. They finally conclude that both ensemble and D-optimality learning provide similar predictions, while the former is computationally more expensive due to the training of multiple models. Best et al. combine ACE potentials with conformal prediction for investigating structural and stress properties of bulk silicon. While most of the aforementioned works in this section explore various methods to determine the uncertainty of a model prediction in their active learning scheme, van der Oord et al. worked on hyperactive learning, an accelerated AL scheme and combine it with ACE potentials. Similar to the aforementioned work by Zaverkin et al. and a study by Kulichenko et al., biased MD simulations are used to expedite the exploration of previously unseen parts of the phase space of the AlSi10 alloy and a polyethylene glycol polymer. Other than that, Bartók and Kermode state that in the context of MLIPs, the use of “probabilistic learning methods such as Gaussian process regression (GPR) is currently under-exploited, because of the tendency of the predicted errors to overestimate the true error”. In a proof-of-concept work, they expand on GPR and develop MLIPs for argon dimers and trimers calculated using a CCSD(T) reference. The authors observe significantly improved accuracy of predicted error estimates of the MLIPs when optimizing the hyperparameters of their GPR by maximizing the leave-one-out cross-validation likelihood. Botu et al. have employed kernel-ridge regression (KRR), which is conceptually related to GPR, in the sense that both rely on a kernel function to measure the similarity of data points, but KRR gives the mean prediction of GPR under Gaussian noise, while only GPR provides formal uncertainty estimates. The authors presented a general framework for deriving MLIPs, using the example of elementary aluminum. The key ingredients of their workflow is an UQ approach of deriving confidence intervals for the force field predictions by deriving an expression for the standard deviation of the predicted force error as a function of the discrepancy between the predicted and reference data fingerprints, which works well for the systems considered in their study. 4.1.6. Multiscale Models Multiscale models combine different computational methods to achieve a description of complex molecular systems and materials. There is a vast number of multiscale modeling strategies, which can broadly be categorized into two classes. First, in vertical multiscale models, accurate methods at smaller length and time scales are used to obtain parameters that are then used as input for lower-resolution methods which are able to reach larger length and time scales. This strategy is widespread in multiscale materials modeling. Here, uncertainty quantification can be performed by propagating an uncertainty of these parameters obtained at the smaller scale to the quantity of interest provided by the larger scale model. As the modeling of large-scale materials properties is beyond the scope of this review, we only refer to previous reviews on uncertainty quantification in multiscale materials modeling. − Second, in horizontal multiscale models, computational methods of different accuracy are used to describe different parts of a complex molecular system. Usually, there is a particular region of interest, such as the active center in an enzyme, for which more accurate (and more computationally expensive) methods are used, while its environment is described using less accurate and computationally cheaper methods. The most prominent example of such a strategy are QM/MM methods. − In addition, there is a plethora of QM/QM methods, in which different quantum-chemical methods are combined. ,− Closely related are quantum-chemical fragmentation methods, in which a complex molecular system is split into smaller fragments, which are treated individually using the same quantum-chemical method. , Even though all these methods have been developed with the goal of allowing for the prediction of a quantity of interest (e.g., the activation energy in an enzymatic reaction) with a certain required accuracy, while accepting larger uncertainties for parts of the system that are less relevant for this quantity of interest, there have been hardly any attempts to rigorously quantify these uncertainties. Many-Body Expansion The many-body expansion (MBE) is a prototypical fragmentation method, in which the total energy of a large system is approximated as E tot = ∑ I E I + ∑ I < J Δ E I J + ∑ I < J < K Δ E I J K + ... 87 where E I is the energy of the I th fragment, Δ E IJ = E IJ – E I – E J is the interaction energy of the dimer made up of fragments I and J , Δ E IJK is a trimer interaction energy, and so on. The sum is usually truncated at low order, i.e., after two-body or three-body terms. Herbert and co-workers investigated the propagation of numerical errors in the MBE. Because the individual interaction energies in eq are obtained as differences between total energies, their precision in floating point arithmetics is limited, and the numerical errors in the many small interaction energies can accumulate to substantial errors in the total energies, in particular for higher-order contributions. Several studies have investigated the convergence of the MBE, in particular for molecular clusters, and have devised strategies for screening higher-order contributions, i.e., to neglect selected terms in eq . Such strategies areat least implicitlybased on an estimation of the size of these individual higher-order contributions and of the uncertainty in the total energy resulting from their neglect. The simplest strategy estimates the relevance of individual contributions from the distance between the involved fragments (distance-based screening). − A more sophisticated estimate relying on electrostatic dipole–dipole interactions, which also takes the orientation of the different fragments into account, was developed in ref . Finally, the individual contributions can be estimated by calculating them at a lower level of theory, such as a polarizable force-field or semiempirical quantum mechanics (energy-based screening). , In the many-body expanded full-CI (MBE-FCI) method, an expansion akin to eq is applied to the correlation energy, with the correlated orbitals used as “fragments”. Greiner et al. developed an error estimation scheme for the MBE-FCI method. First, for each orbital it establishes an estimate of the maximum magnitude of higher-order contributions by a fit to the lower-order contributions. Second, for groups of orbitals it determines the distribution of the contributions at a certain order by Monte Carlo sampling. The latter makes it possible to exploit the fact that many contributions of opposite signs will cancel. These error estimates are used to truncate the MBE-FCI expansion, while rigorously controlling the maximum error in the resulting total energy. QM/MM Models In QM/MM models, the system is split into a region of interest, which is treated using quantum-chemical methods (QM region), and its environment, which is treated using a classical force field (MM region). Despite the long history of QM/MM methods − and their broad application, in particular for studying enzyme catalysis, − the accuracy of such models is hard to assess. This was recently highlighted by Giudetti et al., who found that the QM/MM reaction energies obtained with implementations in different software packages can show very large differences, even if the same structural models and QM/MM methods are used. These difference could partly be traced back to minor differences in technical settings, which propagate to substantial differences in the calculated reaction energies. This underlines the need for systematic uncertainty quantification and error control in QM/MM calculations. This becomes even more pressing when trying to construct suitable QM/MM models for a specific system of interest, which involves numerous choices by the computational scientist. ,, Most relevant is the choice of the QM region (i.e., selecting which parts of the system are treated quantum-chemically). Systematic studies have demonstrated that the convergence with respect to the size of the QM region is slow and in many cases not monotonic. − Several schemes have been developed to assist or automate the choice of the QM region in biomolecular QM/MM calculations (for a review, see ref ). The first group is based on some form of sensitivity analysis, in which the effect of the change in some parameter related to each amino acid on the property of interest (usually an energy difference) is used to select those with the largest sensitivity for inclusion in the QM region. In the charge-deletion analysis, , the effect of deleting the MM charges of individual amino acids is used. In the point-charge variation analysis, , this is extended by using a local sensitivity analysis with respect to MM charges. A second group of approaches uses different descriptors to predict the importance of individual amino acids in the QM region, often based on calculations using large QM regions. The charge-shift and Fukui shift analysis compare the change in the charges between the apo and the holo form. Brunken and Reiher compare the energy gradient between large and small QM regions. Cisnero and co-workers proposed an approach based on protein sequence similarity, and in ref , protein network centralities were explored as descriptors in the construction of the QM region. 4.2. Uncertainties in Trajectories and Related Properties While the works discussed in section pertained to solving the stationary (time-independent) Schrödinger equation or approximations thereof, this section is devoted to works exploring the dynamics of matter i.e., solving time-dependent problems. The most popular method of dynamically traversing potential energy landscapes are molecular dynamics (MD) simulations, which propagate an ensemble of particles in time by solving Newton’s equations of motion. Given an ensemble of N particles with masses m i and nuclear positions R i , the force F i acting on particle i is given by F i ( t ) = m i d 2 R i ( t ) d t 2 = − ∇ R i V ( R 1 , ... , R N ) 88 where t is the time, and V is the potential that explicitly depends on the positions of all N particles. We note that for classical force-field-based MD simulations, V is usually equivalent to E FF from eq . Unfortunately, the above equation cannot be solved analytically for N > 2 and is thus numerically integrated in small intervals δ t , typically 0.5 to 1 fs. During a simulation, the phase space Γ( t ), a function of the positions and velocities, is sampled, and these quantities are stored as a trajectory. For completeness, we also mention the concept of shadowing , which has gained attention in the simulation community recently. − It is known that trajectories with slightly different initial conditions (positions or velocities) diverge exponentially in phase space. As MD simulations rely on numerical integration for solving Newton’s equations of motion, the conservation of energy, and maintaining the “true” path in phase space are critical for producing reliable predictions. It was found that modern symplectic integrators from the Verlet family usually produce trajectories that shadow the true trajectory, meaning that the simulated trajectory stays “close” in phase space to the true one. It is evident that the complexity of the integration is largely dependent on V in fact, different flavors of MD simulation are distinguished by certain classes of functional forms V can adopt. Among the most common forms, force-field-based MD simulations and ab initio MD (AIMD) simulations have evolved. The former typically rely on a pairwise-additive and computationally simple force field (see section ), thus allowing for simulations involving several thousand atoms and time scales up to the μs range. , AIMD simulations feature a more complex potential, mostly by means of DFT (see section ), and therefore, come with drastically increased computational cost, but provide information on electronic structure of a dynamic system, enabling, for example, the computation of vibrational spectra from time-correlation functions of the molecular dipole moments. − For a more detailed introduction, the reader is referred to refs , , , , and . MD simulations generally suffer from both aleatoric and epistemic uncertainties (refer to section ), the former primarily being caused by incomplete sampling due to finite computational resources, and the latter arising through the lack of information on the physical/chemical model. Aleatoric uncertainties can often be compensated for by investing more computing time and performing more or longer simulations, and applying enhanced sampling techniques. The simulation community has become increasingly aware of this matter, and efforts to promote standardized simulation protocols for reproducible simulations have been made (see ref and references therein). Specifically, Coveney and co-workers emphasize repeatedly that MD is intrinsically chaotic and therefore, ensemble simulation methods are needed independent of the duration that is carried out, because MD is extremely sensitive to the initial conditions, “making accurate predictions impossible and one-off observations largely unreproducible even though their underlying dynamics is deterministic”. Consequently, it has been recommended by several researchers ,, to increase the statistical robustness of MD simulation studies by performing an ensemble of simulations (also known as replica simulations), and computing properties as ensemble averages including meaningful uncertainties. Tackling epistemic uncertainties for MD simulations typically involves either (i) improving the physical model, e.g., by adjusting the functional form of the force field, or (ii) updating the parameters of the physical model to reproduce certain reference data more accurately, or (iii) improving the chemical model by including more molecules or a considering a different representation of the system of interest. While measures to treat aleatoric uncertainties sometimes appear straightforward, suitable paths to treating epistemic uncertainties may not always appear obvious. However, various UQ techniques have been applied to both sources of uncertainties in the context of MD simulations, and are reviewed in this section. At the same time, epistemic uncertainties in MD simulations may also arise from issues of code correctness and model adequacy. In the terminology commonly used in VVUQ frameworks, verification refers to ensuring that the numerical implementation correctly solves the underlying equations of the chosen model, whereas validation concerns the extent to which the selected model faithfully represents the physical system of interest. While verification and validation are closely related to epistemic uncertainty, they address different questions than UQ itself and are therefore not treated as primary topics in this section. Besides MD simulations, Monte Carlo (MC) methods constitute the second big branch of sampling-based simulation methods with a widespread use in in silico chemistry. In an MC simulation of an ensemble of particles with a given configuration (geometry), a random step in phase space is proposed by randomly moving one (or possibly more) particles. The energy for the new configuration is calculated and the move to that new configuration will be accepted with a certain probability, typically involving the negative exponential of the energy. Otherwise, a new random step in the phase space is proposed, and the procedure is repeated. Recent advances in this field concern hybrid Monte Carlo, , kinetic Monte Carlo, , and quantum Monte Carlo methods. − It is important to note that similar to MD, MC simulations are inherently stochastic too, as detailed in section and refs and . So far, the literature on UQ for MC simulations remains sparse. 4.2.1. Simulation Parameters Tracing and propagating the uncertainty from input to output (commonly referred to as sensitivity analysis, section ) is an important challenge in the general field of computational simulations and consequently, various approaches to tackle this challenge have been reported. Applied to molecular dynamics simulations, the aforementioned mostly refers to investigating the propagation of uncertainties in the simulation input parameters to the QoIs predicted based on the simulation, such as density, diffusion coefficient, or radial distribution function. Uncertainties in the input parameters are mostly investigated with a focus on (i) force field parameters (bond lengths, partial charges, LJ parameters) or (ii) system parameters (e.g., number of molecules, temperature, simulation box size). For completeness, we will also distinguish local sensitivity analysis ( section ), that considers functional derivates and polynomial approximations thereof for quantifying uncertainties (also known as “forward propagation”), from global sensitivity analysis ( section ), where Sobol’ (or other types of) indices are computed to derive an ordering, indicating how strongly each input parameter affects the output. Most works discussed in this review fall into the first category, likely because the second often is computationally much more demanding. Given that UQ for MD simulations is an emerging, and not yet fully established field, it is understandable that existing works on (local or global) sensitivity analysis for MD simulations mainly address either systems with small molecules or well-defined, simple materials, featuring manageably complicated functional forms and low-dimensional parameter spaces. We will first discuss works with an application to molecular systems. Rizzi et al. , performed a local sensitivity analysis on force field parameters in MD simulations, at the example of the TIP4P water model. By combining polynomial chaos expansions with Bayesian inference (see sections and ), they found that the force field parameters are scarcely sensitive to each other while being very sensitive toward specific observables of the system, such as density and diffusivity. After that, they conducted two studies on the ion transport of NaCl in aqueous solution through silica nanopores. While the first work focused on examining the influence of the pore diameter and the gating charge on the ionic flux, the authors investigated uncertainty in the potential parameters and their impact on the flux in the second work. Molinero and co-workers ,, followed a similar approach, employing polynomial chaos expansions in deriving parameters for a coarse-grained force field model for simulating hydrated anion-exchange membranes and fuel cell membranes with controlled water uptake. Peerless et al. studied how the uncertainty associated with partial atomic charges in MD simulations impacts diffusion, density and solubility in the bulk phase, using the example of liquid acetonitrile, described by the General Amber Force Field (GAFF). The formulation and application of a Gaussian process regression model ( section ), demonstrates a notable speed-up in generating additional sampling and thus, achieves rapid predictions of local sensitivities of bulk phase properties. In the context of high-performance computing, Angelikopoulos et al. introduced a Bayesian probabilistic framework designed to quantify and propagate the uncertainty in parameters of force fields for classical molecular dynamics simulations. At the example of liquid argon, it is demonstrated how uncertainties in the parameters of the intermolecular interaction potential (in this case, the Lennard-Jones potential) are propagated in MD simulations, resulting in confidence intervals for the predicted transport quantities (e.g., diffusion coefficient and viscosity). The authors then build on their previous work and systematically assess the uncertainties of properties predicted in MD simulations of nanofluidic water transport and confined water/carbon interfaces. Messerly et al. investigated the transferability of Mie λ-6 force fields for alkanes, derived at standard conditions, to very high temperatures and pressures. By utilizing Bayesian inference, the authors demonstrated that, for the systems investigated, there is no valid combination of Mie potential parameters that can predict liquid–vapor equilibrium properties (density and pressure) at harsh thermodynamic conditions. From the authors’ viewpoint, that highlights the demand for more accurate potentials in applications with industrial interest. In a follow-up study, the authors further explored finding the optimal force field parameters to describe the aforementioned vapor–liquid equilibrium properties. Two types of surrogate models were explored in sampling the parameter space: One model, multistate Bennet acceptance ratio was found to be promising when exploring remote regions of the parameter space, while the other, pair correlation function rescaling proves to be helpful in the more localized regime. Raabe et al. used Gaussian process regression ( section ) for deriving force field parameters for trans -1,2-dichloroethene, based on its vapor–liquid equilibrium properties. On a related note, Madin et al. examined three types of Lennard-Jones force fields with optional additional quadrupole terms and increasing complexity, targeting at density, saturated vapor pressure and surface tension of Br 2 , F 2 , N 2 , O 2 , C 2 H 2 , C 2 H 4 , C 2 H 6 , and C 2 F 4 . By using Bayesian inference for selecting the type of force field, a model’s complexity and computational effort is weighed against the accuracy and precision of its predictions. In that study, it turned out that an increased model complexity compared to the common Lennard-Jones potential is only worthwhile in a few cases. With the goal of finding the optimal parameters for Mie λ-6 force field for liquid neon, Shanks et al. presented local Gaussian process surrogate models, trained on X-ray/neutron diffraction scattering data. They find that their local surrogate models are much faster in computing the force field parameters because of a better scaling with the number of independent variables compared to standard Gaussian processes. Dutta et al. studied the Bayesian calibration of force field-parameters of water and helium, assuming their uncertainty not to be Gaussian. In particular, they used approximate Bayesian computation, a likelihood-free inference scheme, and its implementation for HPC systems. The application of their method is presented using datasets from neutron and X-ray diffraction measurements and MD simulations. As the proposed method provides access to the entire posterior distribution, uncertainty quantification of the model predictions is possible and thereby, can in principle, calibrate force fields from any type of structural or dynamic property data. Turning to solid-state applications, Tran and Wang introduce a reliable molecular dynamics (R-MD) framework to assess the output of MD simulations, given the input uncertainty in the intermolecular potential and/or the parameters. At the heart of R-MD, there is the formulation of input uncertainty in the potential as intervals, and as a consequence, positions and velocities are interval-valued, too. In a follow-up work, the authors describe four different schemes of uncertainty propagation in the framework of R-MD and apply their developments to predict the tensile uniaxial deformation of aluminum single crystals, described by an embedded atom method (EAM) potential. Similarly, Dhaliwal et al. also studied the local sensitivity of the parameters of EAM potentials of face-centered cubic aluminum, using a Bayesian statistical framework, while Vohra et al. used nonequilibrium molecular dynamics (NEMD) simulations to capture thermal transport in silicon and performed a local sensitivity analysis on the potential parameters. By deriving a surrogate model for the NEMD simulations from a reduced-order polynomial chaos expansion ( section ), the authors could quantify which parameters in the employed Stillinger–Weber potential contribute to the uncertainty of predicted bulk thermal conductivity in Si. Studying solid nickel, Longbottom and Brommer implemented a Bayesian framework for propagating uncertainties in three types of potentials (LJ, Morse, and EAM), derived from DFT reference data, to simulated lattice constants, elastic moduli and thermal expansion coefficients. The work by Kurniawan et al. investigates classical interatomic potentials such as the Lennard-Jones, Morse, or Stillinger–Weber potential for several metals and materials from a conceptual viewpoint and shows that these are mostly sloppy, i.e., insensitive to concerted changes in certain sets of parameters. Lastly, a few applications of local sensitivity analysis with coarse-grained MD simulations are known, mainly by Müller-Plathe and co-workers. They developed a coarse-grained molecular dynamics-finite element coupling approach to investigate the mechanical behaviors of polymers. That approach partitions the system under investigation into an MD region and a continuum (finite element) region, requiring several technical input parameters, whose values cannot always be deduced based on physical principles. Therefore, in their first work, the authors investigated polystyrene and performed a local sensitivity analysis of the polymer structure (density, radius of gyration, end-to-end distance and radial distribution function) on these input parameters. They find that the simulation technique is generally robust with the polymer properties being only weakly dependent on the simulation parameters (number of anchor points, force constant between polymer and anchor points and the size of the molecular dynamics domain). This first work, however, exclusively focused on the MD region and did not consider the continuum region in the UQ. In a follow-up work, the authors then also considered the continuum region in their UQ treatment and identify trustworthy ranges for the simulation parameters. We also reference a rather mathematical description on sensitivity analysis of observables from Langevin dynamics simulations. Further works involving sensitivity analysis in the context of free energy calculations are discussed in section . 4.2.2. Time Correlation Functions In molecular dynamics simulations, dynamical quantities, such as self-diffusion coefficients, viscosities as well as thermal and ionic conductivities, are accessible via Green–Kubo type correlation functions. Those, however, can be subject to large uncertainties caused by inadequate sampling. At this point, replica trajectories in conjunction with bootstrapping ( section ) are suitable to increase the statistical robustness of the calculations. Fischer et al. developed a protocol for bootstrapping correlation functions from simulations of ethane, propane, and dimethyl ether in conjunction with the time decomposition method of Maginn and co-workers. Desbiens et al. also apply bootstrapping to the current and velocity autocorrelation function, highlighting that the increased sampling can compensate for shorter simulation times or smaller number of replica simulations. Kirchner, Frömbgen, and co-workers provided two tutorials on how to reduce the uncertainty in the simulation and analysis of ionic liquid trajectories, , with a focus on dynamic properties, such as self-diffusion coefficients and ionic conductivities. These quantities involve calculating the (individual or collective) mean squared displacement (MSD) of particles, based on correlation functions of the particles’ positions, and performing a subsequent linear regression of the linear regime of the MSD. Rigorously identifying that linear regime is not a straightforward task, especially in the case of the collective MSD which is used for computing the ionic conductivity, due to poor sampling. In avoiding “eyeball statistics”, Frömbgen et al. proposed a systematical procedure for identifying linear regimes in MSD data, by exploiting the fact that the slope of log(MSD) has to equal unity within the linear regime. This procedure was also used in a recent work by Frömbgen et al. that implemented a multifidelity Monte Carlo strategy ( section ) for simulating self-diffusion coefficients of liquid water from MD simulations. By exploiting the dependence of the self-diffusion coefficient on the simulation box size, the authors constructed a multifidelity model hierarchy based on differently sized simulation boxes. It was shown that for cubic simulation boxes, in terms of the mean squared error of the predicted self-diffusion coefficient as a function of the computational budget, any combination of models is superior to solely using a single high-fidelity model. A multifidelity Gaussian process approach (see sections and ) to calculating shear viscosities of small to medium-sized liquid aliphatic alkanes and alcohols, based on the autocorrelation function of the stress tensor, was proposed by Fleck et al. In that work, a few high-fidelity experimental data points is combined with a larger number of low-fidelity data points from MD simulations to train a GP model and predict shear viscosities at various thermodynamic state points (temperature, pressure, density) with high accuracy. 4.2.3. System Size, Time, and Length Scale: Enhanced Sampling and Reweighting Methods The predictions from molecular simulations are often limited in terms of sampling relevant sizes or times of a particular system at interest, occurring frequently in the field of biomolecular simulation. UQ methods targeting such applications are addressed in the following. It should be noted that, as the section on free energy calculations partially overlaps with this section, works relevant to the former topic are discussed in section . In scenarios where only a few samples of QoIs can be calculated for computational reasons, such that the measured probability distribution does not allow for computing meaningful averages and standard deviations, Bayesian bootstrapping can be a promising approach. Mostofian and Zuckerman apply this method to study the rate constants of protein folding, simulated by means of MD, and conclude that it is well-suited method for high log-variance datasets. Xia and Wei also address the challenge of modeling protein folding, specifically by introducing molecular nonlinear dynamics, an approach in which the atoms are represented by chaotic oscillators. The authors then show that folding a protein reduces the chaoticity of the system and finally, build on chaos to devise an algorithm for thermal protein uncertainty quantification. Russo et al. proposed a trajectory reweighting scheme allowing for prediction of various observables from sets of short, unbiased trajectories, and applied it to a 1 μs tryptophan cage folding trajectory. 4.2.4. Free Energy Calculations Free energy calculations are at the heart of understanding various kinds of chemical processes, for example reactions, protein–ligand binding, solvation of particles, and wetting at liquid/solid interfaces. The principles of free energy calculations are rooted in statistical mechanics and derived from the density of states or the partition function. In practice, however, in silico chemistry does not seek to calculate absolute free energy of a system, but the difference in free energy of a system in two states, where 0 is the initial or reference state, and 1 is the final or target state. Chemically, the transition of a system from state 0 to 1 may, for example, correspond to (i) transferring a molecule from the gas phase into a solvent, enabling the calculation of the solvation free energy, or (ii) bringing a protein in touch with a ligand, giving rise to the binding free energy. For a deeper introduction to this matter, we refer the reader to the comprehensive book by Chipot and Pohorille, or existing reviews. − From a bird’s-eye perspective, the most prominent approaches to calculating free energies are based on the free energy perturbation method, thermodynamic integration, or probability distributions and histograms , and will be briefly summarized in the following: Free energy perturbation (FEP) relates the free energy difference between two states to the ensemble average of the exponential of their energy difference. A key advantage of the method is that the energy of both states 0 and 1 is solely evaluated in one of the phase spaces, conventionally in that of 0, and thus, renders the method computationally efficient. It is, however, only suitable when the two states are very similar (i.e., their phase spaces overlap strongly), and becomes unreliable when this overlap is poor, as the exponential averaging introduces large statistical noise. To correct for a poor overlap in phase space between two states, a number of intermediate states can be introduced, ensuring sufficient overlap between adjacent states. Thermodynamic integration (TI) circumvents the overlap issue by introducing a continuous coupling parameter λ which smoothly transforms the system from state 0 to state 1 via a hybrid potential V . The free energy difference Δ F is computed by integrating the average of the derivative of the potential energy with respect to the coupling parameter: Δ F = ∫ 0 1 ⟨ ∂ V ( λ ) ∂ λ ⟩ λ d λ 89 Note that similar to the FEP method, a sufficient phase space overlap between the states or the consideration of intermediate states is required. TI is more robust than FEP across a broader range of transformations but requires careful choice and spacing of coupling points to ensure smooth convergence and accurate numerical integration. In this context, alchemical free energy calculations should be mentioned, that, referring back to ancient alchemists who sought to transform lead into gold, involve transforming (mutating) a chemical species into an alternate one. Thus, the difference in the potential stems from the different chemical species in state 0 and 1. Note that the majority of works referenced in the following is based on TI. Histogram-based methods (or reweighting methods), including umbrella sampling and the weighted histogram analysis method, are particularly useful for systems involving slow transitions along a well-defined reaction coordinate. These methods apply biasing potentials to enhance sampling in poorly explored regions of phase space, generating overlapping histograms of sampled configurations across several windows. While offering a powerful tool, such methods require prior knowledge of a suitable reaction coordinate and can be computationally demanding to achieve sufficient overlap. In practice, the multistate Bennet acceptance ratio (MBAR) estimator by Shirts and Chodera is often used in works that address UQ in the context of histogram methods. It should be noted that, when discussing UQ in terms of free energy calculations, existing works can be roughly allocated to two main categories, namely UQ with respect to (i) the free energy method itself (e.g., those laid out above) and (ii) the sampling of the phase space, which is inevitable to accurate free energy predictions. The second category clearly has an overlap with section , and thus, all works that are relevant to free energy calculations are discussed below. Many studies with respect to the uncertainty quantification in free energy calculations, especially on alchemical free energies, were published by Coveney and co-workers. ,− In ref , Bhati et al. introduce UQ for TI-based alchemical free energy calculations by a method called thermodynamic integration with enhanced sampling (TIES). In traditional TI approaches, the derivative of the potential ( eq ) is typically obtained from a single MD simulation for each discrete window of the coupling parameter, while in TIES, an ensemble of MDs is performed for each window. This procedure has two advantages: first, by performing multiple replica simulations, sampling of the phase space is greatly enhanced and second, the ensemble provides access to statistical quantities and hence, UQ. , The authors have found that ∂V / ∂λ behaves like a Gaussian random variable and therefore, treat the integral as a stochastic one, which is solved numerically and provides a variance of Δ F based on the bootstrapped standard error in each λ window. Combining TIES with several enhanced sampling methods, , Bhati et al. calculate free energies for seven protein–ligand pairs, and (i) find that the accuracy of the results is improved by an ensemble of replica simulations compared to performing a single very long simulation, while (ii) they investigate the performance of the enhanced sampling schemes with respect to their computational effort. Recently, an updated TIES version was introduced by Bieniek et al., that uses a new protocol to superimpose the two transformed species (often ligands) in alchemical free energy calculations. That, for example, allows for aromatic rings to be partially superimposed, which reduces the size of the alchemical region and hence, the TIES error, while improving the precision of the predicted free energies. For completeness, we also mention the work by Wade et al. and Bhati et al. here, which apply the aforementioned UQ methods for alchemical free energies two large benchmark studies on ligand transformations using different MD simulation packages in conjunction with TI and FEP methods. Further, Vassaux et al. performed a large-scale study on simulating the binding free energy of the bromodomain-containing protein 4 and the tetrahydroquinoline ligand with approximated continuum solvent. The authors investigated, how uniformly distributed uncertainty in a set of 14 input parameters, including the simulation temperature, pressure and duration of equilibration/production runs, is propagated through MD simulations to estimate the binding free energy of the protein–ligand pair. Additionally, the influence of the random velocity seeds is investigated by means of replica simulations. Uncertainty is propagated through the simulations by applying a dimension-adaptive version of stochastic collocation , (refer to section ). It is found that the variance in the predicted binding free energies is primarily dominated by six parameters, including temperature, box size and barostat settings. Also, the authors show that the uncertainty is “damped” during the simulation, meaning that the variation of the computed binding free energies around the mean are smaller than for the input parameters. In a follow-up work, Edeling et al. conducted a conceptually related sensitivity analysis, investigating the influence of uncertainties in force field parameters on the stiffness of a polymer material and the binding free energies of two protein–ligand systems, but propagating uncertainties by a Gaussian process method. Note that studies on sensitivity analysis of MD parameters not related to free energies are discussed in section . On the notion of using histogram methods for free energy calculations, the works by Schieber et al. and Bauer and Gross investigate phase diagrams of solid benzene as well as argon, methanol, and water, respectively. Both studies employ the MBAR histogram method and estimate uncertainties in their predictions of the phase coexistence lines using bootstrapping ( section ) or standard error propagation. 4.2.5. Coarse-Grained Simulations The importance of coarse-grained (CG) simulations has been discussed by Noid in a recent perspective, where the author explains that “by representing systems in reduced detail, CG models provide the necessary computational efficiency for simulating length- and time-scales that remain far beyond the scope of conventional atomically detailed simulations”. As the aforementioned “reduced detail” often come along with approximations and (over)simplifications, CG simulations are highly concerned with uncertainty quantification. Molinero and co-workers have worked on UQ for CG simulations of anion exchange membranes , (see section ). Naturally, CG force fields for water are in high demand as well, and thus, it comes as no surprise that these have been investigated by means of UQ. Zavadlav et al. developed a Bayesian framework for the data-driven selection of CG water models for a given application. Most importantly, the authors find that flexible water models do not provide improved accuracy, and consequently, rigid water models are recommended for computational efficiency. Also, the charge distribution within the model is found to be another key quantity, with more complex charge models delivering the best accuracy. A central issue, not only in the CG community but also in the general field of force-field-based modeling, is transferability, i.e., the reusability of force fields beyond the specific chemical system for which they were developed. In their methodological work, Patrone et al. approach this issue for CG simulations specifically by deriving an iterative Bayesian correction algorithm to recalibrate the forces obtained through force-matching against thermodynamic state points. 4.3. Uncertainties in Static Properties and in Electronic Structure Applications 4.3.1. Chemical Reaction Networks A chemical reaction network is a set of interconnected chemical reactions that describe how different species transform into one another. It represents the overall reaction system as a network of reactants, products, and intermediates linked by elementary steps. While handling prediction errors in species-specific properties is a challenge in itself, the situation becomes considerably more complex when the target property depends on a network of interconnected species. These species are coupled through elementary reaction steps, meaning that an inaccuracy in any activation energy across the network may affect the predicted concentration of the desired product over time. This effect can be substantial given that the concentration is a function of rate constants, each of which depends exponentially on its associated energy barrier. Thus, even minor deviations in activation energies can cascade through the reaction network, leading to substantial discrepancies in product concentrations. On the other hand, uncertainties in activation energies of different reaction steps are often correlated, in particular if these energies are obtained from quantum-chemical calculations. Reiher and co-workers demonstrated this cascading behavior for a small model network of the formose reaction (6 species, 10 elementary steps). By generating an ensemble of long-range-corrected PBE0 functionals via Bayesian statistics (cf. section and ref , where BEEFs were combined with reaction networks to study dry reforming of methane in the context of heterogeneous catalysis), distributionsinstead of single valuesof activation energies and hence rate constants were obtained. The resulting species concentrations span up to 23 orders of magnitude in time. This early study on uncertainty quantification for chemical reactions demonstrates how critically sensitive network properties, such as time-dependent concentrations, can be to species-specific energy uncertainties. In a follow-up study, Proppe and Reiher addressed an issue that was later summarized as follows: “The exquisite details that [network] exploration algorithms can generate for any chemical process raise the question of reliability as, in general, no or very little experimental or theoretical reference data will be available. Consequently, uncertainty quantification will become a crucial part of the whole exploration process.” To still demonstrate the effect of uncertainty-equipped activation energies for significantly larger reaction networks, KiNetX and the helper tool AutoNetGen were developed. The latter generates artificial, chemistry-mimicking reaction networks and equips each node (species) of a network with an ensemble of activation energies sampled from a predefined covariance matrix. Each diagonal element of the covariance matrix represents the variance of the corresponding energy barrier, while each off-diagonal element constitutes the correlation between a pair of activation energies. Subsequently, KiNetX propagates this energy uncertainty by solving the underlying ensemble of kinetic models and estimates the kinetic relevance of each species based on its maximum rate of formation. The software then reduces the reaction network by identifying and eliminating kinetically irrelevant vertices and edges through a systematic hierarchy of flux analyses, thus achieving a more compact network representation of the reaction mechanism. Eventually, KiNetX performs a sensitivity analysis to distinguish between model parameters (here, activation energies) that significantly impact the outcomes (here, time-dependent species concentrations) and those that do not. This classification is particularly valuable as it helps in determining the quality of energy barriers obtained from efficient, semiaccurate quantum-chemical methods. If the uncertainty in the concentrations of certain species proves too large to make reliable conclusions about specific aspects of the reaction network, the result of the sensitivity analysis will highlight which activation energies are most critical. These critical parameters can then be re-evaluated using a more sophisticated electronic-structure method, ensuring a higher level of accuracy where needed. In an actual exploration scenario, unlike the artificial networks generated by AutoNetGen , one would leverage uncertainty from the very beginning to guide the exploration process. So far, the discussion has focused on uncertainty in kinetic parameters, but when exploring unknown reaction networks, this uncertainty plays an even more central role. Rather than applying it retroactively, uncertainty actively informs the exploration, helping to prioritize which reactions and species to investigate further. This approach ensures that the exploration is both efficient and comprehensive, reducing the risk of overlooking critical pathways due to minor errors in activation energies. This method has been applied in studies using the Reaction Mechanism Generator (RMG) software, where the propagation of uncertainty through reaction networks was demonstrated for carbon dioxide methanation on Ni(111) , and exhaust gas oxidation on Pt(111). By analyzing ensembles of reaction networks, these studies quantified how uncertainty in activation energies can affect the entire network, ultimately influencing predicted product concentrations. Such analyses highlight the necessity of incorporating uncertainty at every stage of the exploration process, ensuring a more reliable understanding of the network’s behavior. While the ensembles of the RMG-generated networks were obtained from a rule-based approach, which requires well-established reaction families and predefined graph rules, the Kinetics-Interlaced Exploration Algorithm (KIEA) , and Yet Another Kinetics Solver (YAKS) methods circumvent these limitations by utilizing exclusively first-principles-based methods for the network exploration. All of the above examples were based on the implicit assumption of mass action and, hence, reaction networks with deterministic kinetics. When stochastic effects are taken into account, the complexity of uncertainty quantification increases significantly. In reaction networks where stochastic kinetics is more appropriate, uncertainty arises not only from kinetic parameters but also from the inherent randomness of reaction events. This shift requires methods that can handle both types of uncertainty simultaneously. For example, Navarro Jimenez et al. developed a global sensitivity analysis approach that effectively disentangles these different sources of variability, offering deeper insight into the system’s behavior under stochastic conditions. Understanding variability in reaction networks with stochastic kinetics requires not only accurate models but also carefully designed experiments that can maximize the information gained about the system. Ruess and co-workers developed a framework based on Fisher information to optimize experimental setups for identifying key system parameters. By leveraging the inherent stochasticity in reaction rates, their approach provides a means to quantify and reduce uncertainty more effectively. This method underscores the importance of experimental design in stochastic systems, ensuring that the variability in system behavior is accurately captured and used to refine model predictions. Such strategies can be invaluable in cases where stochastic noise dominates. 4.3.2. Reactivity Parameters Instead of focusing on entire reaction networks, the quantification of uncertainties can also be performed at the level of individual elementary reactions. Such uncertainties in individual rate constants can later be propagated through a network of these elementary reactions. Particularly suitable for such quantification are the highly accurate experimental rate constants from Mayr’s laboratory, which have been measured and documented over the past decades for thousands of electrophile–nucleophile reactions. To enable efficient estimation for arbitrary combinations of reaction partners, in 1994 Mayr and Patz developed a three-parameter equation: log k = s N ( N + E ) 90 where E denotes the electrophilicity and N the nucleophilicity, while s N is an additional nucleophile-specific sensitivity factor. The two nucleophilic parameters are also solvent-dependent. The resulting rate constants are defined at a reference temperature of 20 °C. In a 2022 study, Proppe and Kircher were the first to determine uncertainties for such reactivity parameters derived from experimental results, which, through propagation, enable the determination of uncertainties in the corresponding rate constants. In a 2023 review by Vahl and Proppe, various strategies were presented for estimating such reactivity parameters using quantum-chemical calculations. After Proppe and co-workers demonstrated in 2024 how machine learning can be employed to generate quantum-chemically derived reactivity parameters in a high-throughput manner, they applied UQ for the first time in 2025 to determine the uncertainty of a reactivity parameter, namely, the electrophilicity of carbon dioxide. 4.3.3. Theoretical Spectroscopy The calculation of molecular spectra and of spectroscopic properties is another important application of in silico chemistry, in particular of quantum-chemical methods. In many cases, the comparison between computations and experiment can serve as a tool for assigning spectral features and for elucidating molecular structures. This makes it even more pressing to quantify the uncertainties in such calculations in order to confidently draw such conclusions. While the sources of uncertainties in the ground-state energies from quantum-chemical calculations that were discussed in section also apply to quantum-chemical calculations of spectra, the latter usually requires additional steps that can introduce uncertainties. Some of these will be discussed in the following for different types of spectroscopy. Structural Sensitivity of Calculated Spectra Like all quantities of interest in computational chemistry, calculated spectra are subject to many sources of uncertainty. In the quantum-chemical calculation of molecular spectra, one important source of uncertainty is the choice of the molecular structure that is used in these calculations. This is particularly relevant if the comparison of experimental and calculated spectra is used to infer structural parameters, for instance when studying the structure of photosystem II and related model complexes with EXAFS. − Jacob and co-workers proposed a general framework for quantifying uncertainties stemming from structural distortions of the input structure in theoretical spectroscopy. As quantity of interest, they chose the (discretized) spectral intensity as a function of the energy (usually obtained by applying an empirical broadening to the calculated individual transitions) instead of individual transitions themselves. This choice is motivated by the observation that for many types of spectroscopy, there are various close-lying transitions which cannot be resolved in experiment. Because of the large number of degrees of freedom for molecular structures (3 N – 6 for nonlinear molecules, where N is the number of atoms), assessing the structural sensitivity in a local sensitivity analysis is a high-dimensional problem. Bergmann et al. showed that by means of a principal component analysis in combination with a high-dimensional model representation (Sobol’ expansion), it becomes possible to construct a low-dimensional surrogate model, that can then be used to efficiently analyze the propagation of uncertainties due to structural distortions, and to assign corresponding error bars to the calculated spectra. This general framework has been applied in the calculation of X-ray emission, UV/vis, and infrared spectra. For related work on surrogate models for describing the structural sensitivity of X-ray spectra, see refs and . Vibrational Spectroscopy In computational vibrational spectroscopy, it is common practice to apply scaling factors to calculated vibrational frequencies, that correct for errors in the underlying quantum-chemical method as well as additional approximations (such as the commonly applied harmonic approximation). − These scaling factors are determined using a linear fit of calculated frequencies to experimental reference data. Irikura et al. applied uncertainty quantification to scaling factors for harmonic vibrational frequencies and to vibrational zero-point energies by employing the standard deviation of this linear fit. Johnson et al. extended this work to scaling factors for anharmonic frequencies calculated with second-order vibrational perturbation theory. A case study for the X3LYP functional is presented in ref . A discussion of methodological aspects can be found in refs − . Parks et al. applied the BEEF-vdW xc functional (see section ) to calculate harmonic vibrational spectra, and propagated the uncertainties provided within the BEEF framework to the harmonic frequencies. They present an in-depth analysis of the resulting uncertainties, and find that certain types of vibrations (e.g., bending and torsional modes) are prone to higher uncertainties. They further point out that in many cases, in particular for intermolecular complexes, the ensembles obtained for the vibrational frequencies are non-Gaussian. The calculation of anharmonic vibrational spectra generally requires solving the nuclear Schrödinger equation on the full potential energy surface (for reviews, see refs − ). Such calculations are extremely challenging, and the accuracy of the final vibrational transition energies and possibly intensities is intricately determined by multiple sources of errors, such as the accuracy of the underlying potential energy surface (including its representation in suitable coordinates − ) and the solution of the nuclear Schrödinger equation using vibrational correlation methods. − Addressing the latter, Larsson established error bound for the individual vibrational energy levels obtained with a given potential-energy surface of acetonitrile. To this end, they performed an extrapolation with respect to the bond dimension D in their tree tensor network calculations. These uncertainty estimates do, however, not account for uncertainties in the potential energy surface. Some steps toward quantifying the uncertainties related to both the potential energy surface and the solution of the nuclear Schrödinger equation have been made by König and Christiansen. UV/Vis Spectroscopy and Photodynamics There are only a few studies that are related to the quantification of uncertainties in quantum-chemical calculations of electronic excitations. These consider the uncertainties due to the quantum-chemical approximations with respect to experimental data or to high-level computational reference data. In an early study, Edwards et al. considered the prediction of vertical ionization potentials (IPs) of small water clusters with a double-hybrid xc functional. For the water dimer, trimer, tetramer, and pentamer, they determined intervals for the fraction of Hartree–Fock exchange and of MP2 correlation that lead to predictions that are consistent with high-level reference data within the Bound-to-Bound Data Collaboration framework. − Subsequently, they use these intervals for those two parameters of the xc functional to establish an error interval for the vertical IP of different isomers of the water hexamer. The accuracy of TD-DFT predictions of the energies of electronically excited state as well as of the corresponding transition moments widely varies with the employed xc functionals, and the accuracy of different xc functionals is highly dependent on both the considered molecules and on the nature of the relevant excited states. Avagliano et al. employed a huge dataset containing the five lowest singlet excited states for over 20,000 molecules to assign a score to each of 38 xc functional. This score combines the uncertainties in the one-electron transition density matrix, the excited states energies, and the transition dipole moments. They then trained a graph attention neural network to predict these scores from the molecular structure, and to recommend the xc functional with the best expected performance for this specific molecule. In simulations of the nonadiabatic dynamics following photoexcitations, the outcome of the simulations can sensitively depend on the underlying ground and excited state potential energy surfaces. Jíra et al. investigated the photochemistry of cis -stilbene using trajectory surface hopping methods and found a large dependence of the quantum yields for two different products on the underlying electronic structure method. They suggest the use of a biasing potential, which allows one to efficiently quantify this sensitivity. X-ray Spectroscopy In many cases, the calculation of spectra involves additional approximations on top of the quantum-chemical methods that are employed. One such an example is computational X-ray spectroscopy, where one commonly applies the core–valence separation approximation, in which excitations from core orbitals are separated from excitations from valence orbitals. To assess the error arising from the approximation, Herbst and Franson developed a postprocessing step for ADC(2) calculations employing the core–valence separation. They apply Rayleigh quotient iterations to refine the resulting eigenvectors, and to assess the corresponding error in the eigenvalues. X-ray absorption spectra mainly probe the local environment of the absorbing atom. Therefore, machine learning models that are trained to predict them from descriptors of this local chemical environment have been developed. Ghose et al. put forward uncertainty quantification for such models by training neural network ensembles (also refer to section ), considering the N, O, and C K-edge spectra of the small organic molecules in the QM9 dataset. Conversely, neural networks can be trained to classify which structural elements are present in a molecule form its X-ray absorption spectrum. Again, a neural network ensemble can be used to provide uncertainties for such a classifier. For the Fe K-edge X-ray absorption spectra, Verma et al. trained deep neural networks, and estimated the uncertainty of their predictions using bootstrap resampling (see section ). Mössbauer Spectroscopy Most studies on theoretical Mössbauer spectroscopy focus on the isomer shift (via the contact density) and/or the quadrupole splitting (via the electric-field gradient), both of which are spectral features caused by electric hyperfine interactions. Proppe and Reiher were the first to study the problem from a UQ perspective and applied a range of statistical tools, including bootstrapping ( section ) and Bayesian inference ( section ), to estimate confidence intervals for isomer shifts derived from DFT-computed contact densities of iron complexes. These confidence intervals take into account both residuals (model error) and fluctuations in regression coefficients (parameter uncertainty). They found that the average predicted confidence interval is significantly larger than the average experimental uncertainty. More complex (i.e., nonlinear) fitting functions could not resolve this issue. An outlier analysis based on the jackknife-after-bootstrapping method , revealed that five of the 44 studied complexes are responsible for the disagreement. After removal of these outliers, the average predicted confidence interval became representative of the average experimental uncertainty. As pointed out later by Krewald and co-workers, a valid UQ model is not enough to obtain reliable predictions of Mössbauer parameters. If the electronic structure is qualitatively wrong, e.g., due to non-negligible multireference character, predictions of Mössbauer parameters will be compromised by biased contact densities and electric-field gradients. It is therefore crucialnot only for this type of applicationto validate the input fed into a UQ model or any other kind of prediction model. Referring to the example above, one would not know that the five outliers were actually outliers without further analysis if they had not been part of the regression procedure. The UQ method developed by Proppe and Reiher for computational Mössbauer spectroscopy was refined in the work by Krewald and co-workers and has been applied, for instance, to study how myoglobin-catalyzed azide reduction proceeds via an anionic metal amide intermediate. 4.3.4. Quantum Cluster Equilibrium Method In the framework of the quantum cluster equilibrium (QCE) theory, the liquid phase is described by an ensemble of interacting gas phase molecular clusters. Thermodynamic properties of the bulk phase are accessible by combining static quantum-chemical calculations of an ensemble of clusters, differing e.g., in the number of molecules, conformation, composition, etc., with simple statistical mechanics and subsequent cluster weighting or population analysis. − By assuming that the clusters are in equilibrium with each other and relying on mass conservation, the total partition function can be expressed, providing access to thermodynamic functions, such as the energy, entropy, and enthalpy. The QCE theory has been applied to calculate thermodynamic quantities of many condensed phase systems, − and cluster weighting was further employed to obtain spectroscopic − data or conductivities , in the condensed phase. In particular, vibrational circular dichroism spectra, which strongly depend on the conformational sampling could be accessed. Recently, a multicomponent QCE theory has been developed, allowing to describe any type of liquid mixtures. QCE was also used to study uncertainty quantification by Blasius et al. In that work, the dependence of the vaporization enthalpies and entropies of several organic liquids on the uncertainties in the experimental QCE input data was investigated. The vaporization enthalpies and entropies showed a smooth dependence on changes in the reference density and boiling point, and while the density showed little influence on the vaporization thermodynamics, variations in the input boiling point had a larger effect on the vaporization enthalpy, but only little effect on the vaporization entropy. Next, a quantification of uncertainty in thermodynamic functions that originates from inaccuracies in the experimental reference data via the Gauss–Hermite estimator was carried out, finding an uncertainty of 30.95 kcal mol –1 for ( R )-butan-2-ol. 4.3.5. Machine Learning of Static Properties In recent years, a variety of approaches have been developed to quantify and validate uncertainty in machine learning models for molecular property prediction. Ensemble-based methods (also refer to MLIPs in section ) and calibration-oriented methods form one major direction in this context. Busk et al. introduced a framework for calibrated uncertainty estimation in message passing neural networks (MPNNs) trained on QM9 and PC9. They demonstrated that recalibration via isotonic regression improves the reliability of aleatoric and epistemic uncertainty estimates, with validation on out-of-distribution data using the expected normalized calibration error (ENCE). Similarly, Gruich et al. compared k -fold ensembling, Monte Carlo dropout, and evidential regression for crystal graph convolutional neural networks (CGCNNs) trained on the Open Catalyst 2020 dataset, showing that evidential regression combined with scalar recalibration yields particularly trustworthy uncertainty estimates for adsorption energies in catalysis. Both studies highlight the importance of calibration for ensuring meaningful uncertainty estimates in chemical machine learning. A number of works provide systematic analyses and benchmarking studies across different UQ techniques. Heid et al. evaluated ensembles, mean–variance estimation, and conformal prediction using various neural network architectures, including d-MPNNs and SchNet, focusing on the decomposition of epistemic uncertainty into bias and variance contributions. Hirschfeld et al. extended this perspective by benchmarking ensemble-, mean-variance-, distance-, and union-based UQ approaches on multiple datasets, finding that dataset-specific factors strongly influence UQ performance. These studies underscore that no single UQ method is universally superior and that calibration quality and data characteristics are crucial determinants of reliability. Several works explored Bayesian or probabilistic formulations of uncertainty estimation (see section ). Janet and Kulik applied Monte Carlo dropout and Gaussian-process-based variance estimation to neural networks predicting DFT-derived properties of transition metal complexes, demonstrating that such uncertainty estimates help identify unreliable predictions. Ryu et al. proposed a Bayesian graph convolutional network (GCN) to separate epistemic and aleatoric contributions, showing improvements in model reliability for datasets covering bioactivity and photovoltaic properties. These methods illustrate how Bayesian inference can naturally incorporate data and model uncertainty in chemical learning tasks. Latent-space and distance-based methods have been developed as computationally efficient alternatives. Janet et al. introduced a latent space distance metric for neural networks that correlates well with predictive error while outperforming ensemble and dropout methods in calibration. Zhang et al. further applied similarity-based pairing in Siamese neural networks to derive variance-based uncertainty measures, confirming that confidence scores derived from pairwise similarity correspond to actual prediction reliability. Together, these works demonstrate that representations learned by neural networks can provide valuable internal indicators of prediction uncertainty. In addition to these approaches, uncertainty quantification has been integrated into Gaussian process models (see also section ). Musil, Ceriotti, and co-workers proposed a subsampling-based estimator for Gaussian process regression (GPR) that reproduces reliable UQ at lower computational cost compared to standard GPR variance estimation. Wollschläger et al. extended this idea by introducing the localized neural kernel (LNK), a Gaussian-process-inspired extension to graph neural networks, achieving improved calibration and out-of-distribution detection in molecular force-field modeling. These methods illustrate how kernel-based approaches can provide interpretable and computationally tractable uncertainty estimates. Several studies focused explicitly on evaluating or validating UQ metrics ( section ). Rasmussen et al. systematically compared different UQ validation metrics for chemical ML and concluded that error-based calibration analysis offers a more reliable assessment of UQ performance than traditional correlation-based measures. Scalia et al. performed a large-scale comparison of deep ensembles, Monte Carlo dropout, and bootstrapping across chemical and biochemical datasets, finding that ensemble-based approaches generally achieve superior calibration. Both works emphasize that the choice of evaluation metric is crucial for assessing uncertainty quality. Finally, a number of specialized approaches extend uncertainty quantification beyond global property prediction. Tynes et al. introduced pairwise difference regression (PADRE), which reformulates regression tasks into pairwise comparisons, providing intrinsic uncertainty estimates that improve candidate selection. Yang and Li developed an atom-based uncertainty method combining deep ensembles with posthoc calibration, enabling atomic-level attribution of uncertainty and improving interpretability. A promising application of UQ in ML of static properties are multifidelity methods, in which high-quality training data is combined with cheaper and less accurate training data to achieve the accuracy of the costlier level. , Such methods have successfully been employed for ML models of atomization energies, band gaps, , potential energy surfaces, and of spectroscopic properties. Overall, these studies collectively reveal a landscape of complementary approaches to uncertainty quantification in molecular machine learning, spanning ensemble-based calibration, Bayesian inference, latent space metrics, kernel-based methods, and specialized architectures designed for interpretability or data efficiency. 5. Conclusion In this review, we have presented an overview of applications of UQ for in silico chemistry. Mathematical tools developed in applied mathematics for quantifying uncertainties in computational sciences have been increasingly adopted in chemistry in recent years. The methods and applications discussed in this review have the potential to transform in silico chemistry by providing reliable uncertainties for a wide range of quantities of interest. The gained insights will increase the reliability of predictions from in silico methods in comparison to experiments. A particularly important goal of UQ in quantum chemistry is to establish uncertainty estimates in (relative) energies of quantum-chemical calculations for specific molecules. Many different sources of uncertainty in such calculations have been tackled using a wide range of methods. Established tools, such as basis set extrapolation schemes, indicators of multireference character, or automated schemes for the selection of the active space in multireference calculations, can be reinterpreted as uncertainty estimates. However, they will require further validation to arrive at routinely applicable tools. Similarly, developments such as xc functionals, including uncertain parameters, still need to find their way into routine applications of quantum-chemical program packages. We expect particularly large benefits from further developments of UQ in quantum chemistry in the field of multilevel and multifidelity methods, where uncertainty estimates can play a decisive role in guiding the partitioning of complex chemical systems into subsystems. Moreover, selecting suitable methods for each subsystem in such a way that a quantity of interest (e.g., an enzymatic reaction energy or spectroscopic properties of a chromophore embedded in a complex environment) is as accurate as necessary for a specific chemical problem at hand. With regard to MD simulations, two main targets have gained notable attention for developing and applying UQ, namely simulation parameters and trajectory analysis. Various mathematical methods, many of them in the context of local sensitivity analysis, have been employed to investigate the propagation of uncertainties in MD input parameters, covering both force field parameters as well as other parameters such as the system size, simulation time, or temperature. Efforts have been made to transform traditional approaches of finding optimal sets of parameters, which often lack reproducibility and a systematic estimation of uncertainties, into UQ-guided, mathematically sound workflows. Additionally, UQ has been incorporated into postprocessing routines for MD simulations, particularly in the context of free energy calculations, and correlation functions. In that context, replica simulations, in conjunction with methods that enhance or improve sampling of the phase space, thermodynamic functions and other QoIs, have proven instrumental for yielding accurate and precise simulation predictions with interpretable uncertainties. We anticipate that alongside the growing interest of UQ for MD simulations, more advanced UQ tools, such as global sensitivity analysis and multifidelity techniques will be developed for and applied to complex QoIs. In ML, UQ has led to a diverse set of methodological approaches, and the field is still evolving and far from being uniformly established. In chemical applications of ML, most importantly in MLIPs, uncertainty estimates are particularly useful for indicating whether the predictions of a ML model are within the domain of its training data. This is in turn heavily exploited in active learning strategies that can reduce the number of data points required for training. As the training data is usually derived from expensive quantum-chemical calculations, the economical use of training data that is achieved by leveraging UQ is particularly relevant for in silico chemistry. A particularly attractive application of ML in UQ for in silico chemistry, with vast potential for new method development, is the use as surrogate models in multifidelity approaches. For example, by training Δ-ML models of the errors of computational predictions (e.g., with respect to experimental data or to high-level computational reference data), it becomes possible to provide error estimates, and to select computational models (e.g., xc functionals) that are most suitable for a specific problem at hand (e.g., the calculation of excitation energies). ML encompasses a broad range of model classes with different parametrizations, from comparatively low-parametric models to highly parametrized deep neural networks. In many contemporary applications, particularly those based on deep neural networks, the extremely high dimensionality of the parameter space renders classical parameter-based uncertainty analyses impractical, motivating alternative approaches that focus on predictive uncertainty and robustness. A particularly critical and still unresolved challenge is the reliable quantification of uncertainty for inputs that differ substantially from the training data, where predictions rely on extrapolation rather than interpolation. For many ML models, especially those used in chemistry and materials science, UQ is more naturally formulated at the level of predictions rather than model parameters, which often lack direct physical meaning. Even if an effective reduction of the parameter space can be achieved, this does not automatically translate into actionable insight for improving the trustworthiness of predictions in practical applications. The reliability of uncertainty estimates in ML models is further conditioned on implicit structural assumptions about the data and models, such as smoothness or differentiability, which are often difficult to validate in realistic settings. Stochastic effects arising from random initialization can lead to differently trained models even when architectures and training data are identical. Ensemble-based approaches are often used to capture this variability at the level of predictions. Summarizing the applications of UQ in in silico chemistry reviewed here, it is apparent that most previous work focused on UQ of a specific source of uncertainty in a specific chemical application. Given the wide range of both, many previous works have come up with ad hoc approaches in each case. Therefore, it becomes hard to spot the universality across these approaches to UQ in in silico chemistry. While this is to some extent a consequence of the wide range of computational methods that are relevant in chemistry, we expect that the next years will see a consolidation of different approaches to UQ in in silico chemistry. The development of user-friendly, open-source UQ software that provide broadly applicable UQ methods with applicability to in silico chemistry could catalyze such a consolidation and incentivize the community to adopt such of-the-shelf approaches to UQ. So far, most of the work reviewed here has tackled the quantification of a single, specific source of uncertainty. The classification of sources of uncertainty in chemical applications and the identification of the most relevant sources of uncertainty remains an important open topic, given the diversity of in silico chemistry. Future work across all areas of in silico chemistry will further need to combine UQ of different sources of uncertainties. On that account, the development of UQ methods that are able to account for systematic error compensation between different sources of errorson which computational chemistry heavily reliesseems to be a particularly promising avenue. While the past years have seen tremendous progress in the development of UQ methods for in silico chemistry, best practices for UQ in different types of applications will need to be established and adopted by researchers. First steps toward establishing UQ as an integral requirement of computational workflows have been started in the molecular dynamics community. We conclude by emphasizing that uncertainty quantification is far more than “just statistics”: it provides rigorous frameworks that have advanced in silico chemistry and will remain central to its future development. Acknowledgments T.F. received financial support from the Studienstiftung des deutschen Volkes (German Academic Scholarship Foundation). T.F., J.D., and B.K. were funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) as part of CRC 1639 NuMeriQS, Project 511713970. E.S. and J.P. acknowledge financial support by the DFG through Project 512350771. J.P. received funding from the Tenure Track Programme established by the Federal Ministry of Research, Technology and Space (BMFTR). J.D. received support from the Hausdorff Center for Mathematics (HCM) through the DFG under Germany’s Excellence Strategy, Project 390685813. Some of the authors used DeepL and ChatGPT during the writing of the manuscript; they carefully reviewed and revised the AI-generated content and take full responsibility for the final publication. Biographies Tom Frömbgen is a Ph.D. student at the University of Bonn, where he works in the group of Barbara Kirchner at the Mulliken Center for Theoretical Chemistry. He holds a B.Sc. and M.Sc. in Chemistry from the University of Bonn. His research focuses on computational modeling of the liquid phase, including its structure, dynamics and vibrational spectroscopy, by means of force-field-based and ab initio molecular dynamics simulations. In 2024, he was awarded a Ph.D. scholarship by the German Academic Scholarship Foundation (Studienstiftung) to investigate the application of uncertainty quantification in the context of computational bulk phase vibrational spectroscopy. Elizaveta Surzhikova is a Ph.D. student in the group of Jonny Proppe at TU Braunschweig. She studied computer science at the University of Leipzig (B.Sc.) and at TU Braunschweig (M.Sc.). In her Ph.D. project, funded by the German Research Foundation (DFG), she is developing data-driven methods for more efficient chemical property prediction and molecular-design purposes, specifically surrounding carbon capture. Jürgen Dölz is a professor for scientific computing at the Institute for Numerical Simulation of the University of Bonn. After his Ph.D. at the University of Basel, he worked as a postdoc at the Technical University of Darmstadt and was an assistant professor at the University of Twente. His research group works on uncertainty quantification, high-dimensional approximation, and machine learning for parametrized partial differential and operator equations as well as efficient methods for nonlocal operators. Jonny Proppe is a tenure-track professor of data-driven chemistry at TU Braunschweig. He studied chemistry at the University of Hamburg before moving to Zurich, where he obtained a Doctor of Science degree from ETH Zurich. After postdoctoral stays at Harvard University, the University of Toronto, and the University of Göttingen, he assumed his current position in October 2021. His team develops data-driven methodsincluding active learning, generative modeling, and explainable AIfor the design of molecules and new chemical transformations in the contexts of carbon capture and utilization as well as drug design. Barbara Kirchner completed her Ph.D. on classical molecular dynamics simulations at the University of Basel. Before taking up her current position at the Mulliken Center at the University of Bonn as one of the chairs of Theoretical Chemistry, she held a chair of Theoretical Chemistry at the University of Leipzig. She worked as a postdoctoral researcher in various institutes on ab initio molecular dynamics simulations and as a visiting scientist in Brisbane, Australia. She received the Ruth M. Lynden-Bell Award for A Trajectory in Ionic Liquids in 2019. She is currently Deputy Editor for the Journal of Physical Chemistry and Vice Dean for research and promotion of early career researchers. Her work is diverse, ranging from understanding liquids and solvents, through intermolecular forces and processes in the condensed phase, to quantum-chemical analysis of interesting molecules and methodological developments mainly for analyzing trajectories. Christoph R. Jacob is a professor of theoretical chemistry at TU Braunschweig. He studied chemistry and mathematics at the Universities of Marburg and Karlsruhe. After a research stay at the University of Auckland, he obtained his Ph.D. at Vrije Universität Amsterdam. He was a postdoc at ETH Zurich and an independent group leader at Karlsruhe Institute of Technology. His research group develops quantum-chemical methods for the description of complex chemical systems ranging from biomolecules to materials and applies these methods to study spectroscopic properties and to design functional chemical systems for energy conversion, catalysis, and drug discovery. The authors declare no competing financial interest. Footnotes a Some mathematical subtleties have to be taken into account, , but these are of no importance to most readers. b Taking care of a few mathematical subtleties, most of the things can be generalized beyond the case Ω ⊂ R d . References Frömbgen T., Surzhikova E., Dölz J., Proppe J., Kirchner B., Jacob C. R.. Uncertainty Quantification for In Silico Chemistry. ChemRxiv. 2026 doi: 10.26434/chemrxiv-2025-67ck9-v2. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] von Toussaint U.. Bayesian Inference in Physics. Rev. Mod. Phys. 2011;83:943–999. doi: 10.1103/RevModPhys.83.943. [ DOI ] [ Google Scholar ] Lei Z., Dai C., Hallett J., Shiflett M.. Introduction: Ionic Liquids for Diverse Applications. Chem. Rev. 2024;124:7533–7535. doi: 10.1021/acs.chemrev.4c00291. [ DOI ] [ PubMed ] [ Google Scholar ] Shiflett M. B., Scurto A. M.. Ionic Liquids: Current State and Future Directions. ACS Symp. Ser. 2017;1250:1–13. doi: 10.1021/bk-2017-1250.ch001. [ DOI ] [ Google Scholar ] Wilkes J. S.. A Short History of Ionic Liquidsfrom Molten Salts to Neoteric Solvents. Green Chem. 2002;4:73–80. doi: 10.1039/b110838g. [ DOI ] [ Google Scholar ] Dong K., Liu X., Dong H., Zhang X., Zhang S.. Multiscale Studies on Ionic Liquids. Chem. Rev. 2017;117:6636–6695. doi: 10.1021/acs.chemrev.6b00776. [ DOI ] [ PubMed ] [ Google Scholar ] Izgorodina E. I., Seeger Z. L., Scarborough D. L. A., Tan S. Y. S.. Quantum Chemical Methods for the Prediction of Energetic, Physical, and Spectroscopic Properties of Ionic Liquids. Chem. Rev. 2017;117:6696–6754. doi: 10.1021/acs.chemrev.6b00528. [ DOI ] [ PubMed ] [ Google Scholar ] Podgoršek A., Jacquemin J., Pádua A. A. H., Costa Gomes M. F.. Mixing Enthalpy for Binary Mixtures Containing Ionic Liquids. Chem. Rev. 2016;116:6075–6106. doi: 10.1021/acs.chemrev.5b00379. [ DOI ] [ PubMed ] [ Google Scholar ] Kirchner B., Hollóczki O., Canongia Lopes J. N., Pádua A. A. H.. Multiresolution Calculation of Ionic Liquids. Wiley Interdiscip. Rev.: Comput. Mol. Sci. 2015;5:202–214. doi: 10.1002/wcms.1212. [ DOI ] [ Google Scholar ] Hayes R., Warr G. G., Atkin R.. Structure and Nanostructure in Ionic Liquids. Chem. Rev. 2015;115:6357–6426. doi: 10.1021/cr500411q. [ DOI ] [ PubMed ] [ Google Scholar ] Izgorodina E. I.. Towards Large-Scale, Fully Ab Initio Calculations of Ionic Liquids. Phys. Chem. Chem. Phys. 2011;13:4189–4207. doi: 10.1039/c0cp02315a. [ DOI ] [ PubMed ] [ Google Scholar ] Méndez-Morales T., Carrete J., Cabeza O., Gallego L. J., Varela L. M.. Molecular Dynamics Simulations of the Structural and Thermodynamic Properties of Imidazolium-Based Ionic Liquid Mixtures. J. Phys. Chem. B. 2011;115:11170–11182. doi: 10.1021/jp206341z. [ DOI ] [ PubMed ] [ Google Scholar ] Kempter V., Kirchner B.. The Role of Hydrogen Atoms in Interactions Involving Imidazolium-Based Ionic Liquids. J. Mol. Struct. 2010;972:22–34. doi: 10.1016/j.molstruc.2010.02.003. [ DOI ] [ Google Scholar ] Maginn E. J.. Molecular Simulation of Ionic Liquids: Current Status and Future Opportunities. J. Phys.: Condens. Matter. 2009;21:373101. doi: 10.1088/0953-8984/21/37/373101. [ DOI ] [ PubMed ] [ Google Scholar ] Zahn S., Uhlig F., Thar J., Spickermann C., Kirchner B.. Intermolecular Forces in an Ionic Liquid ([Mmim][Cl]) versus Those in a Typical Salt (NaCl) Angew. Chem., Int. Ed. 2008;47:3639–3641. doi: 10.1002/anie.200705526. [ DOI ] [ PubMed ] [ Google Scholar ] Baca K. R.. et al. Ionic Liquids for the Separation of Fluorocarbon Refrigerant Mixtures. Chem. Rev. 2024;124:5167–5226. doi: 10.1021/acs.chemrev.3c00276. [ DOI ] [ PubMed ] [ Google Scholar ] Zhou T., Gui C., Sun L., Hu Y., Lyu H., Wang Z., Song Z., Yu G.. Energy Applications of Ionic Liquids: Recent Developments and Future Prospects. Chem. Rev. 2023;123:12170–12253. doi: 10.1021/acs.chemrev.3c00391. [ DOI ] [ PubMed ] [ Google Scholar ] Kondrat S., Feng G., Bresme F., Urbakh M., Kornyshev A. A.. Theory and Simulations of Ionic Liquids in Nanoconfinement. Chem. Rev. 2023;123:6668–6715. doi: 10.1021/acs.chemrev.2c00728. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Nordness O., Brennecke J. F.. Ion Dissociation in Ionic Liquids and Ionic Liquid Solutions. Chem. Rev. 2020;120:12873–12902. doi: 10.1021/acs.chemrev.0c00373. [ DOI ] [ PubMed ] [ Google Scholar ] Kříž K., Schmidt L., Andersson A. T., Walz M.-M., van der Spoel D.. An Imbalance in the Force: The Need for Standardized Benchmarks for Molecular Simulation. J. Chem. Inf. Model. 2023;63:412–431. doi: 10.1021/acs.jcim.2c01127. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Brooks C. L. III, Case D. A., Plimpton S., Roux B., van der Spoel D., Tajkhorshid E.. Classical Molecular Dynamics. J. Chem. Phys. 2021;154:100401. doi: 10.1063/5.0045455. [ DOI ] [ PubMed ] [ Google Scholar ] Galvez Vallejo J. L.. et al. Toward an Extreme-Scale Electronic Structure System. J. Chem. Phys. 2023;159:044112. doi: 10.1063/5.0156399. [ DOI ] [ PubMed ] [ Google Scholar ] Smith E. L., Abbott A. P., Ryder K. S.. Deep Eutectic Solvents (DESs) and Their Applications. Chem. Rev. 2014;114:11060–11082. doi: 10.1021/cr300162p. [ DOI ] [ PubMed ] [ Google Scholar ] Hansen B. B.. et al. Deep Eutectic Solvents: A Review of Fundamentals and Applications. Chem. Rev. 2021;121:1232–1285. doi: 10.1021/acs.chemrev.0c00385. [ DOI ] [ PubMed ] [ Google Scholar ] MacFarlane D. R., Forsyth M., Howlett P. C., Pringle J. M., Sun J., Annat G., Neil W., Izgorodina E. I.. Ionic Liquids in Electrochemical Devices and Processes: Managing Interfacial Electrochemistry. Acc. Chem. Res. 2007;40:1165–1173. doi: 10.1021/ar7000952. [ DOI ] [ PubMed ] [ Google Scholar ] Li G., Chen K., Lei Z., Wei Z.. Condensable Gases Capture with Ionic Liquids. Chem. Rev. 2023;123:10258–10301. doi: 10.1021/acs.chemrev.3c00175. [ DOI ] [ PubMed ] [ Google Scholar ] Jeanmairet G., Rotenberg B., Salanne M.. Microscopic Simulations of Electrochemical Double-Layer Capacitors. Chem. Rev. 2022;122:10860–10898. doi: 10.1021/acs.chemrev.1c00925. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Paek E., Pak A. J., Hwang G. S.. A Computational Study of the Interfacial Structure and Capacitance of Graphene in [BMIM][PF6] Ionic Liquid. J. Electrochem. Soc. 2013;160:A1. doi: 10.1149/2.019301jes. [ DOI ] [ Google Scholar ] Unsleber J. P., Reiher M.. The Exploration of Chemical Reaction Networks. Annu. Rev. Phys. Chem. 2020;71:121–142. doi: 10.1146/annurev-physchem-071119-040123. [ DOI ] [ PubMed ] [ Google Scholar ] Wen M., Spotte-Smith E. W. C., Blau S. M., McDermott M. J., Krishnapriyan A. S., Persson K. A.. Chemical Reaction Networks and Opportunities for Machine Learning. Nat. Comput. Sci. 2023;3:12–24. doi: 10.1038/s43588-022-00369-z. [ DOI ] [ PubMed ] [ Google Scholar ] Margraf J. T., Jung H., Scheurer C., Reuter K.. Exploring Catalytic Reaction Networks with Machine Learning. Nat. Catal. 2023;6:112–121. doi: 10.1038/s41929-022-00896-y. [ DOI ] [ Google Scholar ] Simm G. N., Türtscher P. L., Reiher M.. Systematic Microsolvation Approach with a Cluster-Continuum Scheme and Conformational Sampling. J. Comput. Chem. 2020;41:1144–1155. doi: 10.1002/jcc.26161. [ DOI ] [ PubMed ] [ Google Scholar ] Steiner M., Holzknecht T., Schauperl M., Podewitz M.. Quantum Chemical Microsolvation by Automated Water Placement. Molecules. 2021;26:1793. doi: 10.3390/molecules26061793. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Shi X., Lin X., Luo R., Wu S., Li L., Zhao Z.-J., Gong J.. Dynamics of Heterogeneous Catalytic Processes at Operando Conditions. JACS Au. 2021;1:2100–2120. doi: 10.1021/jacsau.1c00355. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Poths P., Lai K. C., Cannizzaro F., Scheurer C., Matera S., Reuter K.. ML-Accelerated Automatic Process Exploration Reveals Facile O-Induced Pd Step-Edge Restructuring on Catalytic Time Scales. ACS Catal. 2025;15:514–522. doi: 10.1021/acscatal.4c06414. [ DOI ] [ Google Scholar ] Lewis-Atwell T., Townsend P. A., Grayson M. N.. Machine Learning Activation Energies of Chemical Reactions. Wiley Interdiscip. Rev.: Comput. Mol. Sci. 2022;12:e1593. doi: 10.1002/wcms.1593. [ DOI ] [ Google Scholar ] Roet S., Daub C. D., Riccardi E.. Chemistrees: Data-Driven Identification of Reaction Pathways via Machine Learning. J. Chem. Theory Comput. 2021;17:6193–6202. doi: 10.1021/acs.jctc.1c00458. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Behler J., Parrinello M.. Generalized Neural-Network Representation of High-Dimensional Potential-Energy Surfaces. Phys. Rev. Lett. 2007;98:146401. doi: 10.1103/PhysRevLett.98.146401. [ DOI ] [ PubMed ] [ Google Scholar ] Botu V., Ramprasad R.. Adaptive Machine Learning Framework to Accelerate Ab Initio Molecular Dynamics. Int. J. Quantum Chem. 2015;115:1074–1083. doi: 10.1002/qua.24836. [ DOI ] [ Google Scholar ] Galvelis R., Varela-Rial A., Doerr S., Fino R., Eastman P., Markland T. E., Chodera J. D., De Fabritiis G.. NNP/MM: Accelerating Molecular Dynamics Simulations with Machine Learning Potentials and Molecular Mechanics. J. Chem. Inf. Model. 2023;63:5701–5708. doi: 10.1021/acs.jcim.3c00773. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Westermayr J., Marquetand P.. Machine Learning for Electronically Excited States of Molecules. Chem. Rev. 2021;121:9873–9926. doi: 10.1021/acs.chemrev.0c00749. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Li S., Xie B.-B., Yin B.-W., Liu L., Shen L., Fang W.-H.. Construction of Highly Accurate Machine Learning Potential Energy Surfaces for Excited-State Dynamics Simulations Based on Low-Level Data Sets. J. Phys. Chem. A. 2024;128:5516–5524. doi: 10.1021/acs.jpca.4c02028. [ DOI ] [ PubMed ] [ Google Scholar ] Jacobson L. C., Kirby R. M., Molinero V.. How Short Is Too Short for the Interactions of a Water Potential? Exploring the Parameter Space of a Coarse-Grained Water Model Using Uncertainty Quantification. J. Phys. Chem. B. 2014;118:8190–8202. doi: 10.1021/jp5012928. [ DOI ] [ PubMed ] [ Google Scholar ] Evaluation of Measurement DataGuide to the Expression of Uncertainty in Measurement; Joint Committee for Guides in Metrology, 2008. [ Google Scholar ] Taylor, B. N. ; Kuyatt, C. E. . Guidelines for Evaluating and Expressing the Uncertainty of NIST Measurement Results; NIST Technical Note 1297; National Institute of Standards and Technology, 1994. [ Google Scholar ] Possolo A., Hibbert D. B., Stohner J., Bodnar O., Meija J.. A Brief Guide to Measurement Uncertainty (IUPAC Technical Report) Pure Appl. Chem. 2024;96:113–134. doi: 10.1515/pac-2022-1203. [ DOI ] [ Google Scholar ] Hartland G. V.. Statistical Analysis of Physical Chemistry Data: Errors Are Not Mistakes. J. Phys. Chem. A. 2020;124:2109–2112. doi: 10.1021/acs.jpca.0c01403. [ DOI ] [ PubMed ] [ Google Scholar ] Irikura K. K, Johnson R. D III, Kacker R. N. Uncertainty Associated with Virtual Measurements from Computational Quantum Chemistry Models. Metrologia. 2004;41:369. doi: 10.1088/0026-1394/41/6/003. [ DOI ] [ Google Scholar ] Smith, R. C.
Uncertainty Quantification: Theory, Implementation, and Applications; Computational Science and Engineering Series; Society for Industrial and Applied Mathematics: Philadelphia, 2013. [ Google Scholar ] Lord, G. J. ; Powell, C. E. ; Shardlow, T. . An Introduction to Computational Stochastic PDEs; Cambridge Texts in Applied Mathematics 50; Cambridge University Press: New York, 2014. [ Google Scholar ] Sullivan, T. J.
Introduction to Uncertainty Quantification; Springer Science+Business Media: New York, 2015. [ Google Scholar ] Handbook of Uncertainty Quantification; Ghanem, R. , Higdon, D. , Owhadi, H. , Eds.; Springer International Publishing: Cham, Switzerland, 2017. [ Google Scholar ] Caflisch R. E.. Monte Carlo and Quasi-Monte Carlo Methods. Acta Numer. 1998;7:1–49. doi: 10.1017/S0962492900002804. [ DOI ] [ Google Scholar ] Bungartz H.-J., Griebel M.. Sparse Grids. Acta Numer. 2004;13:147–269. doi: 10.1017/S0962492904000182. [ DOI ] [ Google Scholar ] Stuart A. M.. Inverse Problems: A Bayesian Perspective. Acta Numer. 2010;19:451–559. doi: 10.1017/S0962492910000061. [ DOI ] [ Google Scholar ] Dick J., Kuo F. Y., Sloan I. H.. High-Dimensional Integration: The Quasi-Monte Carlo Way. Acta Numer. 2013;22:133–288. doi: 10.1017/S0962492913000044. [ DOI ] [ Google Scholar ] Giles M. B.. Multilevel Monte Carlo Methods. Acta Numer. 2015;24:259–328. doi: 10.1017/S096249291500001X. [ DOI ] [ Google Scholar ] Peherstorfer B., Willcox K., Gunzburger M.. Survey of Multifidelity Methods in Uncertainty Propagation, Inference, and Optimization. SIAM Rev. 2018;60:550–591. doi: 10.1137/16M1082469. [ DOI ] [ Google Scholar ] Glotzer, S. C. ; Kim, S. ; Cummings, P. T. ; Deshmukh, A. ; Head-Gordon, M. ; Karniadakis, G. ; Petzold, L. ; Sagui, C. ; Shinozuka, M. . WTEC Panel Report on International Assessment of Research and Development in Simulation-Based Engineering and Science; World Technology Evaluation Center: Baltimore, MD, 2013. 10.2172/1088842. [ DOI ] [ Google Scholar ] Soize, C.
Uncertainty Quantification: An Accelerated Course with Advanced Applications in Computational Engineering; Interdisciplinary Applied Mathematics, Vol. 47; Springer International Publishing: Cham, Switzerland, 2017. [ Google Scholar ] Riedmaier S., Danquah B., Schick B., Diermeyer F.. Unified Framework and Survey for Model Verification, Validation and Uncertainty Quantification. Arch. Comput. Methods. Eng. 2021;28:2655–2688. doi: 10.1007/s11831-020-09473-7. [ DOI ] [ Google Scholar ] Xin C., Gang W., Yin Y. Z., Xiaojun W. U.. A review of uncertainty quantification methods for Computational Fluid Dynamics. Acta Aerodyn. Sin. 2021;39:1–13. doi: 10.7638/kqdlxxb-2021.0012. [ DOI ] [ Google Scholar ] Simm G. N., Proppe J., Reiher M.. Error Assessment of Computational Models in Chemistry. Chimia. 2017;71:202–208. doi: 10.2533/chimia.2017.202. [ DOI ] [ PubMed ] [ Google Scholar ] Patrone P. N., Dienstfrey A.. Uncertainty Quantification for Molecular Dynamics. Rev. Comput. Chem. 2018;31:115–169. doi: 10.1002/9781119518068.ch3. [ DOI ] [ Google Scholar ] Grossfield A., Patrone P. N., Roe D. R., Schultz A. J., Siderius D., Zuckerman D. M.. Best Practices for Quantification of Uncertainty and Sampling Quality in Molecular Simulations [Article v1.0] Living J. Comput. Mol. Sci. 2019;1:5067. doi: 10.33011/livecoms.1.1.5067. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Lejaeghere, K.
The uncertainty pyramid for electronic-structure methods. In Uncertainty Quantification in Multiscale Materials Modeling; Wang, Y. , McDowell, D. L. , Eds.; Elsevier Series in Mechanics of Advanced Materials; Woodhead Publishing, 2020; pp 41–76. [ Google Scholar ] Pernot P.. The Long Road to Calibrated Prediction Uncertainty in Computational Chemistry. J. Chem. Phys. 2022;156:114109. doi: 10.1063/5.0084302. [ DOI ] [ PubMed ] [ Google Scholar ] Reiher M.. Molecule-Specific Uncertainty Quantification in Quantum Chemical Studies. Isr. J. Chem. 2022;62:e202100101. doi: 10.1002/ijch.202100101. [ DOI ] [ Google Scholar ] Weymuth, T. ; Reiher, M. . Heuristics and Uncertainty Quantification in Rational and Inverse Compound and Catalyst Design. In Comprehensive Computational Chemistry, 1st ed.; Elsevier, 2024; Vol. 4, pp 485–495. [ Google Scholar ] Dai J., Adhikari S., Wen M.. Uncertainty Quantification and Propagation in Atomistic Machine Learning. Rev. Chem. Eng. 2025;41:333–357. doi: 10.1515/revce-2024-0028. [ DOI ] [ Google Scholar ] Arróyave R., McDowell D. L.. Systems Approaches to Materials Design: Past, Present, and Future. Annu. Rev. Mater. Res. 2019;49:103–126. doi: 10.1146/annurev-matsci-070218-125955. [ DOI ] [ Google Scholar ] O’Malley P. J. J.. et al. Scalable Quantum Simulation of Molecular Energies. Phys. Rev. X. 2016;6:031007. doi: 10.1103/PhysRevX.6.031007. [ DOI ] [ Google Scholar ] Cao Y., Romero J., Olson J. P., Degroote M., Johnson P. D., Kieferová M., Kivlichan I. D., Menke T., Peropadre B., Sawaya N. P. D., Sim S., Veis L., Aspuru-Guzik A.. Quantum Chemistry in the Age of Quantum Computing. Chem. Rev. 2019;119:10856–10915. doi: 10.1021/acs.chemrev.8b00803. [ DOI ] [ PubMed ] [ Google Scholar ] Becerra A., Prabhu A., Rongali M. S., Velpur S. C. S., Debusschere B., Walker E. A.. How a Quantum Computer Could Quantify Uncertainty in Microkinetic Models. J. Phys. Chem. Lett. 2021;12:6955–6960. doi: 10.1021/acs.jpclett.1c01917. [ DOI ] [ PubMed ] [ Google Scholar ] Motta M., Rice J. E.. Emerging Quantum Computing Algorithms for Quantum Chemistry. Wiley Interdiscip. Rev.: Comput. Mol. Sci. 2022;12:e1580. doi: 10.1002/wcms.1580. [ DOI ] [ Google Scholar ] Bickley T. M., Mingare A., Weaving T., Williams de la Bastida M., Wan S., Nibbi M., Seitz P., Ralli A., Love P. J., Chung M., Hernandez Vera M., Schulz L., Coveney P. V.. Extending Quantum Computing through Subspace, Embedding and Classical Molecular Dynamics Techniques. Digital Discovery. 2025;4:3427–3444. doi: 10.1039/D5DD00225G. [ DOI ] [ Google Scholar ] Van Noorden R.. More than 10,000 Research Papers Were Retracted in 2023A New Record. Nature. 2023;624:479–481. doi: 10.1038/d41586-023-03974-8. [ DOI ] [ PubMed ] [ Google Scholar ] Langmuir I.. Pathological Science. Res.-Technol. Manage. 1989;32:11–17. doi: 10.1080/08956308.1989.11670607. [ DOI ] [ Google Scholar ] Lombardo T.. et al. Artificial Intelligence Applied to Battery Research: Hype or Reality? Chem. Rev. 2022;122:10899–10969. doi: 10.1021/acs.chemrev.1c00108. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Ceriotti M., Clementi C., Anatole von Lilienfeld O.. Introduction: Machine Learning at the Atomic Scale. Chem. Rev. 2021;121:9719–9721. doi: 10.1021/acs.chemrev.1c00598. [ DOI ] [ PubMed ] [ Google Scholar ] Coveney P. V., Highfield R. R.. When We Can Trust Computers (and When We Can’t) Philos. Trans. R. Soc. A. 2021;379:20200067. doi: 10.1098/rsta.2020.0067. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Volodina V., Challenor P.. The Importance of Uncertainty Quantification in Model Reproducibility. Philos. Trans. R. Soc. A. 2021;379:20200071. doi: 10.1098/rsta.2020.0071. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Mitchell J. B. O.. Machine Learning Methods in Chemoinformatics. Wiley Interdiscip. Rev.: Comput. Mol. Sci. 2014;4:468–481. doi: 10.1002/wcms.1183. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Peterson A. A., Christensen R., Khorshidi A.. Addressing Uncertainty in Atomistic Machine Learning. Phys. Chem. Chem. Phys. 2017;19:10978–10985. doi: 10.1039/C7CP00375G. [ DOI ] [ PubMed ] [ Google Scholar ] Mater A. C., Coote M. L.. Deep Learning in Chemistry. J. Chem. Inf. Model. 2019;59:2545–2559. doi: 10.1021/acs.jcim.9b00266. [ DOI ] [ PubMed ] [ Google Scholar ] Keith J. A., Vassilev-Galindo V., Cheng B., Chmiela S., Gastegger M., Müller K.-R., Tkatchenko A.. Combining Machine Learning and Computational Chemistry for Predictive Insights Into Chemical Systems. Chem. Rev. 2021;121:9816–9872. doi: 10.1021/acs.chemrev.1c00107. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Unke O. T., Chmiela S., Sauceda H. E., Gastegger M., Poltavsky I., Schütt K. T., Tkatchenko A., Müller K.-R.. Machine Learning Force Fields. Chem. Rev. 2021;121:10142–10186. doi: 10.1021/acs.chemrev.0c01111. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Meuwly M.. Machine Learning for Chemical Reactions. Chem. Rev. 2021;121:10218–10239. doi: 10.1021/acs.chemrev.1c00033. [ DOI ] [ PubMed ] [ Google Scholar ] Koutsoukos S., Philippi F., Malaret F., Welton T.. A Review on Machine Learning Algorithms for the Ionic Liquid Chemical Space. Chem. Sci. 2021;12:6820–6843. doi: 10.1039/D1SC01000J. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Huang B., von Lilienfeld O. A.. Ab Initio Machine Learning in Chemical Compound Space. Chem. Rev. 2021;121:10001–10036. doi: 10.1021/acs.chemrev.0c01303. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Bender A., Schneider N., Segler M., Patrick Walters W., Engkvist O., Rodrigues T.. Evaluation Guidelines for Machine Learning Tools in the Chemical Sciences. Nat. Rev. Chem. 2022;6:428–442. doi: 10.1038/s41570-022-00391-9. [ DOI ] [ PubMed ] [ Google Scholar ] Weymuth T., Reiher M.. The Transferability Limits of Static Benchmarks. Phys. Chem. Chem. Phys. 2022;24:14692–14698. doi: 10.1039/D2CP01725C. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Gecht M., Siggel M., Linke M., Hummer G., Köfinger J.. MDBenchmark: A Toolkit to Optimize the Performance of Molecular Dynamics Simulations. J. Chem. Phys. 2020;153:144105. doi: 10.1063/5.0019045. [ DOI ] [ PubMed ] [ Google Scholar ] Chen, G. ; Chen, P. ; Hsieh, C.-Y. ; Lee, C.-K. ; Liao, B. ; Liao, R. ; Liu, W. ; Qiu, J. ; Sun, Q. ; Tang, J. ; Zemel, R. ; Zhang, S. . Alchemy: A Quantum Chemistry Dataset for Benchmarking AI Models. arXiv (Computer Science.Machine Learning), June 22, 2019, 1906.09427, ver. 1. https://arxiv.org/abs/1906.09427 . Helgaker, T. ; Jorgensen, P. ; Olsen, J. . Molecular Electronic-Structure Theory; John Wiley & Sons, 2014. [ Google Scholar ] Allen, M. P. ; Tildesley, D. J. . Computer Simulation of Liquids, 2nd ed.; Oxford University Press: New York, 2017. [ Google Scholar ] Frenkel, D. ; Smit, B. . Understanding Molecular Simulation: From Algorithms to Applications, 3rd ed.; Academic Press, 2023. [ Google Scholar ] Griebel, M. ; Zumbusch, G. ; Knapeck, S. . Numerical Simulation in Molecular Dynamics; Texts in Computational Science and Engineering, Vol. 5; Springer: Berlin, 2007. [ Google Scholar ] Herbst M. F., Levitt A., Cancès E.. A Posteriori Error Estimation for the Non-Self-Consistent Kohn-Sham Equations. Faraday Discuss. 2020;224:227–246. doi: 10.1039/D0FD00048E. [ DOI ] [ PubMed ] [ Google Scholar ] Leimkuhler, B. ; Matthews, C. . Molecular Dynamics: With Deterministic and Stochastic Numerical Methods; Interdisciplinary Applied Mathematics, Vol. 39; Springer International Publishing: Cham, Switzerland, 2015. [ Google Scholar ] Coveney, P. V. ; Wan, S. . Molecular Dynamics: Probability and Uncertainty; Oxford University Press: New York, 2025. [ Google Scholar ] Cox, D. R.
Principles of Statistical Inference; Cambridge University Press: Cambridge, U.K., 2006. [ Google Scholar ] McCullagh P.. What Is a Statistical Model? Ann. Stat. 2002;30:1225–1310. doi: 10.1214/aos/1035844977. [ DOI ] [ Google Scholar ] Karr, A. F.
Probability; Springer Texts in Statistics; Springer: New York, 1993. [ Google Scholar ] Bayes T.. LII. An Essay towards Solving a Problem in the Doctrine of Chances. By the Late Rev. Mr. Bayes, F. R. S. Communicated by Mr. Price, in a Letter to John Canton, A. M. F. R. S. Philos. Trans. R. Soc. 1763;53:370–418. doi: 10.1098/rstl.1763.0053. [ DOI ] [ Google Scholar ] Rabitz H., Aliş Ö. F.. General Foundations of High-Dimensional Model Representations. J. Math. Chem. 1999;25:197–233. doi: 10.1023/A:1019188517934. [ DOI ] [ Google Scholar ] Kaipio, J. ; Somersalo, E. . Statistical and Computational Inverse Problems; Springer E-Books, Vol. 160; Springer: New York, 2005. [ Google Scholar ] Sobol’ I.. Global Sensitivity Indices for Nonlinear Mathematical Models and Their Monte Carlo Estimates. Math. Comput. Simul. 2001;55:271–280. doi: 10.1016/S0378-4754(00)00270-6. [ DOI ] [ Google Scholar ] Borgonovo, E.
Sensitivity Analysis; International Series in Operations Research & Management Science, Vol. 251; Springer International Publishing: Cham, Switzerland, 2017. [ Google Scholar ] Ghanem, R. G. ; Spanos, P. D. . Stochastic Finite Elements: A Spectral Approach; Dover Publications: Mineola, NY, 2003. [ Google Scholar ] Perturbation Methods; Nayfeh, A. H. , Ed.; Wiley Classics Library; John Wiley & Sons: New York, 2011. [ Google Scholar ] Collins J. D., Thomson W. T.. The Eigenvalue Problem for Structural Systems with Statistical Properties. AIAA J. 1969;7:642–648. doi: 10.2514/3.5180. [ DOI ] [ Google Scholar ] Dölz J., Harbrecht H., Schwab Ch.. Covariance Regularity and H-Matrix Approximation for Rough Random Fields. Numer. Math. 2017;135:1045–1071. doi: 10.1007/s00211-016-0825-y. [ DOI ] [ Google Scholar ] Dölz J., Ebert D.. On Uncertainty Quantification of Eigenvalues and Eigenspaces with Higher Multiplicity. SIAM J. Numer. Anal. 2024;62:422–451. doi: 10.1137/22M1529324. [ DOI ] [ Google Scholar ] Harbrecht H., Schneider R., Schwab Ch.. Multilevel Frames for Sparse Tensor Product Spaces. Numer. Math. 2008;110:199–220. doi: 10.1007/s00211-008-0162-x. [ DOI ] [ Google Scholar ] Dölz, J. ; Ebert, D. . Local Sensitivity Analysis for Bayesian Inverse Problems. arXiv (Mathematics.Numerical Analysis), March 26, 2025, 2503.20526, ver. 1. https://arxiv.org/abs/2503.20526 . Lemieux, C.
Monte Carlo and Quasi-Monte Carlo Sampling; Springer Series in Statistics; Springer: New York, 2009. [ Google Scholar ] Hastings W. K.. Monte Carlo Sampling Methods Using Markov Chains and Their Applications. Biometrika. 1970;57:97–109. doi: 10.1093/biomet/57.1.97. [ DOI ] [ Google Scholar ] Handbook of Markov Chain Monte Carlo, 1st ed.; Brooks, S. , Gelman, A. , Jones, G. , Meng, X.-L. , Eds.; Chapman and Hall/CRC: New York, 2011. [ Google Scholar ] Haario H., Saksman E., Tamminen J.. An Adaptive Metropolis Algorithm. Bernoulli. 2001;7:223–242. doi: 10.2307/3318737. [ DOI ] [ Google Scholar ] Rosenthal, J. S.
Optimal Proposal Distributions and Adaptive MCMC. In Handbook of Markov Chain Monte Carlo, 1st ed.; Chapman and Hall/CRC: New York, 2011; pp 93–112. [ Google Scholar ] Girolami M., Calderhead B.. Riemann Manifold Langevin and Hamiltonian Monte Carlo Methods. J. R. Stat. Soc. Ser. B Stat. Method. 2011;73:123–214. doi: 10.1111/j.1467-9868.2010.00765.x. [ DOI ] [ Google Scholar ] Hoffman M. D., Gelman A.. The No-U-turn Sampler: Adaptively Setting Path Lengths in Hamiltonian Monte Carlo. J. Mach. Learn. Res. 2014;15:1593–1623. [ Google Scholar ] Duane S., Kennedy A., Pendleton B. J., Roweth D.. Hybrid Monte Carlo. Phys. Lett. B. 1987;195:216–222. doi: 10.1016/0370-2693(87)91197-X. [ DOI ] [ Google Scholar ] Fisher, R. A.
The Design of Experiments; Oliver and Boyd: Edinburgh, 1935. [ Google Scholar ] Edgington, E. S.
Randomization Tests; Marcel Dekker: New York, 1987. [ Google Scholar ] Good, P.
Permutation Tests; Springer: New York, 1994. [ Google Scholar ] Quenouille M. H.. Approximate Tests of Correlation in Time-Series. J. R. Stat. Soc. Ser. B Stat. Method. 1949;11:68–84. doi: 10.1111/j.2517-6161.1949.tb00023.x. [ DOI ] [ Google Scholar ] Tukey J. W.. Bias and Confidence in Not Quite Large Samples. Ann. Math. Statist. 1958;29:614. doi: 10.1214/aoms/1177706647. [ DOI ] [ Google Scholar ] Davison, A. C. ; Hinkley, D. V. . Bootstrap Methods and Their Application; Cambridge University Press: Cambridge, U.K., 1997. [ Google Scholar ] Efron B.. Bootstrap Methods: Another Look at the Jackknife. Ann. Stat. 1979;7:1–26. doi: 10.1214/aos/1176344552. [ DOI ] [ Google Scholar ] Efron, B. ; Tibshirani, R. . An Introduction to the Bootstrap; Monographs on Statistics and Applied Probability 57; Chapman & Hall: New York, 1993. [ Google Scholar ] Rodgers J. L.. The Bootstrap, the Jackknife, and the Randomization Test: A Sampling Taxonomy. Multivar. Behav. Res. 1999;34:441–456. doi: 10.1207/S15327906MBR3404_2. [ DOI ] [ PubMed ] [ Google Scholar ] Niederreiter, H.
Random Number Generation and Quasi-Monte Carlo Methods; CBMS-NSF Regional Conference Series in Applied Mathematics 63; Society for Industrial and Applied Mathematics: Philadelphia, PA, 1992. [ Google Scholar ] Gantner, R. N. ; Schwab, C. . Computational Higher Order Quasi-Monte Carlo Integration. In Monte Carlo and Quasi-Monte Carlo Methods (MCQMC), Leuven, Belgium, April 2014; Cools, R. , Nuyens, D. , Eds.; Springer Proceedings in Mathematics & Statistics, Vol. 163; Springer International Publishing: Cham, Switzerland, 2016; pp 271–288. [ Google Scholar ] Nuyens D., Cools R.. Higher Order Quasi-Monte Carlo Methods: A Comparison. AIP Conf. Proc. 2010;1281:553–557. doi: 10.1063/1.3498535. [ DOI ] [ Google Scholar ] Ullrich, M.
On “Upper Error Bounds for Quadrature Formulas on Function Classes” by K.K. Frolov. In Monte Carlo and Quasi-Monte Carlo Methods (MCQMC), Leuven, Belgium, April 2014; Cools, R. , Nuyens, D. , Eds.; Springer Proceedings in Mathematics & Statistics, Vol. 163; Springer International Publishing: Cham, Switzerland, 2016; pp 571–582. [ Google Scholar ] Smolyak S. A.. Quadrature and interpolation formulas for tensor products of certain classes of functions. Dokl. Akad. Nauk SSSR. 1963;148:1042–1045. [ Google Scholar ] Gerstner T., Griebel M.. Numerical Integration Using Sparse Grids. Numer. Algorithms. 1998;18:209–232. doi: 10.1023/A:1019129717644. [ DOI ] [ Google Scholar ] Griebel, M. ; Schneider, M. ; Zenger, C. . A combination technique for the solution of sparse grid problems. In Iterative Methods in Linear Algebra; de Groen, P. , Beauwens, R. , Eds.; IMACS, Elsevier: North Holland, 1992; pp 263–281. [ Google Scholar ] Wasilkowski G., Wozniakowski H.. Explicit Cost Bounds of Algorithms for Multivariate Tensor Product Problems. J. Complex. 1995;11:1–56. doi: 10.1006/jcom.1995.1001. [ DOI ] [ Google Scholar ] Novak E., Ritter K.. High Dimensional Integration of Smooth Functions over Cubes. Numer. Math. 1996;75:79–97. doi: 10.1007/s002110050231. [ DOI ] [ Google Scholar ] Haji-Ali A.-L., Harbrecht H., Peters M., Siebenmorgen M.. Novel Results for the Anisotropic Sparse Grid Quadrature. J. Complex. 2018;47:62–85. doi: 10.1016/j.jco.2018.02.003. [ DOI ] [ Google Scholar ] Zech J., Schwab C.. Convergence Rates of High Dimensional Smolyak Quadrature. ESAIM: M2AN. 2020;54:1259–1307. doi: 10.1051/m2an/2020003. [ DOI ] [ Google Scholar ] Gerstner T., Griebel M.. Dimension Adaptive Tensor Product Quadrature. Computing. 2003;71:65–87. doi: 10.1007/s00607-003-0015-5. [ DOI ] [ Google Scholar ] Garcke, J.
Sparse Grids in a Nutshell. In Sparse Grids and Applications; Garcke, J. , Griebel, M. , Eds.; Lecture Notes in Computational Science and Engineering, Vol. 88; Springer: Berlin, 2012; pp 57–80. [ Google Scholar ] Xiu, D.
Stochastic Collocation Methods: A Survey. In Handbook of Uncertainty Quantification; Ghanem, R. , Higdon, D. , Owhadi, H. , Eds.; Springer International Publishing: Cham, Switzerland, 2016; pp 1–18. [ Google Scholar ] Chkifa A., Cohen A., Schwab C.. High-Dimensional Adaptive Sparse Polynomial Interpolation and Applications to Parametric PDEs. Found. Comput. Math. 2014;14:601–633. doi: 10.1007/s10208-013-9154-z. [ DOI ] [ Google Scholar ] Nobile F., Tamellini L., Tempone R.. Convergence of Quasi-Optimal Sparse-Grid Approximation of Hilbert-space-valued Functions: Application to Random Elliptic PDEs. Numer. Math. 2016;134:343–388. doi: 10.1007/s00211-015-0773-y. [ DOI ] [ Google Scholar ] Wiener N.. The Homogeneous Chaos. Am. J. Math. 1938;60:897. doi: 10.2307/2371268. [ DOI ] [ Google Scholar ] Xiu D., Karniadakis G. E.. The Wiener-Askey Polynomial Chaos for Stochastic Differential Equations. SIAM J. Sci. Comput. 2002;24:619–644. doi: 10.1137/S1064827501387826. [ DOI ] [ Google Scholar ] Debusschere B., Najm H., Pébay P., Knio O., Ghanem R., Le Maître O.. Numerical Challenges in the Use of Polynomial Chaos Representations for Stochastic Processes. SIAM J. Sci. Comput. 2004;26:698–719. doi: 10.1137/S1064827503427741. [ DOI ] [ Google Scholar ] Adcock, B. ; Brugiapaglia, S. ; Webster, C. G. . Sparse Polynomial Approximation of High-Dimensional Functions; Society for Industrial and Applied Mathematics: Philadelphia, PA, 2022. [ Google Scholar ] Jakeman J. D., Franzelin F., Narayan A., Eldred M., Plfüger D.. Polynomial Chaos Expansions for Dependent Random Variables. Comput. Methods Appl. Mech. Eng. 2019;351:643–666. doi: 10.1016/j.cma.2019.03.049. [ DOI ] [ Google Scholar ] Lüthen N., Marelli S., Sudret B.. Sparse Polynomial Chaos Expansions: Literature Survey and Benchmark. SIAM/ASA J. Uncertainty Quantification. 2021;9:593–649. doi: 10.1137/20M1315774. [ DOI ] [ Google Scholar ] Blatman G., Sudret B.. Adaptive Sparse Polynomial Chaos Expansion Based on Least Angle Regression. J. Comput. Phys. 2011;230:2345–2367. doi: 10.1016/j.jcp.2010.12.021. [ DOI ] [ Google Scholar ] Sudret B.. Global Sensitivity Analysis Using Polynomial Chaos Expansions. Reliab. Eng. Syst. Saf. 2008;93:964–979. doi: 10.1016/j.ress.2007.04.002. [ DOI ] [ Google Scholar ] Heinrich, S.
Multilevel Monte Carlo Methods. In Large-Scale Scientific Computing; Margenov, S. , Waśniewski, J. , Yalamov, P. , Eds.; Lecture Notes in Computer Science, Vol. 2179; Springer: Berlin, 2001; pp 58–67. [ Google Scholar ] Harbrecht, H. ; Peters, M. ; Siebenmorgen, M. . On Multilevel Quadrature for Elliptic Stochastic Partial Differential Equations. In Sparse Grids and Applications; Garcke, J. , Griebel, M. , Eds.; Lecture Notes in Computational Science and Engineering, Vol. 88; pp 161–179. [ Google Scholar ] Haji-Ali A.-L., Nobile F., Tempone R.. Multi-Index Monte Carlo: When Sparsity Meets Sampling. Numer. Math. 2016;132:767–806. doi: 10.1007/s00211-015-0734-5. [ DOI ] [ Google Scholar ] Abdar M., Pourpanah F., Hussain S., Rezazadegan D., Liu L., Ghavamzadeh M., Fieguth P., Cao X., Khosravi A., Acharya U. R., Makarenkov V., Nahavandi S.. A Review of Uncertainty Quantification in Deep Learning: Techniques, Applications and Challenges. Inf. Fusion. 2021;76:243–297. doi: 10.1016/j.inffus.2021.05.008. [ DOI ] [ Google Scholar ] Hüllermeier E., Waegeman W.. Aleatoric and Epistemic Uncertainty in Machine Learning: An Introduction to Concepts and Methods. Mach. Learn. 2021;110:457–506. doi: 10.1007/s10994-021-05946-3. [ DOI ] [ Google Scholar ] Rasmussen, C. E. ; Williams, C. K. I. . Gaussian Processes for Machine Learning; Adaptive Computation and Machine Learning Series; MIT Press: Cambridge, MA, 2006. [ Google Scholar ] Cressie, N. A. C.
Statistics for Spatial Data, 1st ed.; Wiley Series in Probability and Statistics; Wiley, 1993. [ Google Scholar ] Stein, M. L.
Interpolation of Spatial Data: Some Theory for Kriging; Springer: New York, 2013. [ Google Scholar ] Saunders, C. ; Gammerman, A. ; Vovk, V. . Ridge Regression Learning Algorithm in Dual Variables. In ICML ’98: Proceedings of the Fifteenth International Conference on Machine Learning, San Francisco, CA, July 24–27, 1998; Morgan Kaufmann Publishers, 1998; pp 515–521. [ Google Scholar ] Schölkopf, B. ; Smola, A. J. . Learning with Kernels: Support Vector Machines, Regularization, Optimization, and Beyond; Adaptive Computation and Machine Learning Series; MIT Press: Cambridge, MA, 2002. [ Google Scholar ] Deringer V. L., Bartók A. P., Bernstein N., Wilkins D. M., Ceriotti M., Csányi G.. Gaussian Process Regression for Materials and Molecules. Chem. Rev. 2021;121:10073–10141. doi: 10.1021/acs.chemrev.1c00022. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Pernot P.. Prediction Uncertainty Validation for Computational Chemists. J. Chem. Phys. 2022;157:144103. doi: 10.1063/5.0109572. [ DOI ] [ PubMed ] [ Google Scholar ] Ruscic B.. Uncertainty Quantification in Thermochemistry, Benchmarking Electronic Structure Computations, and Active Thermochemical Tables. Int. J. Quantum Chem. 2014;114:1097–1101. doi: 10.1002/qua.24605. [ DOI ] [ Google Scholar ] Wright D. W.. et al. Building Confidence in Simulation: Applications of EasyVVUQ. Adv. Theory Simul. 2020;3:1900246. doi: 10.1002/adts.201900246. [ DOI ] [ Google Scholar ] Bohn, B. ; Garcke, J. ; Griebel, M. . Algorithmic Mathematics in Machine Learning; Society for Industrial and Applied Mathematics: Philadelphia, 2024. [ Google Scholar ] Edeling W.. On the Deep Active-Subspace Method. SIAM/ASA J. Uncertainty Quantification. 2023;11:62–90. doi: 10.1137/21M1463240. [ DOI ] [ Google Scholar ] van der Maaten, L. ; Postma, E. ; Herik, H. . Dimensionality Reduction: A Comparative Review; Tilburg Centre for Creative Computing: Tilburg, The Netherlands, 2009. [ Google Scholar ] Tripathy R. K., Bilionis I.. Deep UQ: Learning Deep Neural Network Surrogate Models for High Dimensional Uncertainty Quantification. J. Comput. Phys. 2018;375:565–588. doi: 10.1016/j.jcp.2018.08.036. [ DOI ] [ Google Scholar ] Feinberg, J.
Chaospy, 2023. Adams, B. M.
et al. Dakota, A Multilevel Parallel Object-Oriented Framework for Design Optimization, Parameter Estimation, Uncertainty Quantification, and Sensitivity Analysis: Version 6.13 User’s Manual, 2020. Richardson R. A., Wright D. W., Edeling W., Jancauskas V., Lakhlili J., Coveney P. V.. EasyVVUQ: A Library for Verification, Validation and Uncertainty Quantification in High Performance Computing. J. Open Res. Software. 2020;8:11. doi: 10.5334/jors.303. [ DOI ] [ Google Scholar ] Parno M., Davis A., Seelinger L.. MUQ: The MIT Uncertainty Quantification Library. J. Open Source Software. 2021;6:3076. doi: 10.21105/joss.03076. [ DOI ] [ Google Scholar ] Lafarge, T. ; Newton, D. ; Koepke, A. ; Possolo, A. . NIST Uncertainty Machine; National Institute of Standards and Technology, 2023. [ Google Scholar ] Baudin, M. ; Dutfoy, A. ; Iooss, B. ; Popelin, A.-L. . OpenTURNS: An Industrial Software for Uncertainty Quantification in Simulation. In Handbook of Uncertainty Quantification; Ghanem, R. , Higdon, D. , Owhadi, H. , Eds.; Springer International Publishing: Cham, Switzerland, 2015; pp 1–38. [ Google Scholar ] Merz P. T., Shirts M. R.. Testing for Physical Validity in Molecular Simulations. PLoS One. 2018;13:e0202764. doi: 10.1371/journal.pone.0202764. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Abraham M. J., Murtola T., Schulz R., Páll S., Smith J. C., Hess B., Lindahl E.. GROMACS: High Performance Molecular Simulations through Multi-Level Parallelism from Laptops to Supercomputers. SoftwareX. 2015;1–2:19–25. doi: 10.1016/j.softx.2015.06.001. [ DOI ] [ Google Scholar ] Jakeman J. D.. PyApprox: A Software Package for Sensitivity Analysis, Bayesian Inference, Optimal Experimental Design, and Multi-Fidelity Uncertainty Quantification and Surrogate Modeling. Environ. Model. Software. 2023;170:105825. doi: 10.1016/j.envsoft.2023.105825. [ DOI ] [ Google Scholar ] Herman J., Usher W.. SALib: An Open-Source Python Library for Sensitivity Analysis. J. Open Source Software. 2017;2:97. doi: 10.21105/joss.00097. [ DOI ] [ Google Scholar ] Groen D.. et al. VECMAtk: A Scalable Verification, Validation and Uncertainty Quantification Toolkit for Scientific Simulations. Philos. Trans. R. Soc. A. 2021;379:20200221. doi: 10.1098/rsta.2020.0221. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Leser, P. ; Wang, M. . SMCPy - Sequential Monte Carlo with Python; NASA, 2023. [ Google Scholar ] Piazzola C., Tamellini L.. Algorithm1040 The Sparse Grids Matlab Kit - a Matlab Implementation of Sparse Grids for High-Dimensional Function Approximation and Uncertainty Quantification. ACM Trans. Math. Software. 2024;50:7. doi: 10.1145/3630023. [ DOI ] [ Google Scholar ] Narayan A., Liu Z., Bergquist J. A., Charlebois C., Rampersad S., Rupp L., Brooks D., White D., Tate J., MacLeod R. S.. UncertainSCI: Uncertainty Quantification for Computational Models in Biomedicine and Bioengineering. Comput. Biol. Med. 2023;152:106407. doi: 10.1016/j.compbiomed.2022.106407. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Seelinger L., Cheng-Seelinger V., Davis A., Parno M., Reinarz A.. UM-Bridge: Uncertainty Quantification and Modeling Bridge. J. Open Source Software. 2023;8:4748. doi: 10.21105/joss.04748. [ DOI ] [ Google Scholar ] Marelli, S. ; Sudret, B. . UQLab: A Framework for Uncertainty Quantification in Matlab. In Vulnerability, Uncertainty, and Risk: Quantification, Mitigation, and Management; ASCE, 2014; pp 2554–2563. 10.1061/9780784413609.257 [ DOI ] [ Google Scholar ] Debusschere, B. ; Sargsyan, K. ; Safta, C. ; Chowdhary, K. . Uncertainty Quantification Toolkit (UQTk). In Handbook of Uncertainty Quantification; Ghanem, R. , Higdon, D. , Owhadi, H. , Eds.; Springer, 2017; pp 1807–1827. [ Google Scholar ] Blanchard J.-B., Damblin G., Martinez J.-M., Arnaud G., Gaudier F.. The Uranie Platform: An Open-Source Software for Optimisation, Meta-Modelling and Uncertainty Analysis. EPJ Nucl. Sci. Technol. 2019;5:4. doi: 10.1051/epjn/2018050. [ DOI ] [ Google Scholar ] Zwier M. C., Adelman J. L., Kaus J. W., Pratt A. J., Wong K. F., Rego N. B., Suárez E., Lettieri S., Wang D. W., Grabe M., Zuckerman D. M., Chong L. T.. WESTPA: An Interoperable, Highly Scalable Software Package for Weighted Ensemble Simulation and Analysis. J. Chem. Theory Comput. 2015;11:800–809. doi: 10.1021/ct5010615. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Russo J. D.. et al. WESTPA 2.0: High-Performance Upgrades for Weighted Ensemble Simulations and Analysis of Longer-Timescale Applications. J. Chem. Theory Comput. 2022;18:638–649. doi: 10.1021/acs.jctc.1c01154. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Bonomi M.. et al. Promoting Transparency and Reproducibility in Enhanced Molecular Simulations. Nat. Methods. 2019;16:670–673. doi: 10.1038/s41592-019-0506-8. [ DOI ] [ PubMed ] [ Google Scholar ] Roca-Sanjuán D., Aquilante F., Lindh R.. Multiconfiguration Second-Order Perturbation Theory Approach to Strong Electron Correlation in Chemistry and Photochemistry. Wiley Interdiscip. Rev.: Comput. Mol. Sci. 2012;2:585–603. doi: 10.1002/wcms.97. [ DOI ] [ Google Scholar ] Szalay P. G., Müller T., Gidofalvi G., Lischka H., Shepard R.. Multiconfiguration Self-Consistent Field and Multireference Configuration Interaction Methods and Applications. Chem. Rev. 2012;112:108–181. doi: 10.1021/cr200137a. [ DOI ] [ PubMed ] [ Google Scholar ] Lischka H., Nachtigallová D., Aquino A. J. A., Szalay P. G., Plasser F., Machado F. B. C., Barbatti M.. Multireference Approaches for Excited States of Molecules. Chem. Rev. 2018;118:7293–7361. doi: 10.1021/acs.chemrev.8b00244. [ DOI ] [ PubMed ] [ Google Scholar ] Baiardi A., Reiher M.. The Density Matrix Renormalization Group in Chemistry and Molecular Physics: Recent Developments and New Challenges. J. Chem. Phys. 2020;152:040903. doi: 10.1063/1.5129672. [ DOI ] [ PubMed ] [ Google Scholar ] Hättig C., Klopper W., Köhn A., Tew D. P.. Explicitly Correlated Electrons in Molecules. Chem. Rev. 2012;112:4–74. doi: 10.1021/cr200168z. [ DOI ] [ PubMed ] [ Google Scholar ] Kołos W., Wolniewicz L.. Accurate Adiabatic Treatment of the Ground State of the Hydrogen Molecule. J. Chem. Phys. 1964;41:3663–3673. doi: 10.1063/1.1725796. [ DOI ] [ Google Scholar ] Cencek W., Szalewicz K.. Ultra-High Accuracy Calculations for Hydrogen Molecule and Helium Dimer. Int. J. Quantum Chem. 2008;108:2191–2198. doi: 10.1002/qua.21740. [ DOI ] [ Google Scholar ] Pachucki K., Komasa J.. Schrödinger Equation Solved for the Hydrogen Molecule with Unprecedented Accuracy. J. Chem. Phys. 2016;144:164306. doi: 10.1063/1.4948309. [ DOI ] [ PubMed ] [ Google Scholar ] Helgaker T., Klopper W., Tew D. P.. Quantitative Quantum Chemistry. Mol. Phys. 2008;106:2107–2143. doi: 10.1080/00268970802258591. [ DOI ] [ Google Scholar ] Ramabhadran R. O., Raghavachari K.. Extrapolation to the Gold-Standard in Quantum Chemistry: Computationally Efficient and Accurate CCSD(T) Energies for Large Molecules Using an Automated Thermochemical Hierarchy. J. Chem. Theory Comput. 2013;9:3986–3994. doi: 10.1021/ct400465q. [ DOI ] [ PubMed ] [ Google Scholar ] Lang J., Przybytek M., Lesiuk M.. Estimating the Complete Basis Set Extrapolation Error through Random Walks. J. Phys. Chem. Lett. 2025;16:4952–4961. doi: 10.1021/acs.jpclett.5c00749. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Schuurman M. S., Muir S. R., Allen W. D., Schaefer H. F. III.. Toward Subchemical Accuracy in Computational Thermochemistry: Focal Point Analysis of the Heat of Formation of NCO and [H,N,C,O] Isomers. J. Chem. Phys. 2004;120:11586–11599. doi: 10.1063/1.1707013. [ DOI ] [ PubMed ] [ Google Scholar ] Hajgató B., Huzak M., Deleuze M. S.. Focal Point Analysis of the Singlet-Triplet Energy Gap of Octacene and Larger Acenes. J. Phys. Chem. A. 2011;115:9282–9293. doi: 10.1021/jp2043043. [ DOI ] [ PubMed ] [ Google Scholar ] Bakowies D.. Estimating Systematic Error and Uncertainty in Ab Initio Thermochemistry. I. Atomization Energies of Hydrocarbons in the ATOMIC(Hc) Protocol. J. Chem. Theory Comput. 2019;15:5230–5251. doi: 10.1021/acs.jctc.9b00343. [ DOI ] [ PubMed ] [ Google Scholar ] Bakowies D.. Estimating Systematic Error and Uncertainty in Ab Initio Thermochemistry: II. ATOMIC(Hc) Enthalpies of Formation for a Large Set of Hydrocarbons. J. Chem. Theory Comput. 2020;16:399–426. doi: 10.1021/acs.jctc.9b00974. [ DOI ] [ PubMed ] [ Google Scholar ] Lang J., Garberoglio G., Przybytek M., Jeziorska M., Jeziorski B.. Three-Body Potential and Third Virial Coefficients for Helium Including Relativistic and Nuclear-Motion Effects. Phys. Chem. Chem. Phys. 2023;25:23395–23416. doi: 10.1039/D3CP01794J. [ DOI ] [ PubMed ] [ Google Scholar ] Larsson H. R., Zhai H., Umrigar C. J., Chan G. K.-L.. The Chromium Dimer: Closing a Chapter of Quantum Chemistry. J. Am. Chem. Soc. 2022;144:15932–15937. doi: 10.1021/jacs.2c06357. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Stein C. J., Reiher M.. Automated Selection of Active Orbital Spaces. J. Chem. Theory Comput. 2016;12:1760–1771. doi: 10.1021/acs.jctc.6b00156. [ DOI ] [ PubMed ] [ Google Scholar ] Sayfutyarova E. R., Sun Q., Chan G. K.-L., Knizia G.. Automated Construction of Molecular Active Spaces from Atomic Valence Orbitals. J. Chem. Theory Comput. 2017;13:4063–4078. doi: 10.1021/acs.jctc.7b00128. [ DOI ] [ PubMed ] [ Google Scholar ] Stein C. J., Reiher M.. autoCAS: A Program for Fully Automated Multiconfigurational Calculations. J. Comput. Chem. 2019;40:2216–2226. doi: 10.1002/jcc.25869. [ DOI ] [ PubMed ] [ Google Scholar ] Jeong W., Stoneburner S. J., King D., Li R., Walker A., Lindh R., Gagliardi L.. Automation of Active Space Selection for Multireference Methods via Machine Learning on Chemical Bond Dissociation. J. Chem. Theory Comput. 2020;16:2389–2399. doi: 10.1021/acs.jctc.9b01297. [ DOI ] [ PubMed ] [ Google Scholar ] Coe J. P., Paterson M. J.. Investigating Multireference Character and Correlation in Quantum Chemistry. J. Chem. Theory Comput. 2015;11:4189–4196. doi: 10.1021/acs.jctc.5b00543. [ DOI ] [ PubMed ] [ Google Scholar ] Duan C., Liu F., Nandy A., Kulik H. J.. Semi-Supervised Machine Learning Enables the Robust Detection of Multireference Character at Low Cost. J. Phys. Chem. Lett. 2020;11:6640–6648. doi: 10.1021/acs.jpclett.0c02018. [ DOI ] [ PubMed ] [ Google Scholar ] Karton A., Rabinovich E., Martin J. M. L., Ruscic B.. W4 Theory for Computational Thermochemistry: In Pursuit of Confident Sub-kJ/Mol Predictions. J. Chem. Phys. 2006;125:144108. doi: 10.1063/1.2348881. [ DOI ] [ PubMed ] [ Google Scholar ] Karton A., Daon S., Martin J. M. L.. W4–11: A High-Confidence Benchmark Dataset for Computational Thermochemistry Derived from First-Principles W4 Data. Chem. Phys. Lett. 2011;510:165–178. doi: 10.1016/j.cplett.2011.05.007. [ DOI ] [ Google Scholar ] Xu X., Soriano-Agueda L., López X., Ramos-Cordoba E., Matito E.. How Many Distinct and Reliable Multireference Diagnostics Are There? J. Chem. Phys. 2025;162:124102. doi: 10.1063/5.0250636. [ DOI ] [ PubMed ] [ Google Scholar ] Lee T. J., Taylor P. R.. A Diagnostic for Determining the Quality of Single-Reference Electron Correlation Methods. Int. J. Quantum Chem. 1989;36:199–207. doi: 10.1002/qua.560360824. [ DOI ] [ Google Scholar ] Janssen C. L., Nielsen I. M. B.. New Diagnostics for Coupled-Cluster and Møller-Plesset Perturbation Theory. Chem. Phys. Lett. 1998;290:423–430. doi: 10.1016/S0009-2614(98)00504-1. [ DOI ] [ Google Scholar ] Nielsen I. M. B., Janssen C. L.. Double-Substitution-Based Diagnostics for Coupled-Cluster and Møller-Plesset Perturbation Theory. Chem. Phys. Lett. 1999;310:568–576. doi: 10.1016/S0009-2614(99)00770-8. [ DOI ] [ Google Scholar ] Bartlett R. J., Park Y. C., Bauman N. P., Melnichuk A., Ranasinghe D., Ravi M., Perera A.. Index of Multi-Determinantal and Multi-Reference Character in Coupled-Cluster Theory. J. Chem. Phys. 2020;153:234103. doi: 10.1063/5.0029339. [ DOI ] [ PubMed ] [ Google Scholar ] Faulstich F. M., Kristiansen H. E., Csirik M. A., Kvaal S., Pedersen T. B., Laestadius A.. S-Diagnostic-An a Posteriori Error Assessment for Single-Reference Coupled-Cluster Methods. J. Phys. Chem. A. 2023;127:9106–9120. doi: 10.1021/acs.jpca.3c01575. [ DOI ] [ PubMed ] [ Google Scholar ] Weflen K. E., Bentley M. R., Thorpe J. H., Franke P. R., Martin J. M. L., Matthews D. A., Stanton J. F.. Exploiting a Shortcoming of Coupled-Cluster Theory: The Extent of Non-Hermiticity as a Diagnostic Indicator of Computational Accuracy. J. Phys. Chem. Lett. 2025;16:5121–5127. doi: 10.1021/acs.jpclett.5c00885. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Tishchenko O., Zheng J., Truhlar D. G.. Multireference Model Chemistries for Thermochemical Kinetics. J. Chem. Theory Comput. 2008;4:1208–1219. doi: 10.1021/ct800077r. [ DOI ] [ PubMed ] [ Google Scholar ] Fogueri U. R., Kozuch S., Karton A., Martin J. M. L.. A Simple DFT-based Diagnostic for Nondynamical Correlation. Theor. Chem. Acc. 2013;132:1291. doi: 10.1007/s00214-012-1291-y. [ DOI ] [ Google Scholar ] Grimme S., Hansen A.. A Practicable Real-Space Measure and Visualization of Static Electron-Correlation Effects. Angew. Chem., Int. Ed. 2015;54:12308–12313. doi: 10.1002/anie.201501887. [ DOI ] [ PubMed ] [ Google Scholar ] Ramos-Cordoba E., Salvador P., Matito E.. Separation of Dynamic and Nondynamic Correlation. Phys. Chem. Chem. Phys. 2016;18:24015–24023. doi: 10.1039/C6CP03072F. [ DOI ] [ PubMed ] [ Google Scholar ] Ramos-Cordoba E., Matito E.. Local Descriptors of Dynamic and Nondynamic Correlation. J. Chem. Theory Comput. 2017;13:2705–2711. doi: 10.1021/acs.jctc.7b00293. [ DOI ] [ PubMed ] [ Google Scholar ] Xu X., Soriano-Agueda L., López X., Ramos-Cordoba E., Matito E.. All-Purpose Measure of Electron Correlation for Multireference Diagnostics. J. Chem. Theory Comput. 2024;20:721–727. doi: 10.1021/acs.jctc.3c01073. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Sprague, M. K. ; Irikura, K. K. . Quantitative estimation of uncertainties from wavefunction diagnostics. In Thom H. Dunning, Jr.: A Festschrift from Theoretical Chemistry Accounts; Wilson, A. K. , Peterson, K. A. , Woon, D. E. , Eds.; Springer: Berlin, 2015; pp 307–318. [ Google Scholar ] Kesharwani M. K., Sylvetsky N., Köhn A., Tew D. P., Martin J. M. L.. Do CCSD and Approximate CCSD-F12 Variants Converge to the Same Basis Set Limits? The Case of Atomization Energies. J. Chem. Phys. 2018;149:154109. doi: 10.1063/1.5048665. [ DOI ] [ PubMed ] [ Google Scholar ] Duan C., Liu F., Nandy A., Kulik H. J.. Data-Driven Approaches Can Overcome the Cost-Accuracy Trade-Off in Multireference Diagnostics. J. Chem. Theory Comput. 2020;16:4373–4387. doi: 10.1021/acs.jctc.0c00358. [ DOI ] [ PubMed ] [ Google Scholar ] Liu F., Duan C., Kulik H. J.. Rapid Detection of Strong Correlation with Machine Learning for Transition-Metal Complex High-Throughput Screening. J. Phys. Chem. Lett. 2020;11:8067–8076. doi: 10.1021/acs.jpclett.0c02288. [ DOI ] [ PubMed ] [ Google Scholar ] Duan C., Chu D. B. K., Nandy A., Kulik H. J.. Detection of Multi-Reference Character Imbalances Enables a Transfer Learning Approach for Virtual High Throughput Screening with Coupled Cluster Accuracy at DFT Cost. Chem. Sci. 2022;13:4962–4971. doi: 10.1039/D2SC00393G. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Vuckovic S., Fabiano E., Gori-Giorgi P., Burke K.. MAP: An MP2 Accuracy Predictor for Weak Interactions from Adiabatic Connection Theory. J. Chem. Theory Comput. 2020;16:4141–4149. doi: 10.1021/acs.jctc.0c00049. [ DOI ] [ PubMed ] [ Google Scholar ] Seidl M., Perdew J. P., Levy M.. Strictly Correlated Electrons in Density-Functional Theory. Phys. Rev. A. 1999;59:51–54. doi: 10.1103/PhysRevA.59.51. [ DOI ] [ Google Scholar ] Seidl M., Perdew J. P., Kurth S.. Density Functionals for the Strong-Interaction Limit. Phys. Rev. A. 2000;62:012502. doi: 10.1103/PhysRevA.62.012502. [ DOI ] [ PubMed ] [ Google Scholar ] Saad, Y.
Numerical Methods for Large Eigenvalue Problems, revised ed.; Society for Industrial and Applied Mathematics: Philadelphia, Pa, 2011. [ Google Scholar ] Shavitt, I.
The Method of Configuration Interaction. In Methods of Electronic Structure Theory; Schaefer, H. F. , Ed.; Springer: Boston, MA, 1977; pp 189–275. [ Google Scholar ] Kutzelnigg W.. Error Analysis and Improvements of Coupled-Cluster Theory. Theoret. Chim. Acta. 1991;80:349–386. doi: 10.1007/BF01117418. [ DOI ] [ Google Scholar ] Rohwedder T., Schneider R.. Error Estimates for the Coupled Cluster Method. ESAIM: M2AN. 2013;47:1553–1582. doi: 10.1051/m2an/2013075. [ DOI ] [ Google Scholar ] Hohenberg P., Kohn W.. Inhomogeneous Electron Gas. Phys. Rev. 1964;136:B864–B871. doi: 10.1103/PhysRev.136.B864. [ DOI ] [ Google Scholar ] Kohn W., Sham L. J.. Self-Consistent Equations Including Exchange and Correlation Effects. Phys. Rev. 1965;140:A1133–A1138. doi: 10.1103/PhysRev.140.A1133. [ DOI ] [ Google Scholar ] Parr, R. G. ; Yang, W. . Density-Functional Theory of Atoms and Molecules; Oxford University Press: Oxford, U.K., 1989. [ Google Scholar ] Koch, W. ; Holthausen, M. C. . A Chemist’s Guide to Density Functional Theory, 2nd ed.; Wiley-VCH: Weinheim, Germany, 2001. [ Google Scholar ] Burke K., Wagner L. O.. DFT in a Nutshell. Int. J. Quantum Chem. 2013;113:96–101. doi: 10.1002/qua.24259. [ DOI ] [ Google Scholar ] Lejaeghere K.. et al. Reproducibility in Density Functional Theory Calculations of Solids. Science. 2016;351:aad3000. doi: 10.1126/science.aad3000. [ DOI ] [ PubMed ] [ Google Scholar ] Bosoni E.. et al. How to Verify the Precision of Density-Functional-Theory Implementations via Reproducible and Universal Workflows. Nat. Rev. Phys. 2024;6:45–58. doi: 10.1038/s42254-023-00655-3. [ DOI ] [ Google Scholar ] Chen H., Gong X., He L., Yang Z., Zhou A.. Numerical Analysis of Finite Dimensional Approximations of Kohn-Sham Models. Adv. Comput. Math. 2013;38:225–256. doi: 10.1007/s10444-011-9235-y. [ DOI ] [ Google Scholar ] Cancès E., Dusson G., Maday Y., Stamm B., Vohralík M.. A Perturbation-Method-Based Post-Processing for the Planewave Discretization of Kohn-Sham Models. J. Comput. Phys. 2016;307:446–459. doi: 10.1016/j.jcp.2015.12.012. [ DOI ] [ Google Scholar ] Cancès E., Chakir R., Maday Y.. Numerical Analysis of the Planewave Discretization of Some Orbital-Free and Kohn-Sham Models. ESAIM: M2AN. 2012;46:341–388. doi: 10.1051/m2an/2011038. [ DOI ] [ Google Scholar ] Chen H., Schneider R.. Error Estimates of Some Numerical Atomic Orbitals in Molecular Simulations. Commun. Comput. Phys. 2015;18:125–146. doi: 10.4208/cicp.170414.231214a. [ DOI ] [ Google Scholar ] Chen H., Schneider R.. Numerical Analysis of Augmented Plane Wave Methods for Full-Potential Electronic Structure Calculations. ESAIM: M2AN. 2015;49:755–785. doi: 10.1051/m2an/2014052. [ DOI ] [ Google Scholar ] Rohwedder T., Schneider R.. An Analysis for the DIIS Acceleration Method Used in Quantum Chemistry Calculations. J. Math. Chem. 2011;49:1889–1914. doi: 10.1007/s10910-011-9863-y. [ DOI ] [ Google Scholar ] Chupin M., Dupuy M.-S., Legendre G., Séré É.. Convergence Analysis of Adaptive DIIS Algorithms with Application to Electronic Ground State Calculations. ESAIM: M2AN. 2021;55:2785–2825. doi: 10.1051/m2an/2021069. [ DOI ] [ Google Scholar ] Cancès É., Ehrlacher V., Gontier D., Levitt A., Lombardi D.. Numerical Quadrature in the Brillouin Zone for Periodic Schrödinger Operators. Numer. Math. 2020;144:479–526. doi: 10.1007/s00211-019-01096-w. [ DOI ] [ Google Scholar ] Cohen A. J., Mori-Sánchez P., Yang W.. Challenges for Density Functional Theory. Chem. Rev. 2012;112:289–320. doi: 10.1021/cr200107z. [ DOI ] [ PubMed ] [ Google Scholar ] Becke A. D.. Perspective: Fifty Years of Density-Functional Theory in Chemical Physics. J. Chem. Phys. 2014;140:18A301. doi: 10.1063/1.4869598. [ DOI ] [ PubMed ] [ Google Scholar ] Toulouse, J.
Review of Approximations for the Exchange–Correlation Energy in Density-Functional Theory. In Density Functional Theory: Modeling, Mathematical Analysis, Computational Methods, and Applications; Cancès, E. , Friesecke, G. , Eds.; Springer International Publishing: Cham, Switzerland, 2023; pp 1–90. [ Google Scholar ] Goerigk L., Hansen A., Bauer C., Ehrlich S., Najibi A., Grimme S.. A Look at the Density Functional Theory Zoo with the Advanced GMTKN55 Database for General Main Group Thermochemistry, Kinetics and Noncovalent Interactions. Phys. Chem. Chem. Phys. 2017;19:32184–32215. doi: 10.1039/C7CP04913G. [ DOI ] [ PubMed ] [ Google Scholar ] Mardirossian N., Head-Gordon M.. Thirty Years of Density Functional Theory in Computational Chemistry: An Overview and Extensive Assessment of 200 Density Functionals. Mol. Phys. 2017;115:2315–2372. doi: 10.1080/00268976.2017.1333644. [ DOI ] [ Google Scholar ] Schultz N. E., Zhao Y., Truhlar D. G.. Density Functionals for Inorganometallic and Organometallic Chemistry. J. Phys. Chem. A. 2005;109:11127–11143. doi: 10.1021/jp0539223. [ DOI ] [ PubMed ] [ Google Scholar ] Mortensen J. J., Kaasbjerg K., Frederiksen S. L., Nørskov J. K., Sethna J. P., Jacobsen K. W.. Bayesian Error Estimation in Density-Functional Theory. Phys. Rev. Lett. 2005;95:216401. doi: 10.1103/PhysRevLett.95.216401. [ DOI ] [ PubMed ] [ Google Scholar ] Aldegunde M., Kermode J. R., Zabaras N.. Development of an Exchange-Correlation Functional with Uncertainty Quantification Capabilities for Density Functional Theory. J. Comput. Phys. 2016;311:173–195. doi: 10.1016/j.jcp.2016.01.034. [ DOI ] [ Google Scholar ] Wellendorff J., Lundgaard K. T., Møgelhøj A., Petzold V., Landis D. D., Nørskov J. K., Bligaard T., Jacobsen K. W.. Density Functionals for Surface Science: Exchange-correlation Model Development with Bayesian Error Estimation. Phys. Rev. B. 2012;85:235149. doi: 10.1103/PhysRevB.85.235149. [ DOI ] [ Google Scholar ] Wellendorff J., Lundgaard K. T., Jacobsen K. W., Bligaard T.. mBEEF: An Accurate Semi-Local Bayesian Error Estimation Density Functional. J. Chem. Phys. 2014;140:144107. doi: 10.1063/1.4870397. [ DOI ] [ PubMed ] [ Google Scholar ] Parks H. L., McGaughey A. J. H., Viswanathan V.. Uncertainty Quantification in First-Principles Predictions of Harmonic Vibrational Frequencies of Molecules and Molecular Complexes. J. Phys. Chem. C. 2019;123:4072–4084. doi: 10.1021/acs.jpcc.8b11689. [ DOI ] [ Google Scholar ] Guan P.-W., Houchins G., Viswanathan V.. Uncertainty Quantification of DFT-predicted Finite Temperature Thermodynamic Properties within the Debye Model. J. Chem. Phys. 2019;151:244702. doi: 10.1063/1.5132332. [ DOI ] [ PubMed ] [ Google Scholar ] Medford A. J., Wellendorff J., Vojvodic A., Studt F., Abild-Pedersen F., Jacobsen K. W., Bligaard T., Nørskov J. K.. Assessing the Reliability of Calculated Catalytic Ammonia Synthesis Rates. Science. 2014;345:197–200. doi: 10.1126/science.1253486. [ DOI ] [ PubMed ] [ Google Scholar ] Pandey M., Jacobsen K. W.. Heats of Formation of Solids with Error Estimation: The mBEEF Functional with and without Fitted Reference Energies. Phys. Rev. B. 2015;91:235201. doi: 10.1103/PhysRevB.91.235201. [ DOI ] [ Google Scholar ] Deshpande S., Kitchin J. R., Viswanathan V.. Quantifying Uncertainty in Activity Volcano Relationships for Oxygen Reduction Reaction. ACS Catal. 2016;6:5251–5259. doi: 10.1021/acscatal.6b00509. [ DOI ] [ Google Scholar ] Wang B., Chen S., Zhang J., Li S., Yang B.. Propagating DFT Uncertainty to Mechanism Determination, Degree of Rate Control, and Coverage Analysis: The Kinetics of Dry Reforming of Methane. J. Phys. Chem. C. 2019;123:30389–30397. doi: 10.1021/acs.jpcc.9b08755. [ DOI ] [ Google Scholar ] Christensen, R. ; Bligaard, T. ; Jacobsen, K. W. . Bayesian error estimation in density functional theory. In Uncertainty Quantification in Multiscale Materials Modeling; Wang, Y. , McDowell, D. L. , Eds.; Elsevier Series in Mechanics of Advanced Materials; Woodhead Publishing, 2020; pp 77–91. [ Google Scholar ] Yang K., Yang B.. Addressing the Uncertainty of DFT-determined Hydrogenation Mechanisms over Coinage Metal Surfaces. Faraday Discuss. 2021;229:50–61. doi: 10.1039/C9FD00122K. [ DOI ] [ PubMed ] [ Google Scholar ] Tsiverioti L. M., Kavalsky L., Viswanathan V.. Robust Analysis of 4e– versus 6e– Reduction of Nitrogen on Metal Surfaces and Single-Atom Alloys. J. Phys. Chem. C. 2022;126:12994–13003. doi: 10.1021/acs.jpcc.2c01630. [ DOI ] [ Google Scholar ] Kreitz B., Lott P., Studt F., Medford A. J., Deutschmann O., Goldsmith C. F.. Automated Generation of Microkinetics for Heterogeneously Catalyzed Reactions Considering Correlated Uncertainties. Angew. Chem., Int. Ed. 2023;62:e202306514. doi: 10.1002/anie.202306514. [ DOI ] [ PubMed ] [ Google Scholar ] Pernot P.. The Parameter Uncertainty Inflation Fallacy. J. Chem. Phys. 2017;147:104102. doi: 10.1063/1.4994654. [ DOI ] [ PubMed ] [ Google Scholar ] Pernot P., Cailliez F.. A Critical Review of Statistical Calibration/Prediction Models Handling Data Inconsistency and Model Inadequacy. AIChE J. 2017;63:4642–4665. doi: 10.1002/aic.15781. [ DOI ] [ Google Scholar ] Proppe J., Husch T., Simm G. N., Reiher M.. Uncertainty Quantification for Quantum Chemical Models of Complex Reaction Networks. Faraday Discuss. 2016;195:497–520. doi: 10.1039/C6FD00144K. [ DOI ] [ PubMed ] [ Google Scholar ] Walker E., Ammal S. C., Terejanu G. A., Heyden A.. Uncertainty Quantification Framework Applied to the Water-Gas Shift Reaction over Pt-Based Catalysts. J. Phys. Chem. C. 2016;120:10328–10339. doi: 10.1021/acs.jpcc.6b01348. [ DOI ] [ Google Scholar ] Walker E. A., Mitchell D., Terejanu G. A., Heyden A.. Identifying Active Sites of the Water-Gas Shift Reaction over Titania Supported Platinum Catalysts under Uncertainty. ACS Catal. 2018;8:3990–3998. doi: 10.1021/acscatal.7b03531. [ DOI ] [ Google Scholar ] Bello M., Bamidele O. H., Terejanu G., Heyden A.. Investigation of Ethane Dehydrogenation and Hydrogenolysis on Pt(111), Pt(211), and Pt(100): Bayesian Quantification and Correction of DFT-Based Enthalpic and Entropic Uncertainties. ACS Catal. 2024;14:15528–15544. doi: 10.1021/acscatal.4c03455. [ DOI ] [ Google Scholar ] Ramakrishnan R., Dral P. O., Rupp M., von Lilienfeld O. A.. Big Data Meets Quantum Chemistry Approximations: The Δ-Machine Learning Approach. J. Chem. Theory Comput. 2015;11:2087–2096. doi: 10.1021/acs.jctc.5b00099. [ DOI ] [ PubMed ] [ Google Scholar ] Lejaeghere K., Van Speybroeck V., Van Oost G., Cottenier S.. Error Estimates for Solid-State Density-Functional Theory Predictions: An Overview by Means of the Ground-State Elemental Crystals. Crit. Rev. Solid State Mater. Sci. 2014;39:1–24. doi: 10.1080/10408436.2013.772503. [ DOI ] [ Google Scholar ] Pernot P., Civalleri B., Presti D., Savin A.. Prediction Uncertainty of Density Functional Approximations for Properties of Crystals with Cubic Symmetry. J. Phys. Chem. A. 2015;119:5288–5304. doi: 10.1021/jp509980w. [ DOI ] [ PubMed ] [ Google Scholar ] Hautier G., Ong S. P., Jain A., Moore C. J., Ceder G.. Accuracy of Density Functional Theory in Predicting Formation Energies of Ternary Oxides from Binary Oxides and Its Implication on Phase Stability. Phys. Rev. B. 2012;85:155208. doi: 10.1103/PhysRevB.85.155208. [ DOI ] [ Google Scholar ] Yu Y., Aykol M., Wolverton C.. Reaction Thermochemistry of Metal Sulfides with GGA and GGA+U Calculations. Phys. Rev. B. 2015;92:195118. doi: 10.1103/PhysRevB.92.195118. [ DOI ] [ Google Scholar ] Grindy S., Meredig B., Kirklin S., Saal J. E., Wolverton C.. Approaching Chemical Accuracy with Density Functional Calculations: Diatomic Energy Corrections. Phys. Rev. B. 2013;87:075150. doi: 10.1103/PhysRevB.87.075150. [ DOI ] [ Google Scholar ] Wang A., Kingsbury R., McDermott M., Horton M., Jain A., Ong S. P., Dwaraknath S., Persson K. A.. A Framework for Quantifying Uncertainty in DFT Energy Corrections. Sci. Rep. 2021;11:15496. doi: 10.1038/s41598-021-94550-5. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Wang L., Maxisch T., Ceder G.. Oxidation Energies of Transition Metal Oxides within the GGA+U Framework. Phys. Rev. B. 2006;73:195107. doi: 10.1103/PhysRevB.73.195107. [ DOI ] [ Google Scholar ] Jain A., Hautier G., Ong S. P., Moore C. J., Fischer C. C., Persson K. A., Ceder G.. Formation Enthalpies by Mixing GGA and GGA+U Calculations. Phys. Rev. B. 2011;84:045115. doi: 10.1103/PhysRevB.84.045115. [ DOI ] [ Google Scholar ] Simm G. N., Reiher M.. Error-Controlled Exploration of Chemical Reaction Networks with Gaussian Processes. J. Chem. Theory Comput. 2018;14:5238–5248. doi: 10.1021/acs.jctc.8b00504. [ DOI ] [ PubMed ] [ Google Scholar ] Yuk S. F., Sargin I., Meyer N., Krogel J. T., Beckman S. P., Cooper V. R.. Putting Error Bars on Density Functional Theory. Sci. Rep. 2024;14:20219. doi: 10.1038/s41598-024-69194-w. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Alfonso-Ramos J. E., Adamo C., Brémond É., Stuyver T.. Improving the Reliability of, and Confidence in, DFT Functional Benchmarking through Active Learning. J. Chem. Theory Comput. 2025;21:1752–1761. doi: 10.1021/acs.jctc.4c01729. [ DOI ] [ PubMed ] [ Google Scholar ] Duan C., Nandy A., Meyer R., Arunachalam N., Kulik H. J.. A Transferable Recommender Approach for Selecting the Best Density Functional Approximations in Chemical Discovery. Nat. Comput. Sci. 2023;3:38–47. doi: 10.1038/s43588-022-00384-0. [ DOI ] [ PubMed ] [ Google Scholar ] Grimme S.. Density Functional Theory with London Dispersion Corrections. Wiley Interdiscip. Rev.: Comput. Mol. Sci. 2011;1:211–228. doi: 10.1002/wcms.30. [ DOI ] [ Google Scholar ] Grimme S., Hansen A., Brandenburg J. G., Bannwarth C.. Dispersion-Corrected Mean-Field Electronic Structure Methods. Chem. Rev. 2016;116:5105–5154. doi: 10.1021/acs.chemrev.5b00533. [ DOI ] [ PubMed ] [ Google Scholar ] Grimme S.. Semiempirical GGA-type Density Functional Constructed with a Long-Range Dispersion Correction. J. Comput. Chem. 2006;27:1787–1799. doi: 10.1002/jcc.20495. [ DOI ] [ PubMed ] [ Google Scholar ] Tkatchenko A., Scheffler M.. Accurate Molecular Van Der Waals Interactions from Ground-State Electron Density and Free-Atom Reference Data. Phys. Rev. Lett. 2009;102:073005. doi: 10.1103/PhysRevLett.102.073005. [ DOI ] [ PubMed ] [ Google Scholar ] Hanke F.. Sensitivity Analysis and Uncertainty Calculation for Dispersion Corrected Density Functional Theory. J. Comput. Chem. 2011;32:1424–1430. doi: 10.1002/jcc.21724. [ DOI ] [ PubMed ] [ Google Scholar ] Grimme S., Antony J., Ehrlich S., Krieg H.. A Consistent and Accurate Ab Initio Parametrization of Density Functional Dispersion Correction (DFT-D) for the 94 Elements H-Pu. J. Chem. Phys. 2010;132:154104. doi: 10.1063/1.3382344. [ DOI ] [ PubMed ] [ Google Scholar ] Weymuth T., Proppe J., Reiher M.. Statistical Analysis of Semiclassical Dispersion Corrections. J. Chem. Theory Comput. 2018;14:2480–2494. doi: 10.1021/acs.jctc.8b00078. [ DOI ] [ PubMed ] [ Google Scholar ] Proppe J., Reiher M.. Mechanism Deduction from Noisy Chemical Reaction Networks. J. Chem. Theory Comput. 2019;15:357–370. doi: 10.1021/acs.jctc.8b00310. [ DOI ] [ PubMed ] [ Google Scholar ] Settles, B.
Active Learning; Synthesis Lectures on Artificial Intelligence and Machine Learning, Vol. 18; Morgan & Claypool: San Rafael, CA, 2012. [ Google Scholar ] Surzhikova E., Proppe J.. regAL: Python Package for Active Learning of Regression Problems. Mach. Learn.: Sci. Technol. 2025;6:025064. doi: 10.1088/2632-2153/addf11. [ DOI ] [ Google Scholar ] Thiel W.. Semiempirical Quantum-Chemical Methods. Wiley Interdiscip. Rev.: Comput. Mol. Sci. 2014;4:145–157. doi: 10.1002/wcms.1161. [ DOI ] [ Google Scholar ] Cui Q., Elstner M.. Density Functional Tight Binding: Values of Semi-Empirical Methods in an Ab Initio Era. Phys. Chem. Chem. Phys. 2014;16:14368–14377. doi: 10.1039/C4CP00908H. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Bannwarth C., Caldeweyher E., Ehlert S., Hansen A., Pracht P., Seibert J., Spicher S., Grimme S.. Extended Tight-Binding Quantum Chemistry Methods. Wiley Interdiscip. Rev.: Comput. Mol. Sci. 2021;11:e1493. doi: 10.1002/wcms.1493. [ DOI ] [ Google Scholar ] Oreluk J., Liu Z., Hegde A., Li W., Packard A., Frenklach M., Zubarev D.. Diagnostics of Data-Driven Models: Uncertainty Quantification of PM7 Semi-Empirical Quantum Chemical Method. Sci. Rep. 2018;8:13248. doi: 10.1038/s41598-018-31677-y. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Frenklach, M. ; Packard, A. ; Seiler, P. . Prediction Uncertainty from Models and Data. In Proceedings of the 2002 American Control Conference; IEEE, 2002; Vol. 5, pp 4135–4140. 10.1109/ACC.2002.1024578. [ DOI ] [ Google Scholar ] Seiler P., Frenklach M., Packard A., Feeley R.. Numerical Approaches for Collaborative Data Processing. Optim. Eng. 2006;7:459–478. doi: 10.1007/s11081-006-0350-4. [ DOI ] [ Google Scholar ] Russi T., Packard A., Frenklach M.. Uncertainty Quantification: Making Predictions of Complex Reaction Systems Reliable. Chem. Phys. Lett. 2010;499:1–8. doi: 10.1016/j.cplett.2010.09.009. [ DOI ] [ Google Scholar ] Leach, A. R.
Molecular Modelling: Principles and Applications; Pearson Education, 2001. [ Google Scholar ] Brunken C., Reiher M.. Self-Parametrizing System-Focused Atomistic Models. J. Chem. Theory Comput. 2020;16:1646–1665. doi: 10.1021/acs.jctc.9b00855. [ DOI ] [ PubMed ] [ Google Scholar ] Friederich P., Häse F., Proppe J., Aspuru-Guzik A.. Machine-Learned Potentials for next-Generation Matter Simulations. Nat. Mater. 2021;20:750–761. doi: 10.1038/s41563-020-0777-6. [ DOI ] [ PubMed ] [ Google Scholar ] Behler J.. Four Generations of High-Dimensional Neural Network Potentials. Chem. Rev. 2021;121:10037–10072. doi: 10.1021/acs.chemrev.0c00868. [ DOI ] [ PubMed ] [ Google Scholar ] Zhang L., Han J., Wang H., Car R., E W.. Deep Potential Molecular Dynamics: A Scalable Model with the Accuracy of Quantum Mechanics. Phys. Rev. Lett. 2018;120:143001. doi: 10.1103/PhysRevLett.120.143001. [ DOI ] [ PubMed ] [ Google Scholar ] Gaiduk A. P., Gustafson J., Gygi F., Galli G.. First-Principles Simulations of Liquid Water Using a Dielectric-Dependent Hybrid Functional. J. Phys. Chem. Lett. 2018;9:3068–3073. doi: 10.1021/acs.jpclett.8b01017. [ DOI ] [ PubMed ] [ Google Scholar ] Zaverkin V., Holzmüller D., Steinwart I., Kästner J.. Exploring Chemical and Conformational Spaces by Batch Mode Deep Active Learning. Digit. Discovery. 2022;1:605–620. doi: 10.1039/D2DD00034B. [ DOI ] [ Google Scholar ] Seung, H. S. ; Opper, M. ; Sompolinsky, H. . Query by Committee. In Proceedings of the Fifth Annual Workshop on Computational Learning Theory; Association for Computing Machinery: New York, 1992; pp 287–294. [ Google Scholar ] Kulichenko M., Nebgen B., Lubbers N., Smith J. S., Barros K., Allen A. E. A., Habib A., Shinkle E., Fedik N., Li Y. W., Messerly R. A., Tretiak S.. Data Generation for Machine Learning Interatomic Potentials and Beyond. Chem. Rev. 2024;124:13681–13714. doi: 10.1021/acs.chemrev.4c00572. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Vazquez-Salazar L. I., Boittier E. D., Meuwly M.. Uncertainty Quantification for Predictions of Atomistic Neural Networks. Chem. Sci. 2022;13:13068–13084. doi: 10.1039/D2SC04056E. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Soleimany A. P., Amini A., Goldman S., Rus D., Bhatia S. N., Coley C. W.. Evidential Deep Learning for Guided Molecular Property Prediction and Discovery. ACS Cent. Sci. 2021;7:1356–1367. doi: 10.1021/acscentsci.1c00546. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Drautz R.. Atomic Cluster Expansion for Accurate and Transferable Interatomic Potentials. Phys. Rev. B. 2019;99:014104. doi: 10.1103/PhysRevB.99.014104. [ DOI ] [ Google Scholar ] Lysogorskiy Y., van der Oord C., Bochkarev A., Menon S., Rinaldi M., Hammerschmidt T., Mrovec M., Thompson A., Csányi G., Ortner C., Drautz R.. Performant Implementation of the Atomic Cluster Expansion (PACE) and Application to Copper and Silicon. npj Comput. Mater. 2021;7:97. doi: 10.1038/s41524-021-00559-9. [ DOI ] [ Google Scholar ] Witt W. C., Van Der Oord C., Gelžinytė E., Järvinen T., Ross A., Darby J. P., Ho C. H., Baldwin W. J., Sachs M., Kermode J., Bernstein N., Csányi G., Ortner C.. ACEpotentials.jl: A Julia Implementation of the Atomic Cluster Expansion. J. Chem. Phys. 2023;159:164101. doi: 10.1063/5.0158783. [ DOI ] [ PubMed ] [ Google Scholar ] Gawlikowski J., Tassi C. R. N., Ali M., Lee J., Humt M., Feng J., Kruspe A., Triebel R., Jung P., Roscher R., Shahzad M., Yang W., Bamler R., Zhu X. X.. A Survey of Uncertainty in Deep Neural Networks. Artif. Intell. Rev. 2023;56:1513–1589. doi: 10.1007/s10462-023-10562-9. [ DOI ] [ Google Scholar ] Grasselli F., Kapil V., Bonfanti S., Rossi K., Chong S.. Uncertainty in the Era of Machine Learning for Atomistic Modeling. Digit. Discovery. 2025;4:2654–2675. doi: 10.1039/D5DD00102A. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Musil F., Willatt M. J., Langovoy M. A., Ceriotti M.. Fast and Accurate Uncertainty Estimation in Chemical Machine Learning. J. Chem. Theory Comput. 2019;15:906–915. doi: 10.1021/acs.jctc.8b00959. [ DOI ] [ PubMed ] [ Google Scholar ] Raimbault N., Grisafi A., Ceriotti M., Rossi M.. Using Gaussian Process Regression to Simulate the Vibrational Raman Spectra of Molecular Crystals. New J. Phys. 2019;21:105001. doi: 10.1088/1367-2630/ab4509. [ DOI ] [ Google Scholar ] Rossi K., Jurásková V., Wischert R., Garel L., Corminbœuf C., Ceriotti M.. Simulating Solvation and Acidity in Complex Mixtures with First-Principles Accuracy: The Case of CH3SO3H and H2O2 in Phenol. J. Chem. Theory Comput. 2020;16:5139–5149. doi: 10.1021/acs.jctc.0c00362. [ DOI ] [ PubMed ] [ Google Scholar ] Imbalzano G., Zhuang Y., Kapil V., Rossi K., Engel E. A., Grasselli F., Ceriotti M.. Uncertainty Estimation for Molecular Dynamics and Sampling. J. Chem. Phys. 2021;154:074102. doi: 10.1063/5.0036522. [ DOI ] [ PubMed ] [ Google Scholar ] Zheng P., Yang W., Wu W., Isayev O., Dral P. O.. Toward Chemical Accuracy in Predicting Enthalpies of Formation with General-Purpose Data-Driven Methods. J. Phys. Chem. Lett. 2022;13:3479–3491. doi: 10.1021/acs.jpclett.2c00734. [ DOI ] [ PubMed ] [ Google Scholar ] Kellner M., Ceriotti M.. Uncertainty Quantification by Direct Propagation of Shallow Ensembles. Mach. Learn.: Sci. Technol. 2024;5:035006. doi: 10.1088/2632-2153/ad594a. [ DOI ] [ Google Scholar ] Carrete J., Montes-Campos H., Wanzenböck R., Heid E., Madsen G. K. H.. Deep Ensembles vs Committees for Uncertainty Estimation in Neural-Network Force Fields: Comparison and Application to Active Learning. J. Chem. Phys. 2023;158:204801. doi: 10.1063/5.0146905. [ DOI ] [ PubMed ] [ Google Scholar ] Lin Q., Zhang L., Zhang Y., Jiang B.. Searching Configurations in Uncertainty Space: Active Learning of High-Dimensional Neural Network Reactive Potentials. J. Chem. Theory Comput. 2021;17:2691–2701. doi: 10.1021/acs.jctc.1c00166. [ DOI ] [ PubMed ] [ Google Scholar ] Gal Y., Ghahramani Z.. Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning. Proc. Mach. Learn. Res. 2016;48:1050–1059. [ Google Scholar ] Wen M., Tadmor E. B.. Uncertainty Quantification in Molecular Simulations with Dropout Neural Network Potentials. npj Comput. Mater. 2020;6:124. doi: 10.1038/s41524-020-00390-8. [ DOI ] [ Google Scholar ] Zaverkin V., Holzmüller D., Christiansen H., Errica F., Alesiani F., Takamoto M., Niepert M., Kästner J.. Uncertainty-Biased Molecular Dynamics for Learning Uniformly Accurate Interatomic Potentials. npj Comput. Mater. 2024;10:83. doi: 10.1038/s41524-024-01254-1. [ DOI ] [ Google Scholar ] Lysogorskiy Y., Bochkarev A., Mrovec M., Drautz R.. Active Learning Strategies for Atomic Cluster Expansion Models. Phys. Rev. Materials. 2023;7:043801. doi: 10.1103/PhysRevMaterials.7.043801. [ DOI ] [ Google Scholar ] Bochkarev A., Lysogorskiy Y., Menon S., Qamar M., Mrovec M., Drautz R.. Efficient Parametrization of the Atomic Cluster Expansion. Phys. Rev. Mater. 2022;6:013804. doi: 10.1103/PhysRevMaterials.6.013804. [ DOI ] [ Google Scholar ] Best I. R., Sullivan T. J., Kermode J. R.. Uncertainty Quantification in Atomistic Simulations of Silicon Using Interatomic Potentials. J. Chem. Phys. 2024;161:064112. doi: 10.1063/5.0214590. [ DOI ] [ PubMed ] [ Google Scholar ] van der Oord C., Sachs M., Kovács D. P., Ortner C., Csányi G.. Hyperactive Learning for Data-Driven Interatomic Potentials. npj Comput. Mater. 2023;9:168. doi: 10.1038/s41524-023-01104-6. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Kulichenko M., Barros K., Lubbers N., Li Y. W., Messerly R., Tretiak S., Smith J. S., Nebgen B.. Uncertainty-Driven Dynamics for Active Learning of Interatomic Potentials. Nat. Comput. Sci. 2023;3:230–239. doi: 10.1038/s43588-023-00406-5. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Bartók, A. P. ; Kermode, J. R. . Improved Uncertainty Quantification for Gaussian Process Regression Based Interatomic Potentials. arXiv (Condensed Matter.Materials Science), June 17, 2022, 2206.08744, ver. 1. https://arxiv.org/abs/2206.08744 . Botu V., Batra R., Chapman J., Ramprasad R.. Machine Learning Force Fields: Construction, Validation, and Outlook. J. Phys. Chem. C. 2017;121:511–522. doi: 10.1021/acs.jpcc.6b10908. [ DOI ] [ Google Scholar ] Chernatynskiy A., Phillpot S. R., LeSar R.. Uncertainty Quantification in Multiscale Simulation of Materials: A Prospective. Annu. Rev. Mater. Res. 2013;43:157–182. doi: 10.1146/annurev-matsci-071312-121708. [ DOI ] [ Google Scholar ] Wang, Y. ; McDowell, D. L. . Uncertainty quantification in materials modeling. In Uncertainty Quantification in Multiscale Materials Modeling; Wang, Y. , McDowell, D. L. , Eds.; Elsevier Series in Mechanics of Advanced Materials; Woodhead Publishing, 2020; pp 1–40. [ Google Scholar ] Honarmandi P., Arróyave R.. Uncertainty Quantification and Propagation in Computational Materials Science and Simulation-Assisted Materials Design. Integr. Mater. Manuf. Innov. 2020;9:103–143. doi: 10.1007/s40192-020-00168-2. [ DOI ] [ Google Scholar ] Houchins G., Krishnamurthy D., Viswanathan V.. The Role of Uncertainty Quantification and Propagation in Accelerating the Discovery of Electrochemical Functional Materials. MRS Bull. 2019;44:204–212. doi: 10.1557/mrs.2019.45. [ DOI ] [ Google Scholar ] Ye D., Veen L., Nikishova A., Lakhlili J., Edeling W., Luk O. O., Krzhizhanovskaya V. V., Hoekstra A. G.. Uncertainty Quantification Patterns for Multiscale Models. Philos. Trans. R. Soc. A. 2021;379:20200072. doi: 10.1098/rsta.2020.0072. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Warshel A., Levitt M.. Theoretical Studies of Enzymic Reactions: Dielectric, Electrostatic and Steric Stabilization of the Carbonium Ion in the Reaction of Lysozyme. J. Mol. Biol. 1976;103:227–249. doi: 10.1016/0022-2836(76)90311-9. [ DOI ] [ PubMed ] [ Google Scholar ] Field M. J., Bash P. A., Karplus M.. A Combined Quantum Mechanical and Molecular Mechanical Potential for Molecular Dynamics Simulations. J. Comput. Chem. 1990;11:700–733. doi: 10.1002/jcc.540110605. [ DOI ] [ Google Scholar ] Gao J.. Methods and Applications of Combined Quantum Mechanical and Molecular Mechanical Potentials. Rev. Comput. Chem. 1996;7:119–185. doi: 10.1002/9780470125847.ch3. [ DOI ] [ Google Scholar ] Severo Pereira Gomes A., Jacob C. R.. Quantum-Chemical Embedding Methods for Treating Local Electronic Excitations in Complex Chemical Systems. Annu. Rep. Prog. Chem., Sect. C: Phys. Chem. 2012;108:222. doi: 10.1039/c2pc90007f. [ DOI ] [ Google Scholar ] Chung L. W., Sameera W. M. C., Ramozzi R., Page A. J., Hatanaka M., Petrova G. P., Harris T. V., Li X., Ke Z., Liu F., Li H.-B., Ding L., Morokuma K.. The ONIOM Method and Its Applications. Chem. Rev. 2015;115:5678–5796. doi: 10.1021/cr5004419. [ DOI ] [ PubMed ] [ Google Scholar ] Jacob C. R., Neugebauer J.. Subsystem Density-Functional Theory (Update) Wiley Interdiscip. Rev.: Comput. Mol. Sci. 2024;14:e1700. doi: 10.1002/wcms.1700. [ DOI ] [ Google Scholar ] Gordon M. S., Fedorov D. G., Pruitt S. R., Slipchenko L. V.. Fragmentation Methods: A Route to Accurate Calculations on Large Systems. Chem. Rev. 2012;112:632–672. doi: 10.1021/cr200093j. [ DOI ] [ PubMed ] [ Google Scholar ] Herbert J. M.. Fantasy versus Reality in Fragment-Based Quantum Chemistry. J. Chem. Phys. 2019;151:170901. doi: 10.1063/1.5126216. [ DOI ] [ PubMed ] [ Google Scholar ] Richard R. M., Lao K. U., Herbert J. M.. Understanding the Many-Body Expansion for Large Systems. I. Precision Considerations. J. Chem. Phys. 2014;141:014108. doi: 10.1063/1.4885846. [ DOI ] [ PubMed ] [ Google Scholar ] Dahlke E. E., Truhlar D. G.. Electrostatically Embedded Many-Body Expansion for Large Systems, with Applications to Water Clusters. J. Chem. Theory Comput. 2007;3:46–53. doi: 10.1021/ct600253j. [ DOI ] [ PubMed ] [ Google Scholar ] Liu J., Qi L.-W., Zhang J. Z. H., He X.. Fragment Quantum Mechanical Method for Large-Sized Ion-Water Clusters. J. Chem. Theory Comput. 2017;13:2021–2034. doi: 10.1021/acs.jctc.7b00149. [ DOI ] [ PubMed ] [ Google Scholar ] Liu K.-Y., Herbert J. M.. Understanding the Many-Body Expansion for Large Systems. III. Critical Role of Four-Body Terms, Counterpoise Corrections, and Cutoffs. J. Chem. Phys. 2017;147:161729. doi: 10.1063/1.4986110. [ DOI ] [ PubMed ] [ Google Scholar ] Ouyang J. F., Bettens R. P. A.. When Are Many-Body Effects Significant? J. Chem. Theory Comput. 2016;12:5860–5867. doi: 10.1021/acs.jctc.6b00864. [ DOI ] [ PubMed ] [ Google Scholar ] Liu K.-Y., Herbert J. M.. Energy-Screened Many-Body Expansion: A Practical Yet Accurate Fragmentation Method for Quantum Chemistry. J. Chem. Theory Comput. 2020;16:475–487. doi: 10.1021/acs.jctc.9b01095. [ DOI ] [ PubMed ] [ Google Scholar ] Broderick D. R., Herbert J. M.. Scalable Generalized Screening for High-Order Terms in the Many-Body Expansion: Algorithm, Open-Source Implementation, and Demonstration. J. Chem. Phys. 2023;159:174801. doi: 10.1063/5.0174293. [ DOI ] [ PubMed ] [ Google Scholar ] Eriksen J. J., Lipparini F., Gauss J.. Virtual Orbital Many-Body Expansions: A Possible Route towards the Full Configuration Interaction Limit. J. Phys. Chem. Lett. 2017;8:4633–4639. doi: 10.1021/acs.jpclett.7b02075. [ DOI ] [ PubMed ] [ Google Scholar ] Greiner J., Gauss J., Eriksen J. J.. Error Control and Automatic Detection of Reference Active Spaces in Many-Body Expanded Full Configuration Interaction. J. Phys. Chem. A. 2024;128:6806–6818. doi: 10.1021/acs.jpca.4c04056. [ DOI ] [ PubMed ] [ Google Scholar ] Senn H. M., Thiel W.. QM/MM Methods for Biological Systems. Top. Curr. Chem. 2007;268:173–290. doi: 10.1007/128_2006_084. [ DOI ] [ Google Scholar ] Senn H. M., Thiel W.. QM/MM Methods for Biomolecular Systems. Angew. Chem., Int. Ed. 2009;48:1198–1229. doi: 10.1002/anie.200802019. [ DOI ] [ PubMed ] [ Google Scholar ] Groenhof G.. Introduction to QM/MM Simulations. Methods Mol. Biol. 2013;924:43–66. doi: 10.1007/978-1-62703-017-5_3. [ DOI ] [ PubMed ] [ Google Scholar ] Ryde U.. QM/MM Calculations on Proteins. Methods Enzymol. 2016;577:119–158. doi: 10.1016/bs.mie.2016.05.014. [ DOI ] [ PubMed ] [ Google Scholar ] Cui Q., Pal T., Xie L.. Biomolecular QM/MM Simulations: What Are Some of the “Burning Issues”? J. Phys. Chem. B. 2021;125:689–702. doi: 10.1021/acs.jpcb.0c09898. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Giudetti G., Polyakov I., Grigorenko B. L., Faraji S., Nemukhin A. V., Krylov A. I.. How Reproducible Are QM/MM Simulations? Lessons from Computational Studies of the Covalent Inhibition of the SARS-CoV-2 Main Protease by Carmofur. J. Chem. Theory Comput. 2022;18:5056–5067. doi: 10.1021/acs.jctc.2c00286. [ DOI ] [ PubMed ] [ Google Scholar ] Ho J., Yu H., Shao Y., Taylor M., Chen J.. How Accurate Are QM/MM Models? J. Phys. Chem. A. 2025;129:1517–1528. doi: 10.1021/acs.jpca.4c06521. [ DOI ] [ PubMed ] [ Google Scholar ] Sumowski C. V., Ochsenfeld C.. A Convergence Study of QM/MM Isomerization Energies with the Selected Size of the QM Region for Peptidic Systems. J. Phys. Chem. A. 2009;113:11734–11741. doi: 10.1021/jp902876n. [ DOI ] [ PubMed ] [ Google Scholar ] Hu L., Söderhjelm P., Ryde U.. On the Convergence of QM/MM Energies. J. Chem. Theory Comput. 2011;7:761–777. doi: 10.1021/ct100530r. [ DOI ] [ PubMed ] [ Google Scholar ] Flaig D., Beer M., Ochsenfeld C.. Convergence of Electronic Structure with the Size of the QM Region: Example of QM/MM NMR Shieldings. J. Chem. Theory Comput. 2012;8:2260–2271. doi: 10.1021/ct300036s. [ DOI ] [ PubMed ] [ Google Scholar ] Jindal G., Warshel A.. Exploring the Dependence of QM/MM Calculations of Enzyme Catalysis on the Size of the QM Region. J. Phys. Chem. B. 2016;120:9913–9921. doi: 10.1021/acs.jpcb.6b07203. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Kulik H. J., Zhang J., Klinman J. P., Martínez T. J.. How Large Should the QM Region Be in QM/MM Calculations? The Case of Catechol O-Methyltransferase. J. Phys. Chem. B. 2016;120:11381–11394. doi: 10.1021/acs.jpcb.6b07814. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Roßbach S., Ochsenfeld C.. Influence of Coupling and Embedding Schemes on QM Size Convergence in QM/MM Approaches for the Example of a Proton Transfer in DNA. J. Chem. Theory Comput. 2017;13:1102–1107. doi: 10.1021/acs.jctc.6b00727. [ DOI ] [ PubMed ] [ Google Scholar ] Mehmood R., Kulik H. J.. Both Configuration and QM Region Size Matter: Zinc Stability in QM/MM Models of DNA Methyltransferase. J. Chem. Theory Comput. 2020;16:3121–3134. doi: 10.1021/acs.jctc.0c00153. [ DOI ] [ PubMed ] [ Google Scholar ] Csizi K.-S., Reiher M.. Universal QM/MM Approaches for General Nanoscale Applications. Wiley Interdiscip. Rev.: Comput. Mol. Sci. 2023;13:e1656. doi: 10.1002/wcms.1656. [ DOI ] [ Google Scholar ] Bash P. A., Field M. J., Davenport R. C., Petsko G. A., Ringe D., Karplus M.. Computer Simulation and Analysis of the Reaction Pathway of Triosephosphate Isomerase. Biochemistry. 1991;30:5826–5832. doi: 10.1021/bi00238a003. [ DOI ] [ PubMed ] [ Google Scholar ] Liao R.-Z., Thiel W.. Convergence in the QM-only and QM/MM Modeling of Enzymatic Reactions: A Case Study for Acetylene Hydratase. J. Comput. Chem. 2013;34:2389–2397. doi: 10.1002/jcc.23403. [ DOI ] [ PubMed ] [ Google Scholar ] Brandt F., Jacob C. R.. Systematic QM Region Construction in QM/MM Calculations Based on Uncertainty Quantification. J. Chem. Theory Comput. 2022;18:2584–2596. doi: 10.1021/acs.jctc.1c01093. [ DOI ] [ PubMed ] [ Google Scholar ] Brandt F., Jacob C. R.. Efficient Automatic Construction of Atom-Economical QM Regions with Point-Charge Variation Analysis. Phys. Chem. Chem. Phys. 2023;25:14484–14495. doi: 10.1039/D3CP01263H. [ DOI ] [ PubMed ] [ Google Scholar ] Karelina M., Kulik H. J.. Systematic Quantum Mechanical Region Determination in QM/MM Simulation. J. Chem. Theory Comput. 2017;13:563–576. doi: 10.1021/acs.jctc.6b01049. [ DOI ] [ PubMed ] [ Google Scholar ] Brunken C., Reiher M.. Automated Construction of Quantum-Classical Hybrid Models. J. Chem. Theory Comput. 2021;17:3797–3813. doi: 10.1021/acs.jctc.1c00178. [ DOI ] [ PubMed ] [ Google Scholar ] Hix M. A., Leddin E. M., Cisneros G. A.. Combining Evolutionary Conservation and Quantum Topological Analyses To Determine Quantum Mechanics Subsystems for Biomolecular Quantum Mechanics/Molecular Mechanics Simulations. J. Chem. Theory Comput. 2021;17:4524–4537. doi: 10.1021/acs.jctc.1c00313. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Brandt F., Jacob C. R.. Protein Network Centralities as Descriptor for QM Region Construction in QM/MM Simulations of Enzymes. Phys. Chem. Chem. Phys. 2023;25:20183–20188. doi: 10.1039/D3CP02713A. [ DOI ] [ PubMed ] [ Google Scholar ] Gillilan R. E., Wilson K. R.. Shadowing, Rare Events, and Rubber Bands. A Variational Verlet Algorithm for Molecular Dynamics. J. Chem. Phys. 1992;97:1757–1772. doi: 10.1063/1.463163. [ DOI ] [ Google Scholar ] Hammonds K. D., Heyes D. M.. Shadow Hamiltonian in Classical NVE Molecular Dynamics Simulations: A Path to Long Time Stability. J. Chem. Phys. 2020;152:024114. doi: 10.1063/1.5139708. [ DOI ] [ PubMed ] [ Google Scholar ] Goff J., Zhang Y., Negre C., Rohskopf A., Niklasson A. M. N.. Shadow Molecular Dynamics and Atomic Cluster Expansions for Flexible Charge Models. J. Chem. Theory Comput. 2023;19:4255–4272. doi: 10.1021/acs.jctc.3c00349. [ DOI ] [ PubMed ] [ Google Scholar ] Stanton R., Kaymak M. C., Niklasson A. M. N.. Shadow Molecular Dynamics for a Charge-Potential Equilibration Model. J. Chem. Theory Comput. 2025;21:4779–4791. doi: 10.1021/acs.jctc.5c00286. [ DOI ] [ PubMed ] [ Google Scholar ] Bedrov D., Piquemal J.-P., Borodin O., MacKerell A. D. J., Roux B., Schröder C.. Molecular Dynamics Simulations of Ionic Liquids and Electrolytes Using Polarizable Force Fields. Chem. Rev. 2019;119:7940–7995. doi: 10.1021/acs.chemrev.8b00763. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Venable R. M., Krämer A., Pastor R. W.. Molecular Dynamics Simulations of Membrane Permeability. Chem. Rev. 2019;119:5954–5997. doi: 10.1021/acs.chemrev.8b00486. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Thomas M., Brehm M., Kirchner B.. Voronoi Dipole Moments for the Simulation of Bulk Phase Vibrational Spectra. Phys. Chem. Chem. Phys. 2015;17:3207–3213. doi: 10.1039/C4CP05272B. [ DOI ] [ PubMed ] [ Google Scholar ] Thomas M., Kirchner B.. Classical Magnetic Dipole Moments for the Simulation of Vibrational Circular Dichroism by Ab Initio Molecular Dynamics. J. Phys. Chem. Lett. 2016;7:509–513. doi: 10.1021/acs.jpclett.5b02752. [ DOI ] [ PubMed ] [ Google Scholar ] Frömbgen T., Drysch K., Tassaing T., Buffeteau T., Hollóczki O., Kirchner B.. Induced Chirality and Vibrational Optical Activity in an Ionic-Liquid Anion. Angew. Chem., Int. Ed. 2025;64:e202502885. doi: 10.1002/anie.202502885. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Car R., Parrinello M.. Unified Approach for Molecular Dynamics and Density-Functional Theory. Phys. Rev. Lett. 1985;55:2471–2474. doi: 10.1103/PhysRevLett.55.2471. [ DOI ] [ PubMed ] [ Google Scholar ] Iftimie R., Minary P., Tuckerman M. E.. Ab Initio Molecular Dynamics: Concepts, Recent Developments, and Future Trends. Proc. Natl. Acad. Sci. U.S.A. 2005;102:6654–6659. doi: 10.1073/pnas.0500193102. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Wan S., Sinclair R. C., Coveney P. V.. Uncertainty Quantification in Classical Molecular Dynamics. Philos. Trans. R. Soc. A. 2021;379:20200082. doi: 10.1098/rsta.2020.0082. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Maginn E. J., Messerly R. A., Carlson D. J., Roe D. R., Elliot J. R.. Best Practices for Computing Transport Properties 1. Self-Diffusivity and Viscosity from Equilibrium Molecular Dynamics [Article v1.0] Liv. J. Comput. Mol. Sci. 2019;1:6324–6324. doi: 10.33011/livecoms.1.1.6324. [ DOI ] [ Google Scholar ] Frömbgen T., Zaby P., Alizadeh V., Da Silva J. L. F., Kirchner B., Lourenço T. C.. Lessons Learned on Obtaining Reliable Dynamic Properties for Ionic Liquids. ChemPhysChem. 2025;26:e202401048. doi: 10.1002/cphc.202401048. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Alizadeh V., Garofalo M., Urbach C., Kirchner B.. A Hybrid Monte Carlo Study of Argon Solidification. Z. Naturforsch. B. 2024;79:283–291. doi: 10.1515/znb-2023-0107. [ DOI ] [ Google Scholar ] Cuppen H. M., Karssemeijer L. J., Lamberts T.. The Kinetic Monte Carlo Method as a Way To Solve the Master Equation for Interstellar Grain Chemistry. Chem. Rev. 2013;113:8840–8871. doi: 10.1021/cr400234a. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Liu D.-J., Garcia A., Wang J., Ackerman D. M., Wang C.-J., Evans J. W.. Kinetic Monte Carlo Simulation of Statistical Mechanical Models and Coarse-Grained Mesoscale Descriptions of Catalytic Reaction-Diffusion Processes: 1D Nanoporous and 2D Surface Systems. Chem. Rev. 2015;115:5979–6050. doi: 10.1021/cr500453t. [ DOI ] [ PubMed ] [ Google Scholar ] Austin B. M., Zubarev D. Y., Lester W. A. J.. Quantum Monte Carlo and Related Approaches. Chem. Rev. 2012;112:263–288. doi: 10.1021/cr2001564. [ DOI ] [ PubMed ] [ Google Scholar ] Maerzke K. A., Yoon T. J., Jadrich R. B., Leiding J. A., Currier R. P.. First-Principles Simulations of CuCl in High-Temperature Water Vapor. J. Phys. Chem. B. 2021;125:4794–4807. doi: 10.1021/acs.jpcb.1c00083. [ DOI ] [ PubMed ] [ Google Scholar ] Chen B., Siepmann J. I.. A Novel Monte Carlo Algorithm for Simulating Strongly Associating Fluids: Applications to Water, Hydrogen Fluoride, and Acetic Acid. J. Phys. Chem. B. 2000;104:8725–8734. doi: 10.1021/jp001952u. [ DOI ] [ Google Scholar ] Rizzi F., Najm H. N., Debusschere B. J., Sargsyan K., Salloum M., Adalsteinsson H., Knio O. M.. Uncertainty Quantification in MD Simulations. Part II: Bayesian Inference of Force-Field Parameters. Multiscale Model. Simul. 2012;10:1460–1492. doi: 10.1137/110853170. [ DOI ] [ Google Scholar ] Rizzi F., Najm H. N., Debusschere B. J., Sargsyan K., Salloum M., Adalsteinsson H., Knio O. M.. Uncertainty Quantification in MD Simulations. Part I: Forward Propagation. Multiscale Model. Simul. 2012;10:1428–1459. doi: 10.1137/110853169. [ DOI ] [ Google Scholar ] Rizzi F., Jones R. E., Debusschere B. J., Knio O. M.. Uncertainty Quantification in MD Simulations of Concentration Driven Ionic Flow through a Silica Nanopore. I. Sensitivity to Physical Parameters of the Pore. J. Chem. Phys. 2013;138:194104. doi: 10.1063/1.4804666. [ DOI ] [ PubMed ] [ Google Scholar ] Rizzi F., Jones R. E., Debusschere B. J., Knio O. M.. Uncertainty Quantification in MD Simulations of Concentration Driven Ionic Flow through a Silica Nanopore. II. Uncertain Potential Parameters. J. Chem. Phys. 2013;138:194105. doi: 10.1063/1.4804669. [ DOI ] [ PubMed ] [ Google Scholar ] Lu J., Miller C., Molinero V.. Parameterization of a Coarse-Grained Model with Short-Ranged Interactions for Modeling Fuel Cell Membranes with Controlled Water Uptake. Phys. Chem. Chem. Phys. 2017;19:17698–17707. doi: 10.1039/C7CP02281F. [ DOI ] [ PubMed ] [ Google Scholar ] Lu J., Jacobson L. C., Perez Sirkin Y. A., Molinero V.. High-Resolution Coarse-Grained Model of Hydrated Anion-Exchange Membranes That Accounts for Hydrophobic and Ionic Interactions through Short-Ranged Potentials. J. Chem. Theory Comput. 2017;13:245–264. doi: 10.1021/acs.jctc.6b00874. [ DOI ] [ PubMed ] [ Google Scholar ] Peerless J. S., Kwansa A. L., Hawkins B. S., Smith R. C., Yingling Y. G.. Uncertainty Quantification and Sensitivity Analysis of Partial Charges on Macroscopic Solvent Properties in Molecular Dynamics Simulations with a Machine Learning Model. J. Chem. Inf. Model. 2021;61:1745–1761. doi: 10.1021/acs.jcim.0c01204. [ DOI ] [ PubMed ] [ Google Scholar ] Wang J., Wolf R. M., Caldwell J. W., Kollman P. A., Case D. A.. Development and Testing of a General Amber Force Field. J. Comput. Chem. 2004;25:1157–1174. doi: 10.1002/jcc.20035. [ DOI ] [ PubMed ] [ Google Scholar ] Angelikopoulos P., Papadimitriou C., Koumoutsakos P.. Bayesian Uncertainty Quantification and Propagation in Molecular Dynamics Simulations: A High Performance Computing Framework. J. Chem. Phys. 2012;137:144103. doi: 10.1063/1.4757266. [ DOI ] [ PubMed ] [ Google Scholar ] Angelikopoulos P., Papadimitriou C., Koumoutsakos P.. Data Driven, Predictive Molecular Dynamics for Nanoscale Flow Simulations under Uncertainty. J. Phys. Chem. B. 2013;117:14808–14816. doi: 10.1021/jp4084713. [ DOI ] [ PubMed ] [ Google Scholar ] Messerly R. A., Razavi S. M., Shirts M. R.. Configuration-Sampling-Based Surrogate Models for Rapid Parameterization of Non-Bonded Interactions. J. Chem. Theory Comput. 2018;14:3144–3162. doi: 10.1021/acs.jctc.8b00223. [ DOI ] [ PubMed ] [ Google Scholar ] Messerly R. A., Shirts M. R., Kazakov A. F.. Uncertainty Quantification Confirms Unreliable Extrapolation toward High Pressures for United-Atom Mie λ-6 Force Field. J. Chem. Phys. 2018;149:114109. doi: 10.1063/1.5039504. [ DOI ] [ PubMed ] [ Google Scholar ] Raabe G., Chheda V. P., Römer U.. Sequential Bayesian Force Field Calibration of Lennard-Jones Parameters with Experimental Data. Ind. Eng. Chem. Res. 2025;64:15109–15119. doi: 10.1021/acs.iecr.5c00708. [ DOI ] [ Google Scholar ] Madin O. C., Boothroyd S., Messerly R. A., Fass J., Chodera J. D., Shirts M. R.. Bayesian-Inference-Driven Model Parametrization and Model Selection for 2CLJQ Fluid Models. J. Chem. Inf. Model. 2022;62:874–889. doi: 10.1021/acs.jcim.1c00829. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Shanks B. L., Sullivan H. W., Shazed A. R., Hoepfner M. P.. Accelerated Bayesian Inference for Molecular Simulations Using Local Gaussian Process Surrogate Models. J. Chem. Theory Comput. 2024;20:3798–3808. doi: 10.1021/acs.jctc.3c01358. [ DOI ] [ PubMed ] [ Google Scholar ] Dutta R., Brotzakis Z. F., Mira A.. Bayesian Calibration of Force-Fields from Experimental Data: TIP4P Water. J. Chem. Phys. 2018;149:154110. doi: 10.1063/1.5030950. [ DOI ] [ PubMed ] [ Google Scholar ] Tran, A. V. ; Wang, Y. . A Molecular Dynamics Simulation Mechanism with Imprecise Interatomic Potentials. In Proceedings of the 3rd World Congress on Integrated Computational Materials Engineering (ICME 2015); John Wiley & Sons, 2015; Chapter 16, pp 131–138. [ Google Scholar ] Tran A. V., Wang Y.. Reliable Molecular Dynamics: Uncertainty Quantification Using Interval Analysis in Molecular Dynamics Simulation. Comput. Mater. Sci. 2017;127:141–160. doi: 10.1016/j.commatsci.2016.10.021. [ DOI ] [ Google Scholar ] Daw M. S., Baskes M. I.. Embedded-Atom Method: Derivation and Application to Impurities, Surfaces, and Other Defects in Metals. Phys. Rev. B. 1984;29:6443–6453. doi: 10.1103/PhysRevB.29.6443. [ DOI ] [ Google Scholar ] Dhaliwal G., Nair P. B., Singh C. V.. Uncertainty and Sensitivity Analysis of Mechanical and Thermal Properties Computed through Embedded Atom Method Potential. Comput. Mater. Sci. 2019;166:30–41. doi: 10.1016/j.commatsci.2019.03.060. [ DOI ] [ Google Scholar ] Vohra M., Nobakht A. Y., Shin S., Mahadevan S.. Uncertainty Quantification in Non-Equilibrium Molecular Dynamics Simulations of Thermal Transport. Int. J. Heat Mass Transfer. 2018;127:297–307. doi: 10.1016/j.ijheatmasstransfer.2018.07.073. [ DOI ] [ Google Scholar ] Stillinger F. H., Weber T. A.. Computer Simulation of Local Order in Condensed Phases of Silicon. Phys. Rev. B. 1985;31:5262–5271. doi: 10.1103/PhysRevB.31.5262. [ DOI ] [ PubMed ] [ Google Scholar ] Longbottom S., Brommer P.. Uncertainty Quantification for Classical Effective Potentials: An Extension to Potfit. Modelling Simul. Mater. Sci. Eng. 2019;27:044001. doi: 10.1088/1361-651X/ab0d75. [ DOI ] [ Google Scholar ] Kurniawan Y., Petrie C. L., Williams K. J., Transtrum M. K., Tadmor E. B., Elliott R. S., Karls D. S., Wen M.. Bayesian, Frequentist, and Information Geometric Approaches to Parametric Uncertainty Quantification of Classical Empirical Interatomic Potentials. J. Chem. Phys. 2022;156:214103. doi: 10.1063/5.0084988. [ DOI ] [ PubMed ] [ Google Scholar ] Liu S., Gerisch A., Rahimi M., Lang J., Böhm M. C., Müller-Plathe F.. Robustness of a New Molecular Dynamics-Finite Element Coupling Approach for Soft Matter Systems Analyzed by Uncertainty Quantification. J. Chem. Phys. 2015;142:104105. doi: 10.1063/1.4914020. [ DOI ] [ PubMed ] [ Google Scholar ] Mao Y., Gerisch A., Lang J., Böhm M. C., Müller-Plathe F.. Uncertainty Quantification Guided Parameter Selection in a Fully Coupled Molecular Dynamics-Finite Element Model of the Mechanical Behavior of Polymers. J. Chem. Theory Comput. 2021;17:3760–3771. doi: 10.1021/acs.jctc.0c01348. [ DOI ] [ PubMed ] [ Google Scholar ] Hall E. J., Katsoulakis M. A., Rey-Bellet L.. Uncertainty Quantification for Generalized Langevin Dynamics. J. Chem. Phys. 2016;145:224108. doi: 10.1063/1.4971433. [ DOI ] [ PubMed ] [ Google Scholar ] Fischer M., Bauer G., Gross J.. Force Fields with Fixed Bond Lengths and with Flexible Bond Lengths: Comparing Static and Dynamic Fluid Properties. J. Chem. Eng. Data. 2020;65:1583–1593. doi: 10.1021/acs.jced.9b01031. [ DOI ] [ Google Scholar ] Zhang Y., Otani A., Maginn E. J.. Reliable Viscosity Calculation from Equilibrium Molecular Dynamics Simulations: A Time Decomposition Method. J. Chem. Theory Comput. 2015;11:3537–3546. doi: 10.1021/acs.jctc.5b00351. [ DOI ] [ PubMed ] [ Google Scholar ] Desbiens N., Arnault P., Weens W., Perrin G., Dubois V.. Bootstrapping Time Correlation Functions of Molecular Dynamics. Phys. Rev. E. 2021;104:055310. doi: 10.1103/PhysRevE.104.055310. [ DOI ] [ PubMed ] [ Google Scholar ] Frömbgen, T. ; Blasius, J. ; Dick, L. ; Drysch, K. ; Alizadeh, V. ; Wylie, L. ; Kirchner, B. . Reducing Uncertainties in and Analysis of Ionic Liquid Trajectories. In Comprehensive Computational Chemistry, 1st ed.; Yáñez, M. , Boyd, R. J. , Eds.; Elsevier: Oxford, U.K., 2024; Vol. 3, pp 692–722. [ Google Scholar ] Frömbgen T., Kuhn A., Dölz J., Kirchner B.. A Multifidelity Monte Carlo Approach for Simulating the Diffusion Coefficient of Water. I. Forward Problem. J. Chem. Phys. 2026;164:024504. doi: 10.1063/5.0308381. [ DOI ] [ PubMed ] [ Google Scholar ] Fleck M., Gross J., Hansen N.. Multifidelity Gaussian Processes for Predicting Shear Viscosity over Wide Ranges of Liquid State Points Based on Molecular Dynamics Simulations. Ind. Eng. Chem. Res. 2024;63:3755–3765. doi: 10.1021/acs.iecr.3c03931. [ DOI ] [ Google Scholar ] Mostofian B., Zuckerman D. M.. Statistical Uncertainty Analysis for Small-Sample, High Log-Variance Data: Cautions for Bootstrapping and Bayesian Bootstrapping. J. Chem. Theory Comput. 2019;15:3499–3509. doi: 10.1021/acs.jctc.9b00015. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Xia K., Wei G.-W.. Molecular Nonlinear Dynamics and Protein Thermal Uncertainty Quantification. Chaos. 2014;24:013103. doi: 10.1063/1.4861202. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Russo, J. D. ; Copperman, J. ; Zuckerman, D. M. . Iterative Trajectory Reweighting for Estimation of Equilibrium and Non-Equilibrium Observables. arXiv (Physics.Computational Physics), June 16, 2020, 2006.09451, ver. 1. https://arxiv.org/abs/2006.09451 . Free Energy Calculations: Theory and Applications in Chemistry and Biology; Chipot, C. , Pohorille, A. , Castleman, A. W. , Toennies, J. P. , Yamanouchi, K. , Zinth, W. , Eds.; Springer Series in Chemical Physics, Vol. 86; Springer: Berlin, 2007. [ Google Scholar ] Kollman P.. Free Energy Calculations: Applications to Chemical and Biochemical Phenomena. Chem. Rev. 1993;93:2395–2417. doi: 10.1021/cr00023a004. [ DOI ] [ Google Scholar ] Simonson T., Archontis G., Karplus M.. Free Energy Simulations Come of Age: Protein-Ligand Recognition. Acc. Chem. Res. 2002;35:430–437. doi: 10.1021/ar010030m. [ DOI ] [ PubMed ] [ Google Scholar ] Rodinger T., Pomès R.. Enhancing the Accuracy, the Efficiency and the Scope of Free Energy Simulations. Curr. Opin. Struct. Biol. 2005;15:164–170. doi: 10.1016/j.sbi.2005.03.001. [ DOI ] [ PubMed ] [ Google Scholar ] Christ C. D., Mark A. E., van Gunsteren W. F.. Basic Ingredients of Free Energy Calculations: A Review. J. Comput. Chem. 2010;31:1569–1582. doi: 10.1002/jcc.21450. [ DOI ] [ PubMed ] [ Google Scholar ] Hansen N., van Gunsteren W. F.. Practical Aspects of Free-Energy Calculations: A Review. J. Chem. Theory Comput. 2014;10:2632–2647. doi: 10.1021/ct500161f. [ DOI ] [ PubMed ] [ Google Scholar ] Bernardi R. C., Melo M. C. R., Schulten K.. Enhanced Sampling Techniques in Molecular Dynamics Simulations of Biological Systems. Biochim. Biophys. Acta, Gen. Subj. 2015;1850:872–877. doi: 10.1016/j.bbagen.2014.10.019. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Skyner R. E., McDonagh J. L., Groom C. R., van Mourik T., Mitchell J. B. O.. A Review of Methods for the Calculation of Solution Free Energies and the Modelling of Systems in Solution. Phys. Chem. Chem. Phys. 2015;17:6174–6191. doi: 10.1039/C5CP00288E. [ DOI ] [ PubMed ] [ Google Scholar ] Mey A. S. J. S., Allen B. K., Bruce Macdonald H. E., Chodera J. D., Hahn D. F., Kuhn M., Michel J., Mobley D. L., Naden L. N., Prasad S., Rizzi A., Scheen J., Shirts M. R., Tresadern G., Xu H.. Best Practices for Alchemical Free Energy Calculations [Article v1.0] Living J. Comput. Mol. Sci. 2020;2:18378. doi: 10.33011/livecoms.2.1.18378. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Zwanzig R. W.. High-Temperature Equation of State by a Perturbation Method. I. Nonpolar Gases. J. Chem. Phys. 1954;22:1420–1426. doi: 10.1063/1.1740409. [ DOI ] [ Google Scholar ] Kirkwood J. G.. Statistical Mechanics of Fluid Mixtures. J. Chem. Phys. 1935;3:300–313. doi: 10.1063/1.1749657. [ DOI ] [ Google Scholar ] Ferrenberg A. M., Swendsen R. H.. Optimized Monte Carlo Data Analysis. Phys. Rev. Lett. 1989;63:1195–1198. doi: 10.1103/PhysRevLett.63.1195. [ DOI ] [ PubMed ] [ Google Scholar ] Kumar S., Rosenberg J. M., Bouzida D., Swendsen R. H., Kollman P. A.. THE Weighted Histogram Analysis Method for Free-Energy Calculations on Biomolecules. I. The Method. J. Comput. Chem. 1992;13:1011–1021. doi: 10.1002/jcc.540130812. [ DOI ] [ Google Scholar ] Shirts M. R., Chodera J. D.. Statistically Optimal Analysis of Samples from Multiple Equilibrium States. J. Chem. Phys. 2008;129:124105. doi: 10.1063/1.2978177. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Bhati A. P., Wan S., Wright D. W., Coveney P. V.. Rapid, Accurate, Precise, and Reliable Relative Free Energy Prediction Using Ensemble Based Thermodynamic Integration. J. Chem. Theory Comput. 2017;13:210–222. doi: 10.1021/acs.jctc.6b00979. [ DOI ] [ PubMed ] [ Google Scholar ] Bhati A. P., Wan S., Hu Y., Sherborne B., Coveney P. V.. Uncertainty Quantification in Alchemical Free Energy Methods. J. Chem. Theory Comput. 2018;14:2867–2880. doi: 10.1021/acs.jctc.7b01143. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Vassaux M., Wan S., Edeling W., Coveney P. V.. Ensembles Are Required to Handle Aleatoric and Parametric Uncertainty in Molecular Dynamics Simulation. J. Chem. Theory Comput. 2021;17:5187–5197. doi: 10.1021/acs.jctc.1c00526. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Bieniek M. K., Bhati A. P., Wan S., Coveney P. V.. TIES 20: Relative Binding Free Energy with a Flexible Superimposition Algorithm and Partial Ring Morphing. J. Chem. Theory Comput. 2021;17:1250–1265. doi: 10.1021/acs.jctc.0c01179. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Wade A. D., Bhati A. P., Wan S., Coveney P. V.. Alchemical Free Energy Estimators and Molecular Dynamics Engines: Accuracy, Precision, and Reproducibility. J. Chem. Theory Comput. 2022;18:3972–3987. doi: 10.1021/acs.jctc.2c00114. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Bhati A. P., Coveney P. V.. Large Scale Study of Ligand-Protein Relative Binding Free Energy Calculations: Actionable Predictions from Statistically Robust Protocols. J. Chem. Theory Comput. 2022;18:2687–2702. doi: 10.1021/acs.jctc.1c01288. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Edeling W., Vassaux M., Yang Y., Wan S., Guillas S., Coveney P. V.. Global Ranking of the Sensitivity of Interaction Potential Contributions within Classical Molecular Dynamics Force Fields. npj Comput. Mater. 2024;10:87. doi: 10.1038/s41524-024-01272-z. [ DOI ] [ Google Scholar ] Wang L., Friesner R. A., Berne B. J.. Replica Exchange with Solute Scaling: A More Efficient Version of Replica Exchange with Solute Tempering (REST2) J. Phys. Chem. B. 2011;115:9431–9438. doi: 10.1021/jp204407d. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Wang L., Berne B. J., Friesner R. A.. On Achieving High Accuracy and Reliability in the Calculation of Relative Protein-Ligand Binding Affinities. Proc. Natl. Acad. Sci. U. S. A. 2012;109:1937–1942. doi: 10.1073/pnas.1114017109. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Wan S., Bhati A. P., Zasada S. J., Wall I., Green D., Bamborough P., Coveney P. V.. Rapid and Reliable Binding Affinity Prediction of Bromodomain Inhibitors: A Computational Study. J. Chem. Theory Comput. 2017;13:784–795. doi: 10.1021/acs.jctc.6b00794. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Schieber N. P., Dybeck E. C., Shirts M. R.. Using Reweighting and Free Energy Surface Interpolation to Predict Solid-Solid Phase Diagrams. J. Chem. Phys. 2018;148:144104. doi: 10.1063/1.5013273. [ DOI ] [ PubMed ] [ Google Scholar ] Bauer G., Gross J.. Phase Equilibria of Solid and Fluid Phases from Molecular Dynamics Simulations with Equilibrium and Nonequilibrium Free Energy Methods. J. Chem. Theory Comput. 2019;15:3778–3792. doi: 10.1021/acs.jctc.8b01023. [ DOI ] [ PubMed ] [ Google Scholar ] Noid W. G.. Perspective: Advances, Challenges, and Insight for Predictive Coarse-Grained Models. J. Phys. Chem. B. 2023;127:4174–4207. doi: 10.1021/acs.jpcb.2c08731. [ DOI ] [ PubMed ] [ Google Scholar ] Zavadlav J., Arampatzis G., Koumoutsakos P.. Bayesian Selection for Coarse-Grained Models of Liquid Water. Sci. Rep. 2019;9:99. doi: 10.1038/s41598-018-37471-0. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Patrone P. N., Rosch T. W., Phelan F. R. Jr. Bayesian Calibration of Coarse-Grained Forces: Efficiently Addressing Transferability. J. Chem. Phys. 2016;144:154101. doi: 10.1063/1.4945380. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Sutton J. E., Guo W., Katsoulakis M. A., Vlachos D. G.. Effects of Correlated Parameters and Uncertainty in Electronic-Structure-Based Chemical Kinetic Modelling. Nat. Chem. 2016;8:331–337. doi: 10.1038/nchem.2454. [ DOI ] [ PubMed ] [ Google Scholar ] Turányi, T. ; Tomlin, A. S. . Analysis of Kinetic Reaction Mechanisms; Springer: Berlin, 2014. [ Google Scholar ] Gao C. W., Allen J. W., Green W. H., West R. H.. Reaction Mechanism Generator: Automatic Construction of Chemical Kinetic Mechanisms. Comput. Phys. Commun. 2016;203:212–225. doi: 10.1016/j.cpc.2016.02.013. [ DOI ] [ Google Scholar ] Kreitz B., Sargsyan K., Blöndal K., Mazeau E. J., West R. H., Wehinger G. D., Turek T., Goldsmith C. F.. Quantifying the Impact of Parametric Uncertainty on Automatic Mechanism Generation for CO2 Hydrogenation on Ni(111) JACS Au. 2021;1:1656–1673. doi: 10.1021/jacsau.1c00276. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Kreitz B., Wehinger G. D., Goldsmith C. F., Turek T.. Microkinetic Modeling of the Transient CO2 Methanation with DFT-Based Uncertainties in a Berty Reactor. ChemCatChem. 2022;14:e202200570. doi: 10.1002/cctc.202200570. [ DOI ] [ Google Scholar ] Bensberg M., Reiher M.. Concentration-Flux-Steered Mechanism Exploration with an Organocatalysis Application. Isr. J. Chem. 2023;63:e202200123. doi: 10.1002/ijch.202200123. [ DOI ] [ Google Scholar ] Bensberg M., Reiher M.. Uncertainty-Aware First-Principles Exploration of Chemical Reaction Networks. J. Phys. Chem. A. 2024;128:4532–4547. doi: 10.1021/acs.jpca.3c08386. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Woulfe M., Savoie B. M.. Chemical Reaction Networks from Scratch with Reaction Prediction and Kinetics-Guided Exploration. J. Chem. Theory Comput. 2025;21:1276–1291. doi: 10.1021/acs.jctc.4c01401. [ DOI ] [ PubMed ] [ Google Scholar ] Navarro Jimenez M., Le Maître O. P., Knio O. M.. Global Sensitivity Analysis in Stochastic Simulators of Uncertain Reaction Networks. J. Chem. Phys. 2016;145:244106. doi: 10.1063/1.4971797. [ DOI ] [ PubMed ] [ Google Scholar ] Ruess J., Milias-Argeitis A., Lygeros J.. Designing Experiments to Understand the Variability in Biochemical Reaction Networks. J. R. Soc. Interface. 2013;10:20130588. doi: 10.1098/rsif.2013.0588. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Komorowski M., Costa M. J., Rand D. A., Stumpf M. P. H.. Sensitivity, Robustness, and Identifiability in Stochastic Chemical Kinetics Models. Proc. Natl. Acad. Sci. U.S.A. 2011;108:8645–8650. doi: 10.1073/pnas.1015814108. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Proppe J., Kircher J.. Uncertainty Quantification of Reactivity Scales. ChemPhysChem. 2022;23:e202200061. doi: 10.1002/cphc.202200061. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Vahl M., Proppe J.. The Computational Road to Reactivity Scales. Phys. Chem. Chem. Phys. 2023;25:2717–2728. doi: 10.1039/D2CP03937K. [ DOI ] [ PubMed ] [ Google Scholar ] Eckhoff M., Diedrich J. V., Mücke M., Proppe J.. Quantitative Structure-Reactivity Relationships for Synthesis Planning: The Benzhydrylium Case. J. Phys. Chem. A. 2024;128:343–354. doi: 10.1021/acs.jpca.3c07289. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Eckhoff M., Bublitz K. L., Proppe J.. Unveiling CO2 Reactivity with Data-Driven Methods. Digit. Discovery. 2025;4:868–878. doi: 10.1039/D5DD00020C. [ DOI ] [ Google Scholar ] Beckwith M. A., Ames W., Vila F. D., Krewald V., Pantazis D. A., Mantel C., Pécaut J., Gennari M., Duboc C., Collomb M.-N., Yano J., Rehr J. J., Neese F., DeBeer S.. How Accurately Can Extended X-ray Absorption Spectra Be Predicted from First Principles? Implications for Modeling the Oxygen-Evolving Complex in Photosystem II. J. Am. Chem. Soc. 2015;137:12815–12834. doi: 10.1021/jacs.5b00783. [ DOI ] [ PubMed ] [ Google Scholar ] Chernev P., Zaharieva I., Rossini E., Galstyan A., Dau H., Knapp E.-W.. Merging Structural Information from X-ray Crystallography, Quantum Chemistry, and EXAFS Spectra: The Oxygen-Evolving Complex in PSII. J. Phys. Chem. B. 2016;120:10899–10922. doi: 10.1021/acs.jpcb.6b05800. [ DOI ] [ PubMed ] [ Google Scholar ] Askerka M., Brudvig G. W., Batista V. S.. The O2-Evolving Complex of Photosystem II: Recent Insights from Quantum Mechanics/Molecular Mechanics (QM/MM), Extended X-ray Absorption Fine Structure (EXAFS), and Femtosecond X-ray Crystallography Data. Acc. Chem. Res. 2017;50:41–48. doi: 10.1021/acs.accounts.6b00405. [ DOI ] [ PubMed ] [ Google Scholar ] Oung S. W., Rudolph J., Jacob C. R.. Uncertainty Quantification in Theoretical Spectroscopy: The Structural Sensitivity of X-ray Emission Spectra. Int. J. Quantum Chem. 2018;118:e25458. doi: 10.1002/qua.25458. [ DOI ] [ Google Scholar ] Bergmann T. G., Welzel M. O., Jacob C. R.. Towards Theoretical Spectroscopy with Error Bars: Systematic Quantification of the Structural Sensitivity of Calculated Spectra. Chem. Sci. 2020;11:1862–1877. doi: 10.1039/C9SC05103A. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Niskanen J., Vladyka A., Niemi J., Sahle C.. Emulator-Based Decomposition for Structural Sensitivity of Core-Level Spectra. R. Soc. Open Sci. 2022;9:220093. doi: 10.1098/rsos.220093. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Eronen E. A., Vladyka A., Sahle Ch. J., Niskanen J.. Structural Sensitivity of N 1s Excitations in N-Methylacetamide Solutions. J. Phys. Chem. Lett. 2025;16:1666–1672. doi: 10.1021/acs.jpclett.4c03487. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Scott A. P., Radom L.. Harmonic Vibrational Frequencies: An Evaluation of Hartree-Fock, Møller-Plesset, Quadratic Configuration Interaction, Density Functional Theory, and Semiempirical Scale Factors. J. Phys. Chem. 1996;100:16502–16513. doi: 10.1021/jp960976r. [ DOI ] [ Google Scholar ] Rauhut G., Pulay P.. Transferable Scaling Factors for Density Functional Derived Vibrational Force Fields. J. Phys. Chem. 1995;99:3093–3100. doi: 10.1021/j100010a019. [ DOI ] [ Google Scholar ] Sinha P., Boesch S. E., Gu C., Wheeler R. A., Wilson A. K.. Harmonic Vibrational Frequencies: Scaling Factors for HF, B3LYP, and MP2Methods in Combination with Correlation Consistent Basis Sets. J. Phys. Chem. A. 2004;108:9213–9217. doi: 10.1021/jp048233q. [ DOI ] [ Google Scholar ] Kashinski D. O., Chase G. M., Nelson R. G., Di Nallo O. E., Scales A. N., VanderLey D. L., Byrd E. F. C.. Harmonic Vibrational Frequencies: Approximate Global Scaling Factors for TPSS, M06, and M11 Functional Families Using Several Common Basis Sets. J. Phys. Chem. A. 2017;121:2265–2273. doi: 10.1021/acs.jpca.6b12147. [ DOI ] [ PubMed ] [ Google Scholar ] Irikura K. K., Johnson R. D., Kacker R. N.. Uncertainties in Scaling Factors for Ab Initio Vibrational Frequencies. J. Phys. Chem. A. 2005;109:8430–8437. doi: 10.1021/jp052793n. [ DOI ] [ PubMed ] [ Google Scholar ] Irikura K. K., Johnson R. D. III, Kacker R. N., Kessel R.. Uncertainties in Scaling Factors for Ab Initio Vibrational Zero-Point Energies. J. Chem. Phys. 2009;130:114102. doi: 10.1063/1.3086931. [ DOI ] [ PubMed ] [ Google Scholar ] Johnson R. D. I., Irikura K. K., Kacker R. N., Kessel R.. Scaling Factors and Uncertainties for Ab Initio Anharmonic Vibrational Frequencies. J. Chem. Theory Comput. 2010;6:2822–2828. doi: 10.1021/ct100244d. [ DOI ] [ PubMed ] [ Google Scholar ] Teixeira F., Melo A., Cordeiro M. N. D. S.. Calibration Sets and the Accuracy of Vibrational Scaling Factors: A Case Study with the X3LYP Hybrid Functional. J. Chem. Phys. 2010;133:114109. doi: 10.1063/1.3493630. [ DOI ] [ PubMed ] [ Google Scholar ] Pernot, P. ; Cailliez, F. . Semi-Empirical Correction of Ab Initio Harmonic Properties by Scaling Factors: A Validated Uncertainty Model for Calibration and Prediction. arXiv (Physics.Chemical Physics), October 27, 2010, 1010.5669, ver. 1. https://arxiv.org/abs/1010.5669 . Pernot P., Cailliez F.. Comment on “Uncertainties in Scaling Factors for Ab Initio Vibrational Zero-Point Energies” [J. Chem. Phys. 130, 114102 (2009)] and “Calibration Sets and the Accuracy of Vibrational Scaling Factors: A Case Study with the X3LYP Hybrid Functional” [J. Chem. Phys. 133, 114109 (2010)] J. Chem. Phys. 2011;134:167101. doi: 10.1063/1.3581022. [ DOI ] [ PubMed ] [ Google Scholar ] Irikura K. K., Johnson R. D. III, Kacker R. N., Kessel R.. Response to “Comment on ‘Uncertainties in Scaling Factors for Ab Initio Vibrational Zero-Point Energies’ and ‘Calibration Sets and the Accuracy of Vibrational Scaling Factors: A Case Study with the X3LYP Hybrid Functional’” [J. Chem. Phys. 134, 167101 (2011)] J. Chem. Phys. 2011;134:167102. doi: 10.1063/1.3581023. [ DOI ] [ PubMed ] [ Google Scholar ] Christiansen O.. Vibrational Structure Theory: New Vibrational Wave Function Methods for Calculation of Anharmonic Vibrational Energies and Vibrational Contributions to Molecular Properties. Phys. Chem. Chem. Phys. 2007;9:2942–2953. doi: 10.1039/b618764a. [ DOI ] [ PubMed ] [ Google Scholar ] Christiansen O.. Selected New Developments in Vibrational Structure Theory: Potential Construction and Vibrational Wave Function Calculations. Phys. Chem. Chem. Phys. 2012;14:6672–6687. doi: 10.1039/c2cp40090a. [ DOI ] [ PubMed ] [ Google Scholar ] Barone V., Biczysko M., Bloino J.. Fully Anharmonic IR and Raman Spectra of Medium-Size Molecular Systems: Accuracy and Interpretation. Phys. Chem. Chem. Phys. 2014;16:1759–1787. doi: 10.1039/C3CP53413H. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Barone V.. The Virtual Multifrequency Spectrometer: A New Paradigm for Spectroscopy. Wiley Interdiscip. Rev.: Comput. Mol. Sci. 2016;6:86–110. doi: 10.1002/wcms.1238. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Yagi K., Keçeli M., Hirata S.. Optimized Coordinates for Anharmonic Vibrational Structure Theories. J. Chem. Phys. 2012;137:204118. doi: 10.1063/1.4767776. [ DOI ] [ PubMed ] [ Google Scholar ] Panek P. T., Jacob C. R.. Efficient Calculation of Anharmonic Vibrational Spectra of Large Molecules with Localized Modes. ChemPhysChem. 2014;15:3365–3377. doi: 10.1002/cphc.201402251. [ DOI ] [ PubMed ] [ Google Scholar ] Cheng X., Steele R. P.. Efficient Anharmonic Vibrational Spectroscopy for Large Molecules Using Local-Mode Coordinates. J. Chem. Phys. 2014;141:104105. doi: 10.1063/1.4894507. [ DOI ] [ PubMed ] [ Google Scholar ] Klinting E. L., König C., Christiansen O.. Hybrid Optimized and Localized Vibrational Coordinates. J. Phys. Chem. A. 2015;119:11007–11021. doi: 10.1021/acs.jpca.5b08496. [ DOI ] [ PubMed ] [ Google Scholar ] Neff M., Rauhut G.. Toward Large Scale Vibrational Configuration Interaction Calculations. J. Chem. Phys. 2009;131:124129. doi: 10.1063/1.3243862. [ DOI ] [ PubMed ] [ Google Scholar ] Oschetzki D., Rauhut G.. Pushing the Limits in Accurate Vibrational Structure Calculations: Anharmonic Frequencies of Lithium Fluoride Clusters (LiF) n , n = 2–10. Phys. Chem. Chem. Phys. 2014;16:16426–16435. doi: 10.1039/C4CP02264E. [ DOI ] [ PubMed ] [ Google Scholar ] Mathea T., Petrenko T., Rauhut G.. Advances in Vibrational Configuration Interaction Theory - Part 2: Fast Screening of the Correlation Space. J. Comput. Chem. 2022;43:6–18. doi: 10.1002/jcc.26764. [ DOI ] [ PubMed ] [ Google Scholar ] Larsson H. R.. Benchmarking Vibrational Spectra: 5000 Accurate Eigenstates of Acetonitrile Using Tree Tensor Network States. J. Phys. Chem. Lett. 2025;16:3991–3997. doi: 10.1021/acs.jpclett.5c00782. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] König C., Christiansen O.. Automatic Determination of Important Mode-Mode Correlations in Many-Mode Vibrational Wave Functions. J. Chem. Phys. 2015;142:144115. doi: 10.1063/1.4916518. [ DOI ] [ PubMed ] [ Google Scholar ] Edwards D. E., Zubarev D. Y., Packard A., Lester W. A., Frenklach M.. Interval Prediction of Molecular Properties in Parametrized Quantum Chemistry. Phys. Rev. Lett. 2014;112:253003. doi: 10.1103/PhysRevLett.112.253003. [ DOI ] [ PubMed ] [ Google Scholar ] Avagliano D., Skreta M., Arellano-Rubach S., Aspuru-Guzik A.. DELFI: A Computer Oracle for Recommending Density Functionals for Excited States Calculations. Chem. Sci. 2024;15:4489–4503. doi: 10.1039/D3SC06440A. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Jíra T., Janoš J., Slavíček P.. Sensitivity Analysis in Photodynamics: How Does the Electronic Structure Control Cis-Stilbene Photodynamics? J. Chem. Theory Comput. 2024;20:10972–10985. doi: 10.1021/acs.jctc.4c01008. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Herbst M. F., Fransson T.. Quantifying the Error of the Core-Valence Separation Approximation. J. Chem. Phys. 2020;153:054114. doi: 10.1063/5.0013538. [ DOI ] [ PubMed ] [ Google Scholar ] Ghose A., Segal M., Meng F., Liang Z., Hybertsen M. S., Qu X., Stavitski E., Yoo S., Lu D., Carbone M. R.. Uncertainty-Aware Predictions of Molecular x-Ray Absorption Spectra Using Neural Network Ensembles. Phys. Rev. Res. 2023;5:013180. doi: 10.1103/PhysRevResearch.5.013180. [ DOI ] [ Google Scholar ] Carbone M. R., Maffettone P. M., Qu X., Yoo S., Lu D.. Accurate, Uncertainty-Aware Classification of Molecular Chemical Motifs from Multimodal X-ray Absorption Spectroscopy. J. Phys. Chem. A. 2024;128:1948–1957. doi: 10.1021/acs.jpca.3c06910. [ DOI ] [ PubMed ] [ Google Scholar ] Verma S., Aznan N. K. N., Garside K., Penfold T. J.. Uncertainty Quantification of Spectral Predictions Using Deep Neural Networks. Chem. Commun. 2023;59:7100–7103. doi: 10.1039/D3CC01988H. [ DOI ] [ PubMed ] [ Google Scholar ] Gütlich, P. ; Bill, E. ; Trautwein, A. . Mössbauer Spectroscopy and Transition Metal Chemistry: Fundamentals and Application; Springer: Berlin, 2011. [ Google Scholar ] Proppe J., Reiher M.. Reliable Estimation of Prediction Uncertainty for Physicochemical Property Models. J. Chem. Theory Comput. 2017;13:3297–3317. doi: 10.1021/acs.jctc.7b00235. [ DOI ] [ PubMed ] [ Google Scholar ] Chernick, M. R.
Bootstrap Methods: A Guide for Practitioners and Researchers, 2nd ed.; Wiley Series in Probability and Statistics; Wiley-Interscience: Hoboken, NJ, 2008. [ Google Scholar ] Riu J., Bro R.. Jack-Knife Technique for Outlier Detection and Estimation of Standard Errors in PARAFAC Models. Chemom. Intell. Lab. Syst. 2003;65:35–49. doi: 10.1016/S0169-7439(02)00090-4. [ DOI ] [ Google Scholar ] Gallenkamp C., Kramm U. I., Proppe J., Krewald V.. Calibration of Computational Mössbauer Spectroscopy to Unravel Active Sites in FeNC Catalysts for the Oxygen Reduction Reaction. Int. J. Quantum Chem. 2021;121:e26394. doi: 10.1002/qua.26394. [ DOI ] [ Google Scholar ] Tinzl M., Diedrich J. V., Mittl P. R. E., Clémancey M., Reiher M., Proppe J., Latour J.-M., Hilvert D.. Myoglobin-Catalyzed Azide Reduction Proceeds via an Anionic Metal Amide Intermediate. J. Am. Chem. Soc. 2024;146:1957–1966. doi: 10.1021/jacs.3c09279. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Weinhold F.. Quantum Cluster Equilibrium Theory of Liquids: Illustrative Application to Water. J. Chem. Phys. 1998;109:373–384. doi: 10.1063/1.476574. [ DOI ] [ Google Scholar ] Weinhold F.. Quantum Cluster Equilibrium Theory of Liquids: General Theory and Computer Implementation. J. Chem. Phys. 1998;109:367–372. doi: 10.1063/1.476573. [ DOI ] [ Google Scholar ] Kirchner B., Spickermann C., Lehmann S. B. C., Perlt E., Langner J., von Domaros M., Reuther P., Uhlig F., Kohagen M., Brüssel M.. What Can Clusters Tell Us about the Bulk?: Peacemaker: Extended Quantum Cluster Equilibrium Calculations. Comput. Phys. Commun. 2011;182:1428–1446. doi: 10.1016/j.cpc.2011.03.011. [ DOI ] [ Google Scholar ] Ludwig R., Weinhold F.. Quantum Cluster Equilibrium Theory of Liquids: Freezing of QCE/3–21G Water to Tetrakaidecahedral “Bucky-ice”. J. Chem. Phys. 1999;110:508–515. doi: 10.1063/1.478136. [ DOI ] [ Google Scholar ] LUDWIG R., WEINHOLD F., FARRAR T. C.. Quantum Cluster Equilibrium Theory of Liquids: Molecular Clusters and Thermodynamics of Liquid Ethanol. Mol. Phys. 1999;97:465–477. doi: 10.1080/00268979909482847. [ DOI ] [ Google Scholar ] Ingenmey J., Blasius J., Marchelli G., Riegel A., Kirchner B.. A Cluster Approach for Activity Coefficients: General Theory and Implementation. J. Chem. Eng. Data. 2019;64:255–261. doi: 10.1021/acs.jced.8b00779. [ DOI ] [ Google Scholar ] Matisz G., Kelterer A.-M., Fabian W. M. F., Kunsági-Máté S.. Structural Properties of Methanol-Water Binary Mixtures within the Quantum Cluster Equilibrium Model. Phys. Chem. Chem. Phys. 2015;17:8467–8479. doi: 10.1039/C4CP05836D. [ DOI ] [ PubMed ] [ Google Scholar ] Blasius J., Drysch K., Pilz F. H., Frömbgen T., Kielb P., Kirchner B.. Efficient Prediction of Mole Fraction Related Vibrational Frequency Shifts. J. Phys. Chem. Lett. 2023;14:10531–10536. doi: 10.1021/acs.jpclett.3c02761. [ DOI ] [ PubMed ] [ Google Scholar ] Blasius J., Kirchner B.. Cluster-Weighting in Bulk Phase Vibrational Circular Dichroism. J. Phys. Chem. B. 2020;124:7272–7283. doi: 10.1021/acs.jpcb.0c06313. [ DOI ] [ PubMed ] [ Google Scholar ] Kirchner B., Blasius J., Esser L., Reckien W.. Predicting Vibrational Spectroscopy for Flexible Molecules and Molecules with Non-Idle Environments. Adv. Theory Simul. 2021;4:2000223. doi: 10.1002/adts.202000223. [ DOI ] [ Google Scholar ] Khanifaev J., Schrader T., Perlt E.. The Effect of Machine Learning Predicted Anharmonic Frequencies on Thermodynamic Properties of Fluid Hydrogen Fluoride. J. Chem. Phys. 2024;160:124302. doi: 10.1063/5.0195386. [ DOI ] [ PubMed ] [ Google Scholar ] Blasius J., Ingenmey J., Perlt E., von Domaros M., Hollóczki O., Kirchner B.. Dissoziation schwacher Säuren über den gesamten Molenbruchbereich. Angew. chem. 2019;131:3245–3249. doi: 10.1002/ange.201811839. [ DOI ] [ Google Scholar ] Blasius J., Ingenmey J., Perlt E., von Domaros M., Hollóczki O., Kirchner B.. Predicting Mole-Fraction-Dependent Dissociation for Weak Acids. Angew. Chem., Int. Ed. 2019;58:3212–3216. doi: 10.1002/anie.201811839. [ DOI ] [ PubMed ] [ Google Scholar ] Taherivardanjani S., Blasius J., Brehm M., Dötzer R., Kirchner B.. Conformer Weighting and Differently Sized Cluster Weighting for Nicotine and Its Phosphorus Derivatives. J. Phys. Chem. A. 2022;126:7070–7083. doi: 10.1021/acs.jpca.2c03133. [ DOI ] [ PubMed ] [ Google Scholar ] Frömbgen T., Drysch K., Zaby P., Dölz J., Ingenmey J., Kirchner B.. Quantum Cluster Equilibrium Theory for Multicomponent Liquids. J. Chem. Theory Comput. 2024;20:1838–1846. doi: 10.1021/acs.jctc.3c00799. [ DOI ] [ PubMed ] [ Google Scholar ] Blasius J., Zaby P., Dölz J., Kirchner B.. Uncertainty Quantification of Phase Transition Quantities from Cluster Weighting Calculations. J. Chem. Phys. 2022;157:014505. doi: 10.1063/5.0093057. [ DOI ] [ PubMed ] [ Google Scholar ] Busk J., Bjørn Jørgensen P., Bhowmik A., Schmidt M. N., Winther O., Vegge T.. Calibrated Uncertainty for Molecular Property Prediction Using Ensembles of Message Passing Neural Networks. Mach. Learn.: Sci. Technol. 2022;3:015012. doi: 10.1088/2632-2153/ac3eb3. [ DOI ] [ Google Scholar ] Ramakrishnan R., Dral P. O., Rupp M., von Lilienfeld O. A.. Quantum Chemistry Structures and Properties of 134 Kilo Molecules. Sci. Data. 2014;1:140022. doi: 10.1038/sdata.2014.22. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Glavatskikh M., Leguy J., Hunault G., Cauchy T., Da Mota B.. Dataset’s Chemical Diversity Limits the Generalizability of Machine Learning Predictions. J. Cheminf. 2019;11:69. doi: 10.1186/s13321-019-0391-2. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Gruich C. J., Madhavan V., Wang Y., Goldsmith B. R.. Clarifying Trust of Materials Property Predictions Using Neural Networks with Distribution-Specific Uncertainty Quantification. Mach. Learn.: Sci. Technol. 2023;4:025019. doi: 10.1088/2632-2153/accace. [ DOI ] [ Google Scholar ] Heid E., McGill C. J., Vermeire F. H., Green W. H.. Characterizing Uncertainty in Machine Learning for Chemistry. J. Chem. Inf. Model. 2023;63:4012–4029. doi: 10.1021/acs.jcim.3c00373. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Schütt K. T., Sauceda H. E., Kindermans P.-J., Tkatchenko A., Müller K.-R.. SchNet - A Deep Learning Architecture for Molecules and Materials. J. Chem. Phys. 2018;148:241722. doi: 10.1063/1.5019779. [ DOI ] [ PubMed ] [ Google Scholar ] Hirschfeld L., Swanson K., Yang K., Barzilay R., Coley C. W.. Uncertainty Quantification Using Neural Networks for Molecular Property Prediction. J. Chem. Inf. Model. 2020;60:3770–3780. doi: 10.1021/acs.jcim.0c00502. [ DOI ] [ PubMed ] [ Google Scholar ] Janet J. P., Kulik H. J.. Predicting Electronic Structure Properties of Transition Metal Complexes with Neural Networks. Chem. Sci. 2017;8:5137–5152. doi: 10.1039/C7SC01247K. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Ryu S., Kwon Y., Kim W. Y.. A Bayesian Graph Convolutional Network for Reliable Prediction of Molecular Properties with Uncertainty Quantification. Chem. Sci. 2019;10:8438–8446. doi: 10.1039/C9SC01992H. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Janet J. P., Duan C., Yang T., Nandy A., Kulik H. J.. A Quantitative Uncertainty Metric Controls Error in Neural Network-Driven Chemical Discovery. Chem. Sci. 2019;10:7913–7922. doi: 10.1039/C9SC02298H. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Zhang Y., Menke J., He J., Nittinger E., Tyrchan C., Koch O., Zhao H.. Similarity-Based Pairing Improves Efficiency of Siamese Neural Networks for Regression Tasks and Uncertainty Quantification. J. Cheminf. 2023;15:75. doi: 10.1186/s13321-023-00744-6. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Wollschläger T., Gao N., Charpentier B., Ketata M. A., Günnemann S.. Uncertainty Estimation for Molecules: Desiderata and Methods. Proc. Mach. Learn. Res. 2023;202:37133–37156. [ Google Scholar ] Rasmussen M. H., Duan C., Kulik H. J., Jensen J. H.. Uncertain of Uncertainties? A Comparison of Uncertainty Quantification Metrics for Chemical Data Sets. J. Cheminf. 2023;15:121. doi: 10.1186/s13321-023-00790-0. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Scalia G., Grambow C. A., Pernici B., Li Y.-P., Green W. H.. Evaluating Scalable Uncertainty Estimation Methods for Deep Learning-Based Molecular Property Prediction. J. Chem. Inf. Model. 2020;60:2697–2717. doi: 10.1021/acs.jcim.9b00975. [ DOI ] [ PubMed ] [ Google Scholar ] Tynes M., Gao W., Burrill D. J., Batista E. R., Perez D., Yang P., Lubbers N.. Pairwise Difference Regression: A Machine Learning Meta-algorithm for Improved Prediction and Uncertainty Quantification in Chemical Search. J. Chem. Inf. Model. 2021;61:3846–3857. doi: 10.1021/acs.jcim.1c00670. [ DOI ] [ PubMed ] [ Google Scholar ] Yang C.-I., Li Y.-P.. Explainable Uncertainty Quantifications for Deep Learning-Based Molecular Property Prediction. J. Cheminf. 2023;15:13. doi: 10.1186/s13321-023-00682-3. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Zaspel P., Huang B., Harbrecht H., von Lilienfeld O. A.. Boosting Quantum Machine Learning Models with a Multilevel Combination Technique: Pople Diagrams Revisited. J. Chem. Theory Comput. 2019;15:1546–1559. doi: 10.1021/acs.jctc.8b00832. [ DOI ] [ PubMed ] [ Google Scholar ] Dral P. O., Owens A., Dral A., Csányi G.. Hierarchical Machine Learning of Potential Energy Surfaces. J. Chem. Phys. 2020;152:204110. doi: 10.1063/5.0006498. [ DOI ] [ PubMed ] [ Google Scholar ] Patrone P. N., Kearsley A. J., Majikes J. M., Liddle J. A.. Analysis and Uncertainty Quantification of DNA Fluorescence Melt Data: Applications of Affine Transformations. Anal. Biochem. 2020;607:113773. doi: 10.1016/j.ab.2020.113773. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Pilania G., Gubernatis J. E., Lookman T.. Multi-Fidelity Machine Learning Models for Accurate Bandgap Predictions of Solids. Comput. Mater. Sci. 2017;129:156–163. doi: 10.1016/j.commatsci.2016.12.004. [ DOI ] [ Google Scholar ] Vinod V., Maity S., Zaspel P., Kleinekathöfer U.. Multifidelity Machine Learning for Molecular Excitation Energies. J. Chem. Theory Comput. 2023;19:7658–7670. doi: 10.1021/acs.jctc.3c00882. [ DOI ] [ PubMed ] [ Google Scholar ] Articles from Chemical Reviews are provided here courtesy of American Chemical Society ACTIONS View on publisher site PDF (6.3 MB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top