ConceptioArchivearXiv CS
arXiv CSopen access

Learning Structure, Energy, and Dynamics: A Survey of Artificial Intelligence for Protein Dynamics

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Learning Structure, Energy, and Dynamics: A Survey of Artificial Intelligence for Protein Dynamics Haocheng Tang1,2,∗,‡ Liang Shi3,∗ Ya-Shi Zhang1,4,∗ Jian Tang1,5,6,† Jiarui Lu1,4,†

arXiv:2604.25244v1 [q-bio.BM] 28 Apr 2026

1

Xixian Liu1,4

Mila – Quebec AI Institute 2 University of Pittsburgh 3 Peking University 4 Université de Montréal 5 HEC Montréal 6 CIFAR AI Chair April 29, 2026

Abstract Protein dynamics underlie many biological functions, yet remain difficult to characterize due to the high computational cost of molecular dynamics simulations and the scarcity of dynamic structural data. This survey reviews recent advances in artificial intelligence for protein dynamics from three perspectives: learning from structural ensembles and trajectories, learning from physical energy signals, and learning to accelerate molecular simulations. We summarize representative methods for conformation ensemble generation, trajectory generation, Boltzmann generators, physics-aware adaptation, machine learning potentials, coarse-grained modeling, and collective variable discovery. We further discuss available datasets and key open challenges, such as scalability, thermodynamic consistency, kinetic fidelity, and integration with experimental constraints.

1

Introduction

Proteins are highly dynamic macromolecules whose biological functions—ranging from enzymatic catalysis and signal transduction to allosteric regulation—are intrinsically tied to their conformational flexibility. While determining a single, static structure provides critical insights, a comprehensive understanding of biological mechanisms demands the characterization of the full conformational ensemble and the dynamic transition pathways between different functional states. Consequently, modeling protein dynamics is a central challenge in structural biology, biophysics, and structure-based drug discovery. Traditionally, Molecular Dynamics (MD) simulation has been the workhorse for investigating protein dynamics at atomic resolution. MD integrates Newton’s equations of motion using empirical or quantum mechanical force fields to simulate the continuous evolution of molecular systems. However, a fundamental limitation of MD is its immense computational cost. Capturing high-frequency atomic vibrations necessitates femtosecond-scale integration time steps, which makes sampling long-timescale biological events—such as protein folding, large-scale conformational rearrangements, and complex ligand binding—computationally prohibitive for most realistic systems. In recent years, the landscape of structural biology has been revolutionized by Deep Learning (DL) and Generative Artificial Intelligence (GenAI). Following the unprecedented success of static protein structure prediction models like AlphaFold and ESMFold, the research frontier has rapidly expanded toward modeling protein dynamics. Generative AI models, including diffusion models, flow matching, and autoregressive models, are now being tailored to sample diverse conformational ensembles, interpolate transition paths, ∗ Equal contribution. † Correspondence to [email protected],

[email protected]

‡ Work conducted during an internship at Mila

1

and even generate complete dynamic trajectories directly. Furthermore, machine learning techniques are increasingly tightly integrated with physical principles, offering new ways to enhance, surrogate, or bypass traditional MD simulations. In this survey, we provide a comprehensive review of the rapidly evolving field of Generative AI and Machine Learning for protein dynamics modeling. We categorize the existing literature into three primary perspectives based on their underlying methodologies and training signals: • Learning from Structural Data (Section 2): We review generative models trained directly on discrete structural ensembles or continuous MD trajectories. This includes methods for sampling equilibrium conformation ensembles, generating sequence-conditional conformers, and developing autoregressive or flow-based models for continuous trajectory generation. • Learning from Energy Signals (Section 3): We discuss methods that rely on physical energy functions, rather than solely on empirical data, to learn the underlying Boltzmann distribution. Models in this category include Boltzmann Generators, which are trained via energy-based objectives or annealed importance sampling, as well as physics-aware parameter fine-tuning and guidance of pre-trained generative models. • Learning for MD Simulations (Section 4): We explore the integration of machine learning to accelerate and refine classical simulations. This encompasses Machine Learning Potentials (MLPs) for high-fidelity force fields, Machine Learning Coarse-Grained (CG) models for dimensionality reduction, and ML-derived Collective Variables (CVs) for enhanced sampling.

Ensemble Generation (§2.1 2.1)

AlphaFlow [Jing Jing et al., al. 2024a 2024a], Str2Str [Lu Lu et al., al. 2024a 2024a], ConfDiff [Wang Wang et al., al. 2024a 2024a], ESMDiff [Lu Lu et al., al. 2025a 2025a]

Trajectory Generation (§2.2 2.2)

MDGen [Jing Jing et al., al. 2024b 2024b], ConfRover [Shen Shen et al., al. 2025 2025], TEMPO [Xu Xu et al., al. 2025a 2025a], BioMD [Feng Feng et al., al. 2026a 2026a]

Boltzmann Generators (§3.1 3.1)

BG [Noé Noé et al., al. 2019 2019], TBG [Klein Klein and Noé, Noé 2024 2024], Tan et al., al. 2025a 2025a], PROSE [Tan Tan et al., al. 2025b 2025b] SBG [Tan

Physics-aware 3.2) Adaptation (§3.2

EPO [Sun Sun et al., al. 2025 2025], EBA [Lu Lu et al., al. 2025b 2025b], Metadiffusion [Lam Lam et al., al. 2026 2026]

Learning from 2) Structure (§2

AI for Protein Dynamics

Learning from Energy (§3 3)

Learning for MD Simulation (§4 4)

ML Force Fields (§4.1 4.1)

AI2 BMD [Wang Wang et al., al. 2024b 2024b], GEMS [Unke Unke et al., al. 2024 2024], Espaloma [Takaba Takaba et al., al. 2024 2024], AQuaRef [Zubatyuk Zubatyuk et al., al. 2025] 2025

Learned CG Models (§4.2 4.2)

CGnets [Wang Wang et al., al. 2019 2019], CGSchNet [Charron Charron et al., al. 2025 2025], PCCG [Ding Ding and Zhang, Zhang 2022 2022], ScoreMD [Plainer Plainer et al., al. 2025 2025]

Learned CVs (§4.3 4.3)

DeepTICA [Bonati Bonati et al., al. 2021 2021], RiD [Wang Wang et al., al. 2021 2021], DiffSim [Sipka Sipka et al., al. 2023 2023], MESA [Chen Chen and Ferguson, Ferguson 2018 2018]

Figure 1: Taxonomy of methods and representative works for each direction. Finally, we catalog publicly available datasets that are instrumental for training and benchmarking these models (Section 5), and conclude with a discussion of current challenges and promising future directions in the field (Section 6).

2

Figure 2: Examples of biomolecular conformational dynamics. This figure illustrates diverse dynamic phenomena observed via computational or experimental methods: (a) the structural ensemble of BPTI (PDB: Shaw et al., al. 2010 2010]; (b) 5PTI), represented by kinetic clusters from millisecond-level molecular dynamics [Shaw observed fold switching in MJ selecase (PDB: 4QHF/4QHH) between monomeric and tetrameric states; (c) the autoinhibited (α-helical hairpin) and the active (β-barrel) states of the RfaH transcription factor C-terminal domain (PDB: 2OUG/6C6S); (d) intrinsically disordered regions (IDR) in the HsLARP6 LaM protein (NMR ensemble, PED00247); and (e) snapshots of reversible folding reaction of the Trp-cage (PDB: 2JOF) sampled from a 100 µs trajectory [Lindorff-Larsen Lindorff-Larsen et al., al. 2011 2011].

2

Learning from Structural Data

This section reviews the generative methods for conformation ensemble generation and trajectory generation. Figure 3 provides an overview of this section.

2.1

Conformation ensemble generation

Problem Formulation. The task of protein ensemble generation involves modeling the distribution p(x) over protein structures x given a protein sequence y. In a physical context, this target distribution corresponds to the Boltzmann distribution p(x) ∝ exp(−E(x)/kB T ) determined by the energy function E(x). Generative models seek to approximate (emulate) this distribution, often leveraging datasets of crystal structures and simulation conformations. Motivation The biological function of a protein is governed not by a single static structure but by an ensemble of interconverting conformations, encompassing ordered fluctuations, coordinated large-scale motions, and transitions between distinct functional states. Accurately modeling this equilibrium distribution p(x) is therefore fundamental to understanding mechanistic biology and guiding molecular design. The gold-standard computational method, Molecular Dynamics (MD) simulation, generates physically rigorous trajectories by integrating Newton’s equations of motion. However, achieving sufficient sampling to cross high energy barriers and converge the Boltzmann distribution is often prohibitively expensive due to the femtosecond-scale time steps required, limiting its practical utility for exploring large-scale conformational landscapes. Deep learning has revolutionized single-state protein structure prediction. This success inspires a natural extension: leveraging deep models to directly approximate the distribution of accessible structures. Initial heuristic approaches, such as perturbing multiple sequence alignments (MSAs) as input to AlphaFold2 [Del Del Alamo et al., al. 2022 2022, Wayment-Steele et al., al. 2024 2024], demonstrated the potential to reveal alternative conformations but suffer from critical limitations [Jing Jing et al., al. 2024a 2024a, Wang et al., al. 2024a 2024a]: (i) they

3

operate only at inference time, preventing direct training on ensembles of known conformations; (ii) the generated conformational diversity is often limited; (iii) they are incompatible with emerging MSA-free predictors like ESMFold that leverage protein language models; and (iv) they provide no theoretical guarantee that the sampled distribution relates to a physically meaningful Boltzmann distribution. To overcome these constraints, a new paradigm employs explicit generative models to learn and sample from p(x|y). This framework enables direct training on diverse structural ensembles, offers the potential for efficient and diverse sampling, and provides a principled statistical foundation for emulating equilibrium protein dynamics. Table 1: Overview of related methods of Learning from Structural Data. Protein Rep. indicates the protein structural representation. Method

Protein Rep.

Training Data

Method

Test Systems

AlphaFlow [Jing Jing et al., al. 2024a 2024a] AlphaFolding [Cheng Cheng et al., al. 2025 2025] AlphaPPImd [Wang Wang et al., al. 2024c 2024c]

All-atom All-atom Backbone

ATLAS, PDB ATLAS, PDB Barnase-Barstar MD

Flow Matching Diffusion Latent LM

AnewSampling [Wang Wang et al., al. 2026 2026]

All-atom

AnewSampling-DB

Diffusion

Janson et al., al. 2025 2025] aSAM [Janson

All-atom

ATLAS, mdCATH

Diffusion

ATLAS, IDP, PDB ATLAS, Fast Folders Protein-protein Complexes held-out PDB systems JACS & Merck in-house MD ATLAS, mdCATH, Fast Folders

Shi et al., al. 2026 2026] ATMOS [Shi

All-atom

BioEmu [Lewis Lewis et al., al. 2025 2025]

Backbone

Feng et al., al. 2026b 2026b] BioKinema [Feng Feng et al., al. 2026a 2026a] BioMD [Feng Boltz2 [Passaro Passaro et al., al. 2025 2025] ConfDiff [Wang Wang et al., al. 2024a 2024a] Shen et al., al. 2025 2025] ConfRover [Shen DeepJump [Costa Costa et al., al. 2025 2025] Zheng et al., al. 2024 2024] DiG [Zheng DynaFold [Fan Fan et al., al. 2025a 2025a] Lu et al., al. 2024b 2024b] DynamicBind [Lu

All-atom All-atom All-atom Backbone All-atom All-atom Backbone All-atom All-atom

PDB, mdCATH, MISATO AFDB, public/in-house MD DD-13M, MISATO DD-13M, MISATO PDB, ATLAS, mdCATH PDB ATLAS mdCATH GPCRmd, PDB, PDBbind AFDB, ATLAS, PDB PDBbind

EquiJump [Costa Costa et al., al. 2024 2024]

All-atom

Fast-folding MD

Lu et al., al. 2025a 2025a] ESMDiff [Lu F3 low [Li Li et al., al. 2024 2024]

Backbone Backbone

PDB Fast-folding MD Peptides, ATLAS, public MD IDRome, PDB CG MD Peptides/Fast-folding mdCATH, Peptides MD Mad2/T4 MD ATLAS, Peptides MD ATLAS CSD, PDB 1HPV/1BRS MD small molecules MD, PDB, PDBBind, ATLAS, mdCATH, MISATO Peptides MD PDB, ATLAS ATLAS PDB ATLAS, mdCATH Peptides PDB MD, PDB

GLDP [Sengar Sengar et al., al. 2026 2026]

All-atom

Zhu et al., al. 2025 2025] IDPFold [Zhu Janson et al., al. 2023 2023] idpGAN [Janson ITO [Schreiner Schreiner et al., al. 2023 2023] Kapuśniak et al., al. 2026 2026] MARS-FM [Kapuśniak MD-LLM-1 [Murtada Murtada et al., al. 2025 2025] Jing et al., al. 2024b 2024b] MDGen [Jing P2DFlow [Jin Jin et al., al. 2025 2025] PLACER [Anishchenko Anishchenko et al., al. 2025 2025] PTraj-Diff [Xu Xu et al., al. 2025b 2025b]

Backbone Cα Cα All-atom Latent All-atom All-atom All-atom Backbone

PVB [Yu Yu et al., al. 2026 2026]

All-atom

ScooBDoob [Zhang Zhang et al., al. 2025 2025] SimpleFold [Wang Wang et al., al. 2025a 2025a] STAR-MD [Shoghi Shoghi et al., al. 2026 2026] Str2Str [Lu Lu et al., al. 2024a 2024a] TEMPO [Xu Xu et al., al. 2025a 2025a] TimeWarp [Klein Klein et al., al. 2023a 2023a] UFConf [Fan Fan et al., al. 2024 2024] UniSim [Yu Yu et al., al. 2025 2025]

CV All-atom All-atom Backbone Backbone All-atom All-atom All-atom

SSM + Diffusion Diffusion Flow Matching Flow Matching Diffusion Diffusion AR + Diffusion Flow matching Diffusion Diffusion Diffusion Stochastic Interpolants Latent LM Flow-matching Encoder-Decoder Diffusion GAN diffusion Flow Matching Latent LM Flow Matching Flow Matching Denoising Diffusion

mdCATH, MISATO Complexes, Fast Folders, IDP Complexes Complexes PDB, ATLAS, mdCATH BPTI, Fast Folders ATLAS Fast-folding MD PDB, IDP ATLAS, Fast Folders Complexes Fast-folding MD PDB, BPTI, IDP Fast-folding MD Peptides, ATLAS, public MD IDP IDP Peptides/Fast-folding Peptides, mdCATH Mad2, T4 Lysozyme ATLAS, Peptides ATLAS, Complexes Complexes, Small Mol. Complexes

Encoder-Decoder

ATLAS, mdCATH, MISATO

Bridge matching Flow Matching Diffusion Diffusion Diffusion Normalizing flow Diffusion Flow Matching

Peptides MD PDB, ATLAS ATLAS, in-house MD BPTI, Fast Folders ATLAS, CATH Peptides Complexes, RAC-47 ATLAS, Small Mol.

Generative Models for Conformational Ensembles. A growing body of work directly employs generative models to learn and sample from the equilibrium distribution of protein conformations. Early approaches such as idpGAN [Janson Janson et al., al. 2023 2023] adopted a transformer-based GAN to generate C-alpha traces, while Str2Str [Lu Lu et al., al. 2024a 2024a] framed sampling as a translation process inspired by simulated annealing, training solely on static PDB structures via denoising score matching. More recent methods leverage pow4

erful pre-trained structure predictors. AlphaFlow [Jing Jing et al., al. 2024a 2024a] fine-tuned AlphaFold/ESMFold using Lu et al., al. 2025a 2025a] adopts a latent flow matching objective for conformation ensemble generation. ESMDiff [Lu language modeling framework by finetuning ESM3 [Hayes Hayes et al., al. 2025 2025] with a discrete diffusion objective, producing efficient sequence-conditioned generative models that require no MSAs during inference. Similarly, Zhu et al., al. 2025 2025] combines a protein language model with a diffusion decoder to predict IDP ensemIDPFold [Zhu bles directly from sequence. Beyond single-state predictors, aSAM [Janson Janson et al., al. 2025 2025] employs an all-atom autoencoder with latent diffusion, and its variant aSAMt conditions generation on thermodynamic variables such as temperature. P2DFlow [Jin Jin et al., al. 2025 2025] injects a perturbed ESMFold prior into flow matching while Wang et al., al. 2025a 2025a] conditioning on energy values to sample distinct conformational states. SimpleFold [Wang takes a minimalist architectural approach, demonstrating that general-purpose transformer blocks trained via flow matching on approximately 9 million distilled structures can achieve competitive folding and ensemble prediction without domain-specific modules such as triangular updates or explicit pair representations. Together, these methods demonstrate the rapid shift from inference-time heuristics to purpose-built generative frameworks that can be trained end-to-end on ensemble data. Tackling Key Challenges: Physical Priors, Data Scarcity, and PDB Bias. Despite their promise, purely data-driven generative models often struggle to produce physically realistic ensembles. ConfDWang et al., al. 2024a 2024a] addresses this by incorporating an MD-derived energy prior as a physics-based guide iff [Wang within a force-guided diffusion process, enforcing better adherence to the Boltzmann distribution. Data scarcity is another major obstacle, especially for systems far from experimental coverage. DiG [Zheng Zheng et al., al. 2024] mitigates this via Physics-Informed Diffusion Pre-training (PIDP), which leverages the Fokker-Planck 2024 equation to supervise training directly with energy functions, removing the requirement for steady-state training data. BioEmu [Lewis Lewis et al., al. 2025 2025] further integrates experimental stability measurements to improve thermodynamic property prediction. Meanwhile, UFConf [Fan Fan et al., al. 2024 2024] confronts the conformational bias of the PDB by reweighting training samples through hierarchical structural clustering, enabling an AlphaFold2-derived diffusion model to explore unseen regions of conformational space. Extending to Protein-Protein and Protein-Ligand Dynamics. Generative dynamics modeling is also being extended beyond single-chain proteins to biomolecular interactions. AlphaPPIMD [Wang Wang et al., al. 2024c] employs a transformer-based latent generative model to sample conformations of protein-protein 2024c complexes, trained and evaluated on MD simulations of the barnase-barstar system. DynamicBind [Lu Lu et al., al. 2024b] introduces an equivariant diffusion model with a morph-like transformation that iteratively refines 2024b AlphaFold-predicted apo structures into ligand-specific holo conformations, effectively performing dynamic docking. PLACER [Anishchenko Anishchenko et al., al. 2025 2025] focuses on ligand dynamics within a protein binding site using a denoising diffusion framework. AnewSampling [Wang Wang et al., al. 2026 2026] curated over 15 million protein-ligand conformations and leveraged a quotient-space framework to sample from all-atom equilibrium distributions.

2.2

Simulation Trajectory generation

Problem Formulation. Consider a biomolecular system consisting of N atoms. We represent the dynamic evolution of this system as a trajectory of coordinates X = (x1 , x2 , . . . , xT ) ∈ RT ×N ×3 , where xt ∈ RN ×3 denotes the 3D Cartesian coordinates of all N atoms at time step t, and T represents the trajectory length, i.e., the number of frames. In addition to coordinates, the system is defined by time-invariant atomic features a, where a ∈ RN ×da represents atom-level attributes including atom type, residue type, moleculelevel identifiers that indicate whether an atom belongs to a protein or ligand, etc. Our objective is to learn a generative model pθ (X|x1 , a) that approximates the true trajectory distribution conditioned on the initial conformation x1 and context a. 2.1, which focus primarily Motivation In contrast to the conformation generation models discussed in Sec 2.1 on generating conformational ensembles that match equilibrium distributions, trajectory generation methods explicitly model the temporal evolution of protein structures. While conformation generation enables highly

5

B

A

One-Shot Generation

Sampled Conformations i.i.d. sampling

MCMC Methods

Learned Distribution Generative Modeling

Auto-Regressive Methods

Real Conformations

𝒕𝒕𝟎𝟎

𝒕𝒕𝟎𝟎 + 𝜹𝜹𝜹𝜹

𝒕𝒕𝟎𝟎 + (𝒏𝒏 − 𝟐𝟐)𝜹𝜹𝜹𝜹

Condition

𝒕𝒕𝟎𝟎 + (𝒏𝒏 − 𝟏𝟏)𝜹𝜹𝜹𝜹

Target

Figure 3: Generative modeling of protein structural dynamics from structural data. (A) Conformation ensemble generation. Generative models learn a distribution over protein structures from structural data. By sampling independent conformations from the learned distribution, the model produces multiple structural realizations that collectively form a conformational ensemble. (B) Trajectory generation. Existing approaches can be broadly categorized into frame transition models that learn long-time molecular dynamics kernels (e.g., MCMC-style transitions), autoregressive dynamics models that predict subsequent conformations sequentially, and one-shot generation models that generate entire trajectories as spatio-temporal sequences. efficient parallel sampling of i.i.d. structures, the absence of temporal dependencies precludes the estimation of kinetic observables such as transition rates and folding pathways. Trajectory generation methods address this limitation by emulating the time-ordered dynamics of molecular systems, serving as surrogate models for expensive Molecular Dynamics (MD) simulations. Markov Chain Monte Carlo Methods. A prominent family of approaches aims to accelerate MD by learning to propose large timestep transitions, effectively acting as surrogate MCMC kernels. These methods typically learn the conditional distribution p(xt+∆t |xt , a), where ∆t is orders of magnitude larger than the femtosecond timestep of MD, enabling rapid exploration of conformational space. ITO [Schreiner Schreiner et al., al. 2023] formulated this problem using conditional denoising diffusion probabilistic models to learn long-time 2023 step transitions directly from MD trajectories. TimeWarp [Klein Klein et al., al. 2023a 2023a] employs a conditional normalizing flow for the same task, crucially incorporating a Metropolis-Hastings correction step to enforce detailed balance. Subsequent works have explored more sophisticated generative backbones: EquiJump [Costa Costa et al., al. 2024] leverages a two-sided stochastic interpolant that bridges between long-interval timesteps, while Deep2024 Jump [Costa Costa et al., al. 2025 2025] adopts a conditional flow matching framework. F3 low [Li Li et al., al. 2024 2024] also utilizes conditional flow matching but implements the conditioning mechanism via classifier-free guidance, treating the previous frame as an "initial guess" combined with noise prior through weighted linear interpolation. TEMPO [Xu Xu et al., al. 2025a 2025a] formulates the transition probability through a one-step SDE that combines de-

6

terministic motion with random noise, further incorporating a two-timescale design where a low-resolution model captures slow collective motions that condition a high-resolution model for fine-grained fluctuations. Yu et al., al. 2026 2026] proposes a unified pretraining Beyond learning transitions directly in coordinate space, PVB [Yu framework that first maps initial structures to a noised latent space via an encoder-decoder, then transports them toward stage-specific targets using augmented bridge matching. This design enables consistent training across both single-structure pretraining and paired-trajectory supervision for step transition finetuning. Sengar et al., al. 2026 2026] similarly operates in a compressed latent representation of all-atom proteins, GLDP [Sengar learning dynamics in this reduced space before decoding back to atomistic trajectories; the authors benchmark three transition operators within this latent space: score-guided Langevin dynamics, a linear Koopman operator, and a standard MLP. A distinct line of MCMC methods focuses specifically on transitions between metastable states rather Zhang et al., al. 2025 2025] formulates a discrete bridge-matching framethan frame-to-frame dynamics. ScooBDoob [Zhang work that tilts Markov state model transition rates via Doob’s h-transform, generating optimal stochastic Kapuśniak et al., al. 2026 2026] similarly samples paths between prescribed initial and terminal ensembles. Mars-FM [Kapuśniak transitions between metastable states defined by an MSM, employing a flow matching framework to generate these state-to-state trajectories without learning every intermediate frame, offering a complementary approach to traditional MD emulation. One-Shot Trajectory Generation. Rather than iteratively stepping through time, an alternative paradigm Jing et al., al. 2024b 2024b], this apgenerates complete trajectories in a single forward pass. Pioneered by MDGen [Jing proach frames molecular dynamics trajectories as "molecular videos" and learns a generative model that produces the full time series of 3D structures—an approach we term one-shot trajectory generation. These models directly generate the entire sequence X = (x1 , x2 , . . . , xT ) conditioned on the initial conformation x1 and atomic context a. This paradigm enables a unified architecture capable of performing multiple tasks beyond forward simulation—such as transition path sampling—through different masking strategies applied during training and inference. Building on this foundation, BioMD [Feng Feng et al., al. 2026a 2026a] extends the video analogy to protein–ligand systems, while PTraj-Diff [Xu Xu et al., al. 2025b 2025b] adapts the framework to a diffusion arKenton and Toutanova, Toutanova 2019 2019] chitecture with tensor product attention for improved efficiency and a BERT [Kenton encoder for enhanced temporal modeling, with a focus on protein–protein complexes. DynaFold [Fan Fan et al., al. 2025a] further advances the trajectory-as-video concept by operating in a compressed latent space, achiev2025a ing greater computational and data efficiency through one-shot generation in the latent trajectory space. AlphaFolding [Cheng Cheng et al., al. 2025 2025] treats the trajectory as a 4D structure and employs a 4D diffusion model conditioned on the initial 3D conformation, directly generating time-ordered coordinate sequences. Autoregressive Methods. A complementary line of work adopts autoregressive formulations for trajectory generation, where each subsequent frame is predicted conditioned on previously generated ones. ConfRover [Shen Shen et al., al. 2025 2025] encodes historical frames into a latent representation sequence and predicts a subsequent latent state, from which the next conformation is decoded by a diffusion-based decoder conditioned on this latent. STAR-MD [Shoghi Shoghi et al., al. 2026 2026] addresses the computational bottlenecks of such architectures by introducing an efficient joint spatio-temporal attention mechanism, enabling scalable auShi et al., al. 2026 2026] adapts State Space Models to this task, capturing toregressive dynamics modeling. ATMOS [Shi long-range temporal dependencies with linear complexity in trajectory length and extending the framework to protein–ligand systems. MD-LLM-1 [Murtada Murtada et al., al. 2025 2025] takes a radically different path by leveraging large language models: it tokenizes protein structures and fine-tunes a modern LLM (Mistral 7B) via LoRA for trajectory generation. BioKinema [Feng Feng et al., al. 2026b 2026b] performs all-atom biomolecular dynamFeng et al., al. 2026a 2026a], yet also supports the ics generation using forecasting-interpolation similar to BioMD [Feng autoregressive paradigm with the Spatial-Temporal attention. UniSim [Yu Yu et al., al. 2025 2025] proposes a unified cross-domain simulator that first learns atomic representations from diverse molecular data via multi-head pretraining, then employs a stochastic interpolant framework with a force guidance module for rapid adaptation, demonstrating competitive performance across small molecules, peptides, and proteins.

7

3

Learning from Energy Signals

Table 2: Overview of Boltzmann generators and related energy-driven samplers. For each method, we report the molecular representation (Repr.), learning objective or model family (Obj.), demonstrated system scale (Scaling), and the force field or energy model used in evaluation. Abbreviations: AA = all-atom; Cart. = Cartesian; int. = internal coordinates; mixed = mixed-coordinate representation; NF = normalizing flow; FM = flow matching. Name

Repr.

Obj.

Scaling

Force Field

BG [Noé Noé et al., al. 2019 2019] Klein and Noé, Noé 2024 2024] TBG [Klein SBG [Tan Tan et al., al. 2025a 2025a] PROSE [Tan Tan et al., al. 2025b 2025b] Klein et al., al. 2023b 2023b] Eq. FM [Klein Smooth NF [Köhler Köhler et al., al. 2021 2021] FAB [Midgley Midgley et al., al. 2023 2023] Scalable BG [Kim Kim et al., al. 2024 2024] Akhound-Sadegh et al., al. 2024 2024] iDEM [Akhound-Sadegh BNEM [OuYang OuYang et al., al. 2025 2025] PITA [Akhound-Sadegh Akhound-Sadegh et al., al. 2025 2025] Schopmans and Friederich, Friederich 2025 2025] TA-BG [Schopmans EWFM (i/a) [Dern Dern et al., al. 2025 2025] von Klitzing et al., al. 2026 2026] CMT [von RegFlow/FORT [Rehman Rehman et al., al. 2026 2026] Schopmans et al., al. 2026 2026] LDR [Schopmans

AA mixed AA Cart. AA Cart. AA Cart. AA Cart. AA int. AA int. AA mixed AA Cart. AA Cart. AA Cart. AA int. AA Cart. AA int. AA Cart. AA mixed

NF FM NF NF FM NF NF NF Diffusion Diffusion Diffusion NF FM NF NF Any

BPTI (58 res.) Dipeptides Up to 10 res. Up to 8 res. Dipeptides, LJ-55 Dipeptides Dipeptides Up to 56 res. LJ-55 LJ-55 Tripeptides Hexapeptides LJ-55 Hexapeptides Tetrapeptides Hexapeptides

Amber99SB-ILDN GFN2-xTB, Amber99SB-ILDN, Amber14 Amber99SB-ILDN, Amber14 Amber14 GFN2-xTB, LJ, Amber99SB-ILDN Amber99SB-ILDN Amber96 Amber14 LJ LJ LJ, Amber14 Amber96, Amber99SB-ILDN LJ Amber99SB-ILDN, Amber96 Pre-trained CNF/OT Amber96, Amber99SB-ILDN

In this section, we provide a definition of the Boltzmann generator and present how different methods attempt to produce models that are able to sample and estimate observables of interest. To put simply, a Boltzmann generator is a conformation generative model that also has its own likelihood model p̃(x) ≈ p(x). Estimating observables of interest is then possible by importance sampling with respect to the true distribution p(x). Problem Formulation. In many protein dynamics settings, we are given an energy function E(x) (e.g., a classical force field, an implicit-solvent model, or a learned potential) that defines a target equilibrium distribution over conformations, Z   1 p(x) = exp − βE(x) , Z = exp − βE(x) dx, Z where β = (kB T )−1 . The goal is to efficiently sample x ∼ p(x) and to estimate observables Ep [f (x)] and freeenergy differences, ideally for protein systems whose energy landscapes are high-dimensional and strongly multimodal. Why learn from energy signals? Molecular dynamics (MD) and MCMC methods provide asymptotically exact sampling but often require prohibitively long trajectories to cross metastable barriers. In contrast, energy evaluations are ubiquitous and physically grounded: for any proposed conformation x, we can typically compute E(x) and often its gradient ∇x E(x) (forces). This enables a complementary paradigm to purely data-driven ensemble learning: train generative samplers using energy/force supervision, either (i) entirely without equilibrium data (simulation-free inner loops), or (ii) as a post-training alignment signal that corrects dataset bias and improves thermodynamic faithfulness. A central practical consideration is that energy and force evaluations can be expensive, so methods are often compared by the number of energy evaluations per (effectively independent) sample. Evaluation and correction with self-normalized importance sampling (SNIS). A useful separation is between proposal generation and thermodynamic correctness. When a learned sampler provides

8

(a) Structure–Energy–Model Interactions

(b) Boltzmann Generator Sampling

Figure 4: Overview of data–energy–model interactions and Boltzmann generator sampling. Top: Structural information from MD trajectories and conformational ensembles provides a structure signal for training generative models, which then sample new structures. Potential energies and force fields can provide a training or inference signal through energy-based losses or guidance, which can also be incorporated during generation via guidance sampling or self-normalized importance sampling (SNIS). Bottom: A latent variable z ∼ N (0, I) is transformed by a learned generator fθ : z 7→ x to produce molecular conformers x drawn from a tractable model distribution (pθ (x)). The model likelihood can be evaluated either with the change-ofvariables formula for a normalizing flow or, for continuous flows, by integrating the divergence of the velocity field. Generated samples are then importance reweighted according to (w̃i = exp[−βE(xi )]/pθ (xi )) and used P to estimate Boltzmann ensemble averages (Ex∼p [Obs(x)] ≈ i ŵi Obs(xi )). a tractable (or accurately estimated) proposal density p̃θ (x), samples {xi }N i=1 ∼ p̃θ can be reweighted to estimate expectations under the Boltzmann distribution without knowing the partition function Z. Define  exp − βE(xi ) w̃i w̃i = , ŵi = PN , p̃θ (xi ) j=1 w̃j so that Ep [f (x)] ≈

N X

ŵi f (xi ).

i=1

By choosing f appropriately, SNIS can estimate mean potential energy, metastable-state populations and free-energy differences, partition-function ratios, heat capacities from energy fluctuations, and free-energy 9

profiles along reaction coordinates. It also supports energy-distribution metrics: for energies ui = E(xi ) and an energy bin B, X pb(u ∈ B) = ŵi , Fb(u ∈ B) = −β −1 log pb(u ∈ B) + C, i: ui ∈B

where C is an arbitrary additive constant. In practice, weighted energy histograms or free-energy curves can be compared against reference simulations, along with discrepancies in low-order moments or 1D marginals. A key diagnostic for the reliability of this reweighting is the effective sample size PN

1 ESS = PN

2 i=1 ŵi

=

2

i=1 w̃i PN 2 i=1 w̃i

,

often reported per energy/force evaluation. Low ESS indicates poor overlap between p̃ and p, leading to high-variance estimates; in such cases, annealing- or SMC/AIS-style corrections provide alternative routes to unbiasedness, at the cost of additional energy evaluations. Scope of this section. We survey methods that explicitly leverage energy signals for equilibrium sampling and (in some cases) dynamics: (i) Boltzmann generators and related likelihood models that learn expressive proposals with (approximate or exact) likelihoods and correct residual bias via reweighting or annealing; (ii) physics-aware adaptation, which incorporates physical energy signals during post-training or at inference time. The former fine-tunes pretrained protein generators using energy-based objectives while the latter uses energies to guide sampling. Key design axes include symmetry handling, the bias–variance tradeoff of reweighting, and scalability.

3.1

Boltzmann Generators

Normalizing flows as likelihood-based generators. Normalizing flows (NFs) are generative models that construct a flexible density on configurations by transforming a simple base distribution through an invertible map. Concretely, let z ∼ N (0, I) and define z = fθ−1 (x),

x = fθ (z),

(1)

where fθ : Rd → Rd is bijective and differentiable. The model density is then available via the change-ofvariables formula, p̃θ (x) = N fθ−1 (x); 0, I



det

∂fθ−1 (x) ∂x

= N (z; 0, I) det

∂fθ (z) ∂z

−1

.

(2)

This combination of fast sampling (z 7→ x) and a tractable likelihood (log p̃θ (x)) enables Boltzmann-generatorstyle reweighting/annealing corrections, e.g. w(x) ∝ exp(−βE(x))/p̃θ (x). Architectures and likelihood cost. Discrete-time flows (e.g., coupling-layer models such as RealNVP) are built so that the Jacobian determinant is cheap to compute, often triangular by design. Continuous-time t flows (CNFs) instead define an ODE dx dt = vθ (xt , t); their likelihood requires integrating a divergence term d dt log p̃t (xt ) = −∇x · vθ (xt , t), which is inaccurate and computationally intractable in general. More recent flow-matching objectives sidestep explicit likelihood training by learning vθ to match a prescribed transport, trading exact likelihood access for improved scalability. For molecular systems, models need coordinate choices (Cartesian vs. internal) and respective symmetry handling (SE(3), E(3)), which impact sample quality and data efficiency non-trivially.

10

Classical BGs and exact-likelihood descendants. Fast direct samples should come with an explicit proposal density that supports principled correction. The original BG framework used RealNVPDinh et al., al. 2016 2016] invertible flows so that generated samples could be corrected by SNIS or annealingstyle [Dinh Noé et al., al. 2019 2019]. After the introduction of TarFlow [Zhai Zhai et al., al. 2025 2025], a transformerbased reweighting [Noé based scalable flow network, SBG [Tan Tan et al., al. 2025a 2025a] utilized it to emphasize inference-time correction, coupling the learned proposal with annealed importance sampling to improve overlap on larger peptide systems. PROSE [Tan Tan et al., al. 2025b 2025b] extends the same exact-likelihood recipe to transferable multi-system training, conditioning on atom-level and sequence-level features so that zero-shot equilibrium sampling and temperature transfer remain compatible with SNIS reweighting. Diffusion-based amortized samplers. Diffusion-based Boltzmann samplers trade tractable likelihoods for flexible, symmetry-aware generators trained directly from energetic supervision. iDEM learns a simulationfree, SE(3)-equivariant diffusion model from energies and gradients through denoising energy matching and a Akhound-Sadegh et al., al. 2024 2024]. replay-buffer outer loop that refreshes informative low-energy conformations [Akhound-Sadegh BNEM preserves the simulation-free inner-loop philosophy but parameterizes time-dependent noised energies OuYang et al., al. 2025 2025]. rather than scores, reducing target variance at the cost of an extra backward pass [OuYang PITA combines diffusion, temperature conditioning, and progressive distillation, using high-temperature MCMC data and Feynman–Kac-based annealing to bootstrap progressively lower-temperature samplers [Akhound-Sadegh Akhound-Sadegh et al., al. 2025 2025]. Energy-guided likelihood-free generative models and stable transport learning. A broad middle ground relaxes exact-likelihood training in favor of more expressive, energy-aware, and scalable transports. TBG replaces MLE-style BG training with EGNN-based, roto-permutation-equivariant flow matching in all-atom Cartesian coordinates, often operating without model likelihood due to high computational costs [Klein Klein and Noé, Noé 2024 2024]. Equivariant Flow Matching pursues a closely related E(3)-equivariant transport formulation that prioritizes geometric inductive bias and scalable flow matching over exact density evaluation [Klein Klein et al., al. 2023b 2023b]. Smooth Normalizing Flows address a complementary concern by constructing infinitely differentiable, compactly supported transformations that behave better for mixed Euclidean and periodic coordinates when stable derivatives or forces matter [Köhler Köhler et al., al. 2021 2021]. FAB shows that useful Boltzmann proposals can be learned from energies alone by coupling a normalizing flow with annealed importance sampling and optimizing a loss tied to importance-weight variance [Midgley Midgley et al., al. 2023]. TA-BG and EWFM tackle the same overlap problem through annealed, energy-aware objectives, 2023 with TA-BG iteratively cooling a high-temperature flow by importance reweighting and forward-KL training and EWFM introducing amortized and annealed energy-reweighted flow-matching losses for rougher landscapes [Schopmans Schopmans and Friederich, Friederich 2025 2025, Dern et al., al. 2025 2025]. Scalable BGs push BG-style learning to substantially larger proteins through long-context-inspired architectures and a staged transition from maximum likeKim et al., al. 2024 2024]. CMT, RegFlow/FORT, lihood to a 2-Wasserstein objective on backbone distance matrices [Kim and LDR address training pathologies from different angles by constraining intermediate transport to reduce mass teleportation, replacing fragile likelihood objectives with regression and self-consistency regularization, and adding off-policy log-dispersion with explicit energy labels to improve data efficiency on fixed datasets [von von Klitzing et al., al. 2026 2026, Rehman et al., al. 2026 2026, Schopmans et al., al. 2026 2026].

3.2

Physics-aware adaptation of generative models

Motivation. Recent protein diffusion and flow models are trained on data distributions that may not match the target thermodynamic ensemble, motivating post hoc physical alignment: starting from a pretrained generator, energies, forces, or constraints can be used to steer samples toward physically meaningful distributions, often without requiring explicit likelihoods. For a base model generating samples from p̃θ (x), we wish to sample from the tilted density padapt (x) ∝ p̃θ (x) exp(−λr(x))

11

for some guidance strength λ > 0 and reward function r(x). This reward function is often chosen to be a surrogate of the energy, validity, or other desirable thermodynamic properties. Post-training alignment. Recent work treats strong pretrained protein generators as starting points Sun et al., al. 2025 2025] and EBA [Lu Lu et al., al. 2025b 2025b] both use and then injects physical information post hoc. EPO [Sun physics-based energies to bias generators toward lower-free-energy ensembles without explicit likelihoods: EPO uses list-wise preference optimization with a tractable upper bound on trajectory likelihood ratios, whereas EBA uses a mini-batch Boltzmann-factor objective that approximately preserves energy-implied probability ratios and an energy-weighted list-wise denoising loss. Inference-time correction and steering. These methods use physical feedback to improve realism, enforce constraints, steer exploration, and extend pretrained generators towards enhanced-sampling or coarse Wang et al., al. 2024a 2024a] adds force guidance from an MD prior, Gauss–Seidel dynamical regimes. ConfDiff [Wang Projection [Chen Chen et al., al. 2026 2026] inserts an implicitly differentiated projection step that enforces steric and Lam et al., al. 2026 2026] biases pretrained genergeometric validity and can reduce denoising steps, Metadiffusion [Lam ators along chosen collective variables or experimental observables, and WT-ASBS [Nam Nam et al., al. 2026 2026] adds a sequential collective-variable bias that promotes rare-event exploration while permitting reweighting. Section summary. Across this section, the field can be read as a spectrum. At one end, exact-likelihood BGs provide the strongest route to unbiased thermodynamic estimation; in the middle, likelihood-free generative models use energy supervision to amortize sampling while relaxing likelihood requirements; at the other end, physics-aware adaptation retrofits thermodynamic information onto pretrained structural generators. The central open problem is to combine the correction guarantees of classical BGs, the symmetry-aware scalability of modern diffusion and flow models, and the transferability of foundation-model-style pretraining within a single framework.

4

Learning for MD simulations

This section provides an overview of machine learning methods for biomolecular dynamics simulation. Current methods mainly focus on basic physical formalism and neural network mapping, enabling the transformation from local atomic environments (e.g., distances rij and angles θjik ) to individual atomic energy contributions Ei and accelerating the exploration of potential energy surface (PES) of complex systems. Table 3: Overview of related methods of machine learning potentials including system scale. Name

Rep.

Training Data

Test Systems

System Size

Espaloma [Takaba Takaba et al., al. 2024 2024]

All-Atom

SPICE, peptides

Scalable

AI2 BMD [Wang Wang et al., al. 2024b 2024b] GEMS [Unke Unke et al., al. 2024 2024] AQuaRef [Zubatyuk Zubatyuk et al., al. 2025 2025] SO3LR [Kabylda Kabylda et al., al. 2025 2025] CGnets [Wang Wang et al., al. 2019 2019] CGSchNet [Charron Charron et al., al. 2025 2025] PCCG [Ding Ding and Zhang, Zhang 2022 2022] TorchMD-Net [Majewski Majewski et al., al. 2023 2023]

All-Atom All-Atom All-Atom All-Atom CG CG CG AA/CG

>10,000 atoms >25,000 atoms Full Protein >200,000 atoms Small proteins Scalable Protein domains Scalable

Flow-Matching [Köhler Köhler et al., al. 2023 2023] InvertibleCG [Chennakesavalu Chennakesavalu et al., al. 2023 2023] 2-for-1-diffusion [Arts Arts et al., al. 2023 2023] ScoreMD [Plainer Plainer et al., al. 2025 2025] GNN-LRP [Bonneau Bonneau et al., al. 2025 2025]

CG CG CG CG Analysis

Protein units GEMS Polypeptides, complexes GEMS, QM7-X, SPICE System-specific mdCATH, peptide System-specific QM9, MD17, Protein folding System-specific System-specific System-specific Fast-folding, dipeptides System-specific

Small molecules, proteins, nucleic acids Proteins Proteins Proteins Complex biosystems Specific proteins Proteins Protein domains Proteins Proteins Specific proteins Proteins Proteins GNN potential

Small to Medium Scalable Scalable Scalable N/A

12

Basic Physical Formalism & NN Mapping

Key MLP Paradigms

Advanced Hybrid ML/MM Methods

ML Collective Variables

Figure 5: Machine Learning Potentials (MLPs) and Collective Variables (CVs) for Molecular Dynamics. Top Left: Mapping of local atomic environments (rij , θjik ) through neural networks to predict total energy EM LP and forces Fi . Top Right: Core paradigms including Full MLP, ∆-learning for accuracy correction, and coarse-grained models for extended timescales. Bottom Left: Hybrid ML/MM schemes for enzymatic reactions, coupling high-precision ML active sites with MM environments. Bottom Right: ML-based dimensionality reduction (e.g., Autoencoders) to identify collective variables and reconstruct Free Energy Surfaces.

4.1

Machine Learning Potentials

Problem Formulation. The fundamental challenge in molecular modeling is to determine the potential energy surface (PES) E(R) and its resulting forces F = −∇E(R) for a system of N atoms with coordinates R ∈ R3N . In the context of protein dynamics, this PES dictates the Boltzmann distribution p(R) ∝ exp(−E(R)/kB T ). The goal of machine learning potentials (MLPs) is to find a functional mapping fθ : {Ri , Zi }N i=1 → E that approximates the high-level quantum mechanical (QM) energy while remaining efficient enough for long-time sampling and stable molecular dynamics (MD), where Zi stands for atom types. Motivation. Historically, the trade-off between accuracy and efficiency has defined the limits of molecular simulation. At one extreme, ab initio QM methods provide high fidelity by solving the Schrödinger equation but are computationally restricted to systems of a few hundred atoms. At the other, empirical force fields (FFs) simplify interactions into fixed functional forms (e.g., harmonic bonds, Lennard-Jones potentials). While FFs enable microsecond-scale simulations, their rigid parameters often fail to capture complex chemical phenomena such as bond breaking, polarization, or many-body effects. Hybrid QM/MM (Quantum Mechanics/Molecular Mechanics) methods attempt to bridge this gap by treating the active site with QM and the environment with MM; however, they still suffer from high computational cost. The emergence of MLPs offers another choice. By leveraging deep neural networks as flexible interpolators, MLPs can learn the underlying physics of QM data with near-QM accuracy while operating at speeds

13

approaching empirical force fields. Specifically, in complex environments like protein-solvent systems, MLPs can explicitly model non-bonded interactions and many-body expansions that are typically neglected or oversimplified in traditional MM. This technique enables the exploration of vast conformational landscapes with chemical accuracy, overcoming the previous limitations of both purely empirical and purely quantum approaches. Methods. The fundamental objective of MLPs is to construct an efficient mapping from the atomic configuration, defined by atomic positions {Ri } and atomic types {Zi }, to the PES. Unlike traditional force fields that rely on fixed functional forms, modern MLPs leverage deep neural networks, usually GNN and MLP, to learn the complex many-body interactions directly from QM data. Most MLPs for biomolecular systems operate on the principle of local energy decomposition, where the total energy Etotal is partitioned into local atomic contributions: N X

Etotal =

εi (Di )

(3)

i=1

where εi is the energy of atom i predicted based on its local chemical environment descriptor Di within a cutoff radius Rc . For large-scale biomolecular systems, all atom models adopt a fragmentation-based approach to balance accuracy and scalability. The total potential energy of a protein system E prot can be decomposed into the internal energies of localized units and their respective long-range non-bonded interactions: E prot = E prot_units +

n−1 X

n X

i=1

j=i+1 i∈A,j ∈A /

Coulomb Eij +

n−1 X

n X

i=1

j=i+1 i∈A,j ∈A /

V DW Eij

(4)

MLPs focus on evaluating E prot_units . Correspondingly, the force Fi acting on atom i is derived as the negative gradient of the total energy, ensuring energy-force consistency: Fiprot = Fiprot_units +

n X j=i+1 j ∈A /

FijCoulomb +

n X

FijV DW

(5)

j=i+1 j ∈A /

For complex environments such as aqueous solutions or biomembrane systems, modern MLPs go beyond local approximations to capture high-order interactions. These approaches include hybrid frameworks like Wang et al., al. 2024b 2024b], which integrates ML with polarizable force fields; ∆-learning schemes such as AI2 BMD [Wang AQuaRef [Zubatyuk Zubatyuk et al., al. 2025 2025] and SO3LR [Kabylda Kabylda et al., al. 2025 2025] that learn quantum corrections (EM L = Ebaseline + ∆E) relative to a baseline; and explicit many-body expansions like GEMS [Unke Unke et al., al. 2024 2024]. These methodologies effectively recover many-body water-solute interactions that are typically overlooked by conventional empirical force fields. Instead of using MLP only for energies or forces, recent work [Bojan Bojan et al., al. 2026 2026] treats MLP intermediate features as general-purpose embeddings of local protein environments, forming a structured manifold of local environments. Machine Learning Potentials in Exploring Molecular Dynamics. MLPs have rapidly evolved from modeling small organic molecules to tackling full-scale biological systems. Early general-purpose models like ANI-1 [Smith Smith et al., al. 2017 2017] reached the accuracy of the underlying quantum mechanical reference data, while equivariant architectures such as TorchMD-NET [Thölke Thölke and Fabritiis, Fabritiis 2022 2022] and NequIP [Batzner Batzner et al., al. 2022] later revolutionized data efficiency and physical fidelity. As the field expands to all-atom biomolecu2022 lar dynamics, diverse paradigms have emerged to handle system complexity. Espaloma [Wang Wang et al., al. 2022 2022, Takaba et al., al. 2024 2024] takes a unique approach by predicting parameters for traditional molecular mechanics force fields rather than directly regressing energy, offering a differentiable bridge to legacy systems. 14

In contrast, AI2 BMD [Wang Wang et al., al. 2024b 2024b] stands out as a pioneering and accurate direct-energy framework, utilizing a fragmentation-based ViSNet architecture to achieve quantum-level precision. Addressing Unke et al., al. 2024 2024] employs a dual “bottom-up” and “top-down” trainthe complexity of solvation, GEMS [Unke Zubatyuk et al., al. 2025 2025] specializes in ing strategy to capture long-range many-body effects, while AQuaRef [Zubatyuk the high-resolution quantum refinement of experimental structures. Finally, SO3LR [Kabylda Kabylda et al., al. 2025 2025] emphasizes physical integration and extreme scalability, enabling stable simulations of explicit solvent and complex biosystems exceeding 200,000 atoms. Together, these advancements signify a computational method shift toward versatile, high-fidelity simulations, combining the physical rigor of QM with the computational throughput essential for capturing the complex dynamics of large-scale biological systems. Recent advances have integrated MLPs into ML/MM frameworks to capture the thermodynamics of enzymatic reactions with ab initio accuracy. To mitigate the computational cost of high-level reference data, ∆-learning schemes have been widely adopted; for instance, Thodika et al. demonstrated that a ∆-MLP model learning the correction between semi-empirical and DFT Hamiltonians exhibits strong transArattu Thodika et al., al. 2025 2025]. Addressing the ferability across different DHFR mutants and environments [Arattu boundary methods and embedding interactions is another improvement: Sha et al. utilized a reweighting mechanical embedding scheme with pseudobonds, correcting for polarization via thermodynamic perturbaSha et al., al. 2025 2025], while Gradisteanu et al. developed an electrostatic machine learning embedding to tion [Sha explicitly model static and induced environmental effects [Gradisteanu Gradisteanu et al., al. 2025 2025]. Furthermore, Wang et al. implemented a robust link-atom boundary scheme combined with metadynamics to explore stereoselectivity in Diels-Alderases [Wang Wang et al., al. 2025b 2025b]. Despite these methodological strides, most current models focus on simple one-step reactions, like D-A reaction and rearrangement reaction, rather than complex multi-step catalytic cycles. Tackling Key Challenges: Long Range Interaction, Generalization and Computation Cost. Despite their promise, applying MLPs to complex biological systems faces significant hurdles, particularly regarding the locality assumption of standard GNNs. Because GNNs typically rely on message passing within a fixed cutoff radius, they struggle to capture long-range physical dependencies, such as electrostatics and dispersion, within a limited number of layers. To overcome this, recent frameworks like AI2 BMD and SO3LR explicitly incorporate classical physical terms or global attention mechanisms to recover these non-local interactions with a favorable balance between accuracy and throughput. Furthermore, achieving robust generalization across the vast chemical space remains a bottleneck; while models excel at standard amino acids, they must also accommodate heterogeneous components such as drug-like ligands, cofactors, and metal ions— a challenge currently addressed through massive, diverse training datasets (e.g., SPICE [Eastman Eastman et al., al. 2023], [Yang 2023 Yang et al., al. 2024a 2024a]) and transfer learning. Rollout stability under distribution shift is another practical concern, since small force errors can accumulate over long simulations; GGND [Hong Hong et al., al. 2026 2026] addresses this with a plug-and-play diffusion-style refinement module that improves extrapolation and stability in long MD rollouts. Finally, the computational overhead of MLPs relative to classical force fields continues to limit access to microsecond-scale dynamics, driving the need for optimized inference architectures and enhanced sampling strategies to bridge the timescale gap.

4.2

Machine Learning Coarse Grained Models

While all-atom MLPs achieve quantum-level precision by explicitly modeling every atomic interaction, they remain fundamentally constrained by the “curse of dimensionality” and the vast separation of timescales in biological processes. Even with high-performance architectures like SO3LR, simulating large-scale phenomena such as protein folding or multi-protein assembly requires sampling billions of time steps—a feat that remains computationally prohibitive when every atom is explicitly propagated. This necessitates a shift from high-fidelity force approximation to high-efficiency manifold learning. Coarse-grained (CG) models address this by strategically partitioning the system into simplified “beads”, effectively smoothing the rugged energy landscape and allowing for much larger integration time steps. By leveraging machine learning to learn the many-body potential of mean force, these models aim to retain the essential thermodynamics and kinetics

15

of the all-atom ensemble while operating at the massive spatial and temporal scales required for biological discovery. Problem Formulation. The central objective of CG modeling is to reduce the dimensionality of a biomolecular system by mapping high-resolution all-atom coordinates r to a reduced set of “beads” R via a mapping operator Ξ(r) = R. The challenge lies in learning a CG potential energy function, UCG (R), that effectively reproduces the thermodynamics of the original system. Mathematically, this corresponds to learning the many-body potential of mean force, defined as the free energy of the coarse-grained variables:  Z e−βUAA (r) δ(Ξ(r) − R)dr + C (6) UCG (R) = −kB T ln where UAA is the all-atom potential and β = 1/kB T . Souza et al., al. 2021 2021] typically rely on predefined, pairMotivation. Classical CG force fields like Martini [Souza wise functional forms (e.g., harmonic bonds, Lennard-Jones potential) that often fail to capture the complex, entropic, and multi-body effects arising from the renormalization of atomic degrees of freedom. Machine learning offers a solution by approximating the complex free energy surface without rigid functional constraints. This enables the capture of accurate folding thermodynamics and kinetics at orders of magnitude lower computational cost than all-atom MD, bridging the gap between accuracy and timescale. Machine Learning Coarse Grained Models in Protein Dynamics. Early deep learning approaches focused on direct force matching. CGnets [Wang Wang et al., al. 2019 2019] pioneered this by using fully connected networks with prior energy regularization to learn multibody interactions from all-atom forces, though they were restricted to fixed topologies. To address transferability, CGSchNet [Husic Husic et al., al. 2020 2020, Charron et al., al. 2025] and TorchMD-Net [Thölke 2025 Thölke and Fabritiis, Fabritiis 2022 2022, Majewski et al., al. 2023 2023] integrated GNNs and equivariant transformers, respectively. By learning features from local chemical environments rather than global geometries, these models allow force fields trained on small peptides to be transferred to unseen protein sequences and larger topologies. Moving beyond supervised force regression, recent methods leverage generative and variational paradigms. PCCG [Ding Ding and Zhang, Zhang 2022 2022] reformulates force field learning as a classification problem using Noise Contrastive Estimation with umbrella sampling, enabling efficient parameterization without expensive force labels. Similarly, Flow-Matching [Köhler Köhler et al., al. 2023 2023] utilizes normalizing flows to minimize relative entropy, achieving high data efficiency solely from structural distributions. The most recent wave of innovation integrates deep learning with physical constraints or interpretable interactions. 2-for-1-diffusion [Arts Arts et al., al. 2023 2023] demonstrates that a single score-based model can serve dual purposes: generating i.i.d. equilibrium samples and driving dynamics via the learned score function (∇ log p(R) ≈ −∇U (R)). Building on this, ScoreMD [Plainer Plainer et al., al. 2025 2025] addresses the physical inconsistency between the denoising distribution and the energy landscape by introducing Fokker-Planck regularization, ensuring the model functions as both a high-quality generator and a stable, conservative force field. Ripken et al., al. 2026 2026] learns Hamiltonian flow maps from instantaneous samples to predict long-time HFM [Ripken dynamics with improved stability and approximate structure preservation. Complementary to these dynamics models, InvertibleCG [Chennakesavalu Chennakesavalu et al., al. 2023 2023] optimizes a state-dependent invertible map to ensure thermodynamic consistency during back-mapping, while GNN-LRP [Bonneau Bonneau et al., al. 2025 2025] applies Layer-wise Relevance Propagation to decompose opaque GNN potentials into interpretable physical interactions.

4.3

Machine Learning Collective Variables

The identification of optimal collective variables (CVs) is a prerequisite for efficient enhanced sampling. Machine learning has revolutionized this field by automating the discovery of CVs through three primary paradigms: dimensionality reduction, reinforcement learning, and generative modeling.

16

Dimensionality Reduction and Kinetic Learning. A foundational approach involves projecting highdimensional atomic coordinates into low-dimensional latent spaces that capture significant conformational Sultan and Pande, Pande 2018 2018] utilized classification algorithms to changes. Early supervised methods like SML [Sultan discriminate between metastable states. Unsupervised approaches, particularly autoencoders, were popuChen and Ferguson, Ferguson 2018 2018, Chen et al., al. 2018 2018], which drives sampling along the nonlinear larized by MESA [Chen manifold learned from MD trajectories. Subsequent frameworks such as DESP [Salawu Salawu, 2021 2021] and FEBILAE [Belkacemi Belkacemi et al., al. 2022 2022] further refined autoencoder-based biasing to explore free energy landscapes. Beyond geometric compression, recent methods incorporate kinetic information to identify slow degrees of Bonati et al., al. 2021 2021] employs neural networks to approximate the eigenfunctions of the freedom. DeepTICA [Bonati transfer operator for variational approach, extracting slow modes directly from biased simulations like OPES. Kleiman and Shukla, Shukla 2023 2023] and TLC [Park Park et al., al. 2025 2025] leverage maximum Similarly, MaxEnt-VAMPNet [Kleiman entropy and time-lagged correlations to resolve kinetic barriers. Other notable dimensionality reduction Sasmal et al., al. 2023 2023], which applies linear discriminant analysis for CV discovery, techniques include LDA [Sasmal and Deep-TDA [Trizio Trizio and Parrinello, Parrinello 2021 2021], which trains a feed-forward neural network with a discrimination criterion on symmetry-invariant physical descriptors (e.g., interatomic distances, angles) collected from short unbiased simulations of metastable basins, projecting them into a low-dimensional space where each basin follows a preassigned distribution, thereby enabling efficient enhanced sampling with fewer CVs even Bafna et al., al. 2026 2026] aims for a lightweight alternative to full MD in multistep chemical processes. DynaProt [Bafna by reducing CV dimension to support residue-level dynamics through multivariate Gaussians at two scales: per-residue covariance matrices and scalar pairwise covariances, directly from a single static structure. Reinforcement Learning and Adaptive Sampling. To actively navigate rugged energy landscapes, RL-based frameworks treat conformational sampling as a decision-making process. REAP [Shamsi Shamsi et al., al. 2018] and its multi-agent extension MA-REAP [Kleiman 2018 Kleiman and Shukla, Shukla 2022 2022] select order parameters to maximize exploration rewards. DeepDriveMD [Lee Lee et al., al. 2019 2019] provides a scalable infrastructure for such adaptive loops. More recent strategies reformulate the exploration problem: AdaptiveBandit [Pérez Pérez et al., al. 2020 2020] treats MSM states as arms in a multi-armed bandit problem to optimize restarting points, while Adaptive CVGen [Shen Shen et al., al. 2024 2024] dynamically updates CV weights using dual RL agents to balance exploration and exploitation. Similarly, policy driven adaptive SMD [Ho Ho et al., al. 2022 2022] uses RL to steer steered molecular dynamics. Distinct from standard RL, RiD [Wang Wang et al., al. 2021 2021, Fan et al., al. 2025b 2025b] operates as a concurrent active learning scheme; it utilizes an ensemble of neural networks to approximate the potential of mean force on-the-fly, using uncertainty quantification to guide the system out of local minima without pre-existing datasets. Generative Models with CV Constraints. The latest wave of methods leverages generative priors and differentiable physics. DiffSim [Sipka Sipka et al., al. 2023 2023] and DeepLNE [Fröhlking Fröhlking et al., al. 2024 2024] utilize end-toend differentiable simulations to learn bias potentials or transition paths by optimizing path-integral loss functions. To address data scarcity, Geodesic Interpolation [Yang Yang et al., al. 2024b 2024b] employs Riemannian manifold operations to synthetically augment transition paths. Furthermore, the integration of well-pretrained Vani et al., al. 2023 2023] use models like AlphaFold (AF) and BioEmu has led to hybrid protocols: AF2-RAVE [Vani multiple AF predictions as initial seeds for Boltzmann-ranked ensembles recovering using deep information bottleneck objectives, while AlphaFold-Metainference [Brotzakis Brotzakis et al., al. 2025 2025] uses predicted distograms as restraints to model disordered regions. BioEmu-CV [Park Park et al., al. 2026 2026] learns CVs by repurposing BioEmu for time-lagged, CV-conditioned generation, encouraging the CV to focus on slow modes. Besides learning CVs, WT-ASBS [Nam Nam et al., al. 2026 2026] adds an explicit exploration bias in CV space (a repulsive term around recent samples) and uses reweighting to recover thermodynamic estimates. In parallel, diffusion-based samplers Benayad and Stirnemann, Stirnemann 2025 2025] and LAST [Tian Tian et al., al. 2023 2023] combine score-based such as DDPM-REST2 [Benayad generation with thermodynamic sampling to improve exploration of configurational space.

17

4.4

Towards Advanced Kinetic and Thermodynamic Modeling

The preceding subsections addressed how machine learning can replace or augment the core components of an MD pipeline, including the potential energy surface, the level of resolution, and the choice of reaction coordinates. Building on these foundations, machine learning is now also transforming downstream tasks that consume MD output, including the identification of transition states and the acceleration of Free Energy Perturbation (FEP) calculations. Learning for transition states. Committor functions are theoretically precise reaction coordinates related to transition pathways, but solving their high-dimensional equations has traditionally been extremely difficult. Many deep learning methods were developed to estimate committors and automatically idenArredondo et al., al. tify transition states in complex conformational changes. Architectures like GVP-GNNs [Arredondo 2025] are now being used to predict committors directly from atomic coordinates. This grants atom-level 2025 interpretability without prior topological assumptions, automatically pinpointing key atoms in the transiChen et al., al. 2023 2023] and iterative variational learntion. Other methods utilizing Siamese neural networks [Chen ing [Megías Megías et al., al. 2025 2025] can now simultaneously learn the committor and its associated transition string, Liu et al., al. 2025 2025] innovatively treats providing a comprehensive view of the reaction mechanism. TS-DAR [Liu transition states as out-of-distribution data in the potential space of the hypersphere. By introducing a regularized loss function of dispersion and variational principle, the method breaks through the previous limitation of finding transition states between two known states, and can automatically and simultaneously identify all unknown transition states of proteins from MD trajectories. Learning for FEP. FEP is the gold standard for predicting binding affinities in structure-based drug design. However, its high computational cost and complex protocol setup traditionally limit its throughput. The latest end-to-end biomolecular foundation models such as Boltz-2 [Passaro Passaro et al., al. 2025 2025] are reshaping this status quo. Boltz-2 is among the first open-source AI models to approach the accuracy of rigorous physical FEP in predicting small molecule-protein binding affinity, while being orders of magnitude faster in inference [Passaro Passaro et al., al. 2025 2025]. The Boltz-ABFE [Thaler Thaler et al., al. 2026 2026] method directly uses the high-quality 3D structure of the protein-ligand complex predicted by Boltz-2 as the initial conformation to initialize the MD simulation of absolute free energy perturbation, which completely expands the application boundary of FEP technology in the amorphous structure scenario. In addition to providing high-quality initial structures, ML is also used for automatic error correction and sampling optimization. Platforms such as FEP Ω [Giannakoulias Giannakoulias et al., al. 2025 2025] combine with downstream ML for error correction for sparse experimental labels. At the same time, the combination of active learning and FEP [van van P and Jespers, Jespers 2025 2025] can be simulated by intelligently selecting the most promising ligands, greatly improving the virtual screening efficiency in massive chemical spaces.

5

Data

The availability of high-quality data is critical for training and benchmarking generative models and machine learning potentials for protein dynamics. We summarize the key datasets in Table 4 and categorize them into three main groups: static structures, molecular dynamics trajectories, and disordered proteins with experimental measurements. Berman et al., al. 2000 2000] serves as the foundational Static Structures. The Protein Data Bank (PDB) [Berman repository for experimentally determined 3D structures of proteins and complex assemblies. Moving beyond experimental data, massive databases of predicted structures have emerged via deep learning distillation. AlphaFold Protein Structure Database (AFDB) [Jumper Jumper et al., al. 2021 2021, Varadi et al., al. 2024 2024] contains over 200 million predicted structures, capturing the structural coverage of almost all cataloged proteins in UniProt [Bateman Bateman et al., al. 2024 2024]. Building on this, the AFESM dataset [Yeo Yeo et al., al. 2025 2025] expands the coverage to metagenomic sequences, providing hundreds of millions of predicted structures for previously unseen 18

Table 4: Overview of public protein datasets: source, scale, and system types. Note that some databases such as PDB are still being updated. Dataset

Source

Scale

PDB [Berman Berman et al., al. 2000 2000] AFDB [Varadi Varadi et al., al. 2024 2024] Yeo et al., al. 2025 2025] AFESM [Yeo CDDB [Reidenbach Reidenbach et al., al. 2025 2025] Teddymer [Didi Didi et al., al. 2026 2026]

Experimental Distillation Distillation Distillation Distillation

ATLAS [Vander Vander Meersche et al., al. 2024 2024] DynamicPDB [Liu Liu et al., al. 2024 2024] DynaRepo [Mokhtari Mokhtari et al., al. 2025 2025] Mirarchi et al., al. 2024 2024] mdCATH [Mirarchi Fast-Folders [Lindorff-Larsen Lindorff-Larsen et al., al. 2011 2011] Shaw2010 [Shaw Shaw et al., al. 2010 2010] Zhou et al., al. 2025 2025] ProteinConformers [Zhou MDRepo [Roy Roy et al., al. 2024 2024] GPCRmd [Rodríguez-Espigares Rodríguez-Espigares et al., al. 2020 2020] Stansfeld et al., al. 2015 2015] MemProtMD [Stansfeld AF-CALVADOS [von von Bülow et al., al. 2025 2025] MISATO [Siebenmorgen Siebenmorgen et al., al. 2024 2024] Li et al., al. 2025 2025] DD-13M [Li ManyPeptidesMD [Tan Tan et al., al. 2025b 2025b]

Simulation (MD) Simulation (MD) Simulation (MD) Simulation (MD) Simulation (MD) Simulation (MD) Simulation (MD) Simulation (MD) Simulation (MD) Simulation (CGMD) Simulation (CGMD) Simulation (MD) Simulation (MD) Simulation (MD)

Systems

Static Structures ∼215K structures 214M structures 821M structures 455k structures 510k clusters

Monomer, Complex Monomer Monomer (Metagenomic) Monomer Dimer

MD Simulation Trajectories 1,390 proteins (100 ns) 12,643 proteins (1 µs) ∼720 systems (500 ns) 5,398 domains (up to 500 ns) 12 domains (up to ms scale) 2 proteins (ms scale) 381K conformations; 87 targets ∼4.2K community trajectories 570 systems (3 × 500 ns) 2,294 proteins (1 µs) 12,483 human proteins ∼17K structures (8 ns) 26,612 trajectories 21,700 peptides (200 ns)

Protein Monomer Protein Monomer Monomer, Complex Monomer (Domains) Mini-protein Mini-protein Protein Monomer Monomer, Complex Membrane (GPCR) Membrane Proteins Monomer, MDPs Protein-Ligand Protein-Ligand Peptides (2-8 AA)

Disordered Proteins & Experimental Measurements BMRB [Hoch Hoch et al., al. 2023 2023] SASBDB [Kikhney Kikhney et al., al. 2020 2020] Piovesan et al., al. 2025 2025] MobiDB [Piovesan Nugnes et al., al. 2025 2025] DisProt [Nugnes

Experimental (NMR) Experimental (SAS) Hybrid Hybrid

>11M chemical shifts >1K entries >200M entries 3,201 proteins

Proteins, Nucleic acids Monomer, Complex IDPs, IDRs IDPs, IDRs

proteins. Additionally, large-scale entirely synthetic, fully atomistic de novo predicted sequences and structures have also been recently adopted to unlock vast structural diversity [Reidenbach Reidenbach et al., al. 2025 2025]. Teddymer also presents more than 500k dimer clusters as a database generated by treating each domain as a chain in AFDB [Didi Didi et al., al. 2026 2026]. Opportunities and Challenges: In vitro crystal structures (PDB) provide ground-truth atomic representations governed by true physical constraints, but they suffer from severe sparsity and known experimental biases (such as crystal packing artifacts). Conversely, distillation and synthetic approaches offer an unprecedented scale and broad sequence-structure coverage, unlocking vast amounts of data to pre-train large-scale generative models. However, these synthetic structures predominantly reflect the single dominant, static state favored by the underlying prediction algorithm, completely missing the ruggedness of genuine thermodynamic ensembles. Furthermore, synthetic data often harbors predictive artifacts or unphysical stereochemistry. Consequently, applying static structures to learn protein dynamics necessitates rigorous filtering, e.g., culling low-confidence predictions via pLDDT scores, ensuring sufficient secondary structures, or enforce consistency [Reidenbach Reidenbach et al., al. 2025 2025], so that models learn physically meaningful relationships rather than overfitting to the unphysical biases inherent in modern structural predictors. MD Simulation Trajectories. While static structures offer single snapshots, characterizing protein dynamics requires continuous trajectories. Existing trajectory datasets can be broadly categorized by their tarVander Meersche et al., al. get systems and timescales. For globular monomeric proteins, datasets like ATLAS [Vander 2024] and DynamicPDB [Liu 2024 Liu et al., al. 2024 2024] provide extensive 100 ns to 1 µs MD simulations for thousands of targets, while the mdCATH dataset [Mirarchi Mirarchi et al., al. 2024 2024] features simulation data at various temperatures for domains. To study long-timescale folding events and conformational diversity, the fastfolding proteins dataset [Lindorff-Larsen Lindorff-Larsen et al., al. 2011 2011], the pioneering millisecond-scale trajectories of BPTI (Shaw2010) [Shaw Shaw et al., al. 2010 2010], and ProteinConformers [Zhou Zhou et al., al. 2025 2025] remain critical benchmarks. Beyond standard monomeric proteins, several specialized datasets focus on complex interactions and en-

19

vironments. For protein-ligand interactions, MISATO [Siebenmorgen Siebenmorgen et al., al. 2024 2024] and DD-13M [Li Li et al., al. 2025] supply valuable QM/MM and unbinding MD trajectories. Membrane proteins and GPCRs are ex2025 tensively covered by MemProtMD [Stansfeld Stansfeld et al., al. 2015 2015] and GPCRmd [Rodríguez-Espigares Rodríguez-Espigares et al., al. 2020 2020], Tan et al., al. 2025b 2025b] respectively. Finally, for bottom-up and scalable ML models, ManyPeptidesMD [Tan provides dense sampling of small peptide dynamics, and AF-CALVADOS [von von Bülow et al., al. 2025 2025] provides coarse-grained MD simulations for large human proteins. Furthermore, community-curated platforms like DynaRepo [Mokhtari Mokhtari et al., al. 2025 2025], MDRepo [Roy Roy et al., al. 2024 2024], and MDDB [Amaro Amaro et al., al. 2025 2025] aggregate and standardize massive simulation records for machine learning. Opportunities and Challenges: MD trajectories capture the time-correlated kinetic behaviors and thermodynamic distribution that static data lacks. Learning from this data gives models the ability to recover transition probabilities, reactive pathways, and sequence-dependent flexibility. However, generating large-scale, diverse, and heavily sampled MD datasets is remarkably expensive. Crucially, the accuracy and reliability of these datasets are fundamentally tied to the quality of the underlying empirical force fields used during simulation, directly impacting the validity of the underlying dynamics landscape explored. Datasets curated from classical simulations inherently carry these modeling inaccuracies. Furthermore, short simulations can easily become trapped in local energy minima, leading to highly auto-correlated distributions that fail to properly represent the global equilibrium. When generative methods rely solely on this data, they are vulnerable to overfitting localized phase spaces and perfectly reproducing the imperfections of the molecular mechanism force fields, highlighting the need for higher-fidelity quantum data or enhanced sampling strategies. Intrinsically Disordered Proteins and Experimental Measurements. Intrinsically Disordered Proteins (IDPs) and Regions (IDRs) present unique challenges due to their highly dynamic, ensemble-like Piovesan et al., al. 2023 2023, 2025 2025] and DisProt [Nugnes Nugnes et al., al. 2025 2025] curate nature. Databases such as MobiDB [Piovesan experimental annotations and predictions of disorder, serving as critical references for models predicting IDP ensembles. Additionally, experimental measurements of dynamics and flexibility, such as NMR chemical shifts from BMRB [Hoch Hoch et al., al. 2023 2023] and small-angle scattering profiles from SASBDB [Kikhney Kikhney et al., al. 2020], provide macroscopic or time-averaged thermodynamic constraints that can steer or validate generative 2020 ensemble models. Opportunities and Challenges: Such ensemble-averaged measurements offer the unique advantage of reflecting flexible macro-states observed under physiological conditions, rather than static artificial snapshots or computationally biased simulations. For AI models, substituting or supplementing force field learning with direct observational signals can correct severe kinetic inaccuracies and guide the sampling towards experimentally validated free-energy basins. This is particularly vital for modeling IDPs, where no true "gold-standard" classical force field yet exists to reliably simulate their highly fluxional ensembles. This distinct lack of a universally accurate physical model provides a profound opportunity for data-driven generative models to learn directly from structural ensembles and experimental macro-signals, bypassing empirical force field bottlenecks altogether. However, the central computational challenge remains: experimental constraints are intrinsically sparse, noisy, and indirect. NMR or SAXS profiles do not provide explicit 3D atomic coordinates. Instead, they entail a mathematically complex, one-to-many relationship where numerous distinct structural ensembles could theoretically reproduce the same average signal. Bridging generative models with experimental constraints thus requires robust, differentiable forward models capable of computing expectation values back to atomic representation, a task complicated by substantial computational overhead and empirical approximations within the forward models themselves.

6

Discussion

The intersection of generative AI and molecular dynamics has catalyzed a paradigm shift in how we model and understand protein dynamics. Throughout this survey, we have explored the rapid evolution from traditional physics-based simulations to sophisticated machine learning frameworks. As detailed in the preceding sections, this progress stems from three primary pillars: learning from vast structural datasets to sample 20

conformational ensembles and simulate trajectories (Section 2), leveraging physical energy signals to steer generative priors toward thermodynamic fidelity (Section 3), and integrating machine learning deeply into MD pipelines via neural potentials, coarse-graining, and learned collective variables (Section 4). Supported by the growing availability of both simulated and experimental datasets (Section 5), these methodologies effectively bridge the critical gap between static atomic snapshots and true, time-aware biomolecular motions. As these foundation models mature, their significance extends far beyond computational biophysics into transformative real-world applications across life sciences. In structure-based drug discovery, characterizing the full conformational ensemble enables the reliable identification of cryptic pockets and elusive allosteric sites that are completely invisible in static crystal structures. Furthermore, generative conformational sampling combined with ML-based potentials may enable faster binding free-energy estimation and improve the practical utility of virtual screening, although achieving reliable chemical accuracy remains system- and protocol-dependent. Anishchenko et al., al. 2025 2025], are already successfully capTargeted generative models, such as PLACER [Anishchenko turing the dynamic behavior of ligands within protein binding sites. In the domain of protein and enzyme engineering, the ability to accurately trace transition states and reaction coordinates empowers the design of highly efficient de novo biocatalysts and dynamic nanosensors, moving the field beyond the engineering of rigid scaffolds toward the programmable design of biological function. Despite these remarkable achievements, the full potential of AI in modeling protein dynamics is hindered by several critical bottlenecks. To overcome these limitations, the research community must actively address the following interconnected future directions: • Generative Force Fields: Traditional Machine Learning Potentials (MLPs, Section 4) act as neural distillations of Density Functional Theory (DFT) or classical Molecular Mechanics Force Fields (MMFFs). While highly efficient, this supervised approach is inherently bottlenecked by the biases, inaccuracies, and simplifications of the underlying empirical or quantum labels. A major future opportunity lies in developing Generative Force Fields: models that establish the energy landscape directly from the data distribution of structural ensembles without relying on potentially biased energy labels. Ensuring that these generative priors naturally entail conservative force fields while adhering to detailed balance and thermodynamic consistency remains a largely unsolved algorithmic challenge, but presents a pathway to bypass legacy label biases entirely. A key challenge is that such data-derived landscapes may not be uniquely identifiable without physical constraints or experimental calibration. • Integration of Sparse Experimental Constraints: As discussed in Section 5, most current ML models still rely heavily on computational MD trajectories or algorithmically distilled synthetic structures, making them vulnerable to empirical modeling biases or unphysical prediction artifacts. A major future direction involves directly incorporating heterogeneous, sparse experimental data—such as NMR chemical shifts, SAXS profiles, and cryo-EM equilibrium densities—as foundational training signals or for rigorous post-training calibration. Using these highly accurate experimental signals to calibrate generated ensembles will critically correct localized kinetic inaccuracies, grounding generative sampling directly in true physiological states in vivo. Bridging this gap will require robust differentiable forward models capable of computing expectation values from atomic coordinates back into macroscopic observables. • Biomolecular Foundation Models: The development of unified computational architectures marks a critical frontier. By coherently integrating diverse modalities of data, encompassing organic molecules, protein and nucleotide sequences, evolutionary profiles, static structural snapshots, and continuous simulation trajectories, future foundation models (e.g., Boltz-2 [Passaro Passaro et al., al. 2025 2025]) hold the potential for robust zero-shot generalization across disparate dynamic tasks, mirroring the profound impact of foundation models in natural language processing. • Scaling Boltzmann Generators: As highlighted in Section 3, applying exact Boltzmann generators to large protein complexes remains exceptionally challenging due to the curse of dimensionality and the

21

high computational cost of the likelihood evaluations required for unbiased reweighting. These difficulties become even more pronounced for protein–ligand systems and other heterogeneous biomolecular assemblies, where binding-mode diversity, interfacial rearrangements, and induced-fit or allosteric couplings substantially enlarge the relevant conformational landscape. Future efforts must therefore focus on architectural innovations that scale these systems (e.g., via efficiently parallelized neural architectures or diffusion-based amortized samplers) while maintaining tractable accept-reject rates during importance sampling. • Trajectory-based Benchmarking: For methods generating continuous time-series trajectories, such as autoregressive steps or Markovian transitions, error accumulation over extended horizons severely restricts models from simulating microsecond phenomena without unphysical structural branching. This methodological hurdle underscores a conspicuous lack of standardized, comprehensive community benchmarks. Establishing rigorous evaluation frameworks to measure the temporal fidelity, kinetic consistency, and long-term stability of generative trajectories is an urgent necessity for standardizing progress in the field. • Uncertainty Quantification: If a protein dynamics model can produce calibrated uncertainty estimates alongside its predicted updates, these signals could indicate when a rollout is leaving poorly explored regions of conformational space and motivate uncertainty-aware adaptive resampling; more generally, limiting rollout length based on model confidence is a standard way to reduce compounding Imbalzano et al., al. 2021 2021]. While AlphaFold’s pLDDT excels at static, pererror in sequential prediction [Imbalzano residue scores [Jumper Jumper et al., al. 2021 2021], we are interested in a calibrated uncertainty over rollout dynamics. Closing Remarks. In summary, generative models has shown great potential in decoupling the sampling of biological phase space from the extreme computational cost of classical Newtonian integration. We are entering a new era of structural biology where the ability to rapidly and accurately simulate protein dynamics will soon rival our current capacity to accurately predict static structures. By increasingly unifying data-driven generative architectures with fundamental physical laws and rigorous experimental calibration, the research community will witness many advances toward an ultimate and transformative goal of the comprehensive, predictive understanding of biomolecular function in motion.

Author Contribution H.T., L.S., and Y.Z. mainly conducted the literature review and wrote Chapter 2, Chapter 3, and Chapter 4, respectively. X.L. contributed to the revision of the manuscript and wrote part of the Chapter 4. J.T. provided overall supervision, reviewed and revised the manuscript. J.L. conceived the structure of the survey, directed the project, conducted the literature review, and wrote the introduction, data & discussion sections. All authors read and approved the final manuscript.

References Bowen Jing, Bonnie Berger, and Tommi Jaakkola. AlphaFold meets flow matching for generating protein ensembles. In Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp, editors, Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 22277–22303. PMLR, 21–27 Jul 2024a. URL https://proceedings.mlr.press/v235/jing24a.html https://proceedings.mlr.press/v235/jing24a.html. Jiarui Lu, Bozitao Zhong, Zuobai Zhang, and Jian Tang. Str2str: A score-based framework for zero-shot protein conformation sampling. In The Twelfth International Conference on Learning Representations, 2024a. URL https://openreview.net/forum?id=C4BikKsgmK https://openreview.net/forum?id=C4BikKsgmK. Yan Wang, Lihao Wang, Yuning Shen, Yiqun Wang, Huizhuo Yuan, Yue Wu, and Quanquan Gu. Protein conformation generation via force-guided se (3) diffusion models. arXiv preprint arXiv:2403.14088, 2024a. 22

Jiarui Lu, Xiaoyin Chen, Stephen Zhewen Lu, Chence Shi, Hongyu Guo, Yoshua Bengio, and Jian Tang. Structure language models for protein conformation generation. In The Thirteenth International Conferhttps://openreview.net/forum?id=OzUNDnpQyd. ence on Learning Representations, 2025a. URL https://openreview.net/forum?id=OzUNDnpQyd Bowen Jing, Hannes Stärk, Tommi Jaakkola, and Bonnie Berger. Generative modeling of molecular dynamics trajectories. Advances in Neural Information Processing Systems, 37:40534–40564, 2024b. Yuning Shen, Lihao Wang, Huizhuo Yuan, Yan Wang, Bangji Yang, and Quanquan Gu. Simultaneous modeling of protein conformation and dynamics via autoregression. In ICML 2025 Generative AI and https://openreview.net/forum?id=BaZL1cQzV0. Biology (GenBio) Workshop, 2025. URL https://openreview.net/forum?id=BaZL1cQzV0 Yaoyao Xu, Di Wang, Zihan Zhou, Tianshu Yu, and Mingchen Chen. TEMPO: Temporal multi-scale autoregressive generation of protein conformational ensembles. In The Thirty-ninth Annual Conference on https://openreview.net/forum?id=0wV5HR7M4P. Neural Information Processing Systems, 2025a. URL https://openreview.net/forum?id=0wV5HR7M4P Bin Feng, Jiying Zhang, Xinni Zhang, Zijing Liu, and Yu Li. BioMD: All-atom generative model for biomolecular dynamics simulation. In The Fourteenth International Conference on Learning Representations, 2026a. URL https://openreview.net/forum?id=LQDeJk6NOr https://openreview.net/forum?id=LQDeJk6NOr. Frank Noé, Simon Olsson, Jonas Köhler, and Hao Wu. Boltzmann generators: Sampling equilibrium states of many-body systems with deep learning. Science, 365(6457):eaaw1147, 2019. Leon Klein and Frank Noé. Transferable boltzmann generators. Advances in Neural Information Processing Systems, 37:45281–45314, 2024. Charlie B. Tan, Joey Bose, Chen Lin, Leon Klein, Michael M. Bronstein, and Alexander Tong. Scalable equilibrium sampling with sequential boltzmann generators. In Forty-second International Conference on Machine Learning, 2025a. URL https://openreview.net/forum?id=U7eMoRDIGi https://openreview.net/forum?id=U7eMoRDIGi. Charlie B. Tan, Majdi Hassan, Leon Klein, Saifuddin Syed, Dominique Beaini, Michael M. Bronstein, Alexander Tong, and Kirill Neklyudov. Amortized sampling with transferable normalizing flows. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025b. URL https://openreview.net/forum?id=JenfC3ovzU. https://openreview.net/forum?id=JenfC3ovzU Yuancheng Sun, Yuxuan Ren, Zhaoming Chen, Xu Han, Kang Liu, and Qiwei Ye. Epo: Diverse and realistic protein ensemble generation via energy preference optimization. arXiv preprint arXiv:2511.10165, 2025. Jiarui Lu, Xiaoyin Chen, Stephen Zhewen Lu, Aurelie Lozano, Vijil Chenthamarakshan, Payel Das, and Jian Tang. Aligning protein conformation ensemble generation with physical feedback. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 40436–40451. PMLR, 13–19 Jul 2025b. URL https://proceedings.mlr.press/v267/lu25b.html. https://proceedings.mlr.press/v267/lu25b.html Hilbert Yuen In Lam, Sebastián Pujalte Ojeda, Michaela Brezinova, Josef Hanke, Xing Er Ong, Yuguang Mu, and Michele Vendruscolo. Metadiffusion: inference-time meta-energy bias10.64898/2026.02.10.704873. URL ing of biomolecular diffusion models. bioRxiv, 2026. doi:10.64898/2026.02.10.704873 https://www.biorxiv.org/content/early/2026/02/11/2026.02.10.704873. https://www.biorxiv.org/content/early/2026/02/11/2026.02.10.704873 Tong Wang, Xinheng He, Mingyu Li, Yatao Li, Ran Bi, Yusong Wang, Chaoran Cheng, Xiangzhen Shen, Jiawei Meng, He Zhang, Haiguang Liu, Zun Wang, Shaoning Li, Bin Shao, and Tie-Yan Liu. Ab initio characterization of protein molecular dynamics with ai2bmd. Nature, 635(8040):1019–1027, 11 2024b. Oliver T. Unke, Martin Stöhr, Stefan Ganscha, Thomas Unterthiner, Hartmut Maennel, Sergii Kashubin, Daniel Ahlin, Michael Gastegger, Leonardo Medrano Sandonas, Joshua T. Berryman, Alexandre Tkatchenko, and Klaus-Robert Müller. Biomolecular dynamics with machine-learned quantum-mechanical force fields trained on diverse chemical fragments. Science Advances, 10(14), 4 2024. 23

Kenichiro Takaba, Anika J Friedman, Chapin E Cavender, Pavan Kumar Behara, Iván Pulido, Michael M Henry, Hugo MacDermott-Opeskin, Christopher R Iacovella, Arnav M Nagle, Alexander Matthew Payne, Michael R Shirts, David L Mobley, John D Chodera, and Yuanqing Wang. Machine-learned molecular mechanics force fields from large-scale quantum chemical data. Chemical science, 15:12861–12878, 6 2024. Roman Zubatyuk, Malgorzata Biczysko, Kavindri Ranasinghe, Nigel W. Moriarty, Hatice Gokcan, Holger Kruse, Billy K. Poon, Paul D. Adams, Mark P. Waller, Adrian E. Roitberg, Olexandr Isayev, and Pavel V. Afonine. Aquaref: machine learning accelerated quantum refinement of protein structures. Nature Communications, 16(1), 10 2025. Jiang Wang, Simon Olsson, Christoph Wehmeyer, Adrià Pérez, Nicholas E. Charron, Gianni de Fabritiis, Frank Noé, and Cecilia Clementi. Machine learning of coarse-grained molecular dynamics force fields. ACS Central Science, 5(5):755–767, 2019. Nicholas E Charron, Klara Bonneau, Aldo S Pasos-Trejo, Andrea Guljas, Yaoyi Chen, Félix Musil, Jacopo Venturin, Daria Gusew, Iryna Zaporozhets, Andreas Krämer, Clark Templeton, Atharva Kelkar, Aleksander E P Durumeric, Simon Olsson, Adrià Pérez, Maciej Majewski, Brooke E Husic, Ankit Patel, Gianni De Fabritiis, Frank Noé, and Cecilia Clementi. Navigating protein landscapes with a machinelearned transferable coarse-grained model. Nature chemistry, 17(8):1284–1292, 8 2025. Xinqiang Ding and Bin Zhang. Contrastive learning of coarse-grained force fields. Journal of Chemical Theory and Computation, 18(10):6334–6344, 2022. Michael Plainer, Hao Wu, Leon Klein, Stephan Günnemann, and Frank Noe. Consistent sampling and simulation: Molecular dynamics with energy-based diffusion models. In The Thirty-ninth Annual Conference on https://openreview.net/forum?id=gzYuvZg28E. Neural Information Processing Systems, 2025. URL https://openreview.net/forum?id=gzYuvZg28E Luigi Bonati, GiovanniMaria Piccini, and Michele Parrinello. Deep learning the slow modes for rare events sampling. Proceedings of the National Academy of Sciences, 118(44):e2113533118, 2021. Dongdong Wang, Yanze Wang, Junhan Chang, Linfeng Zhang, Han Wang, and Weinan E. Efficient sampling of high-dimensional free energy landscapes using adaptive reinforced dynamics. Nature Computational Science, 2(1):20–29, 12 2021. Martin Sipka, Johannes CB Dietschreit, Lukáš Grajciar, and Rafael Gómez-Bombarelli. Differentiable simulations for enhanced sampling of rare events. In International Conference on Machine Learning, pages 31990–32007. PMLR, 2023. Wei Chen and Andrew L. Ferguson. Molecular enhanced sampling with autoencoders: On-the-fly collective variable discovery and accelerated free energy landscape exploration. Journal of Computational Chemistry, 39(25):2079–2102, 9 2018. David E Shaw, Paul Maragakis, Kresten Lindorff-Larsen, Stefano Piana, Ron O Dror, Michael P Eastwood, Joseph A Bank, John M Jumper, John K Salmon, Yibing Shan, et al. Atomic-level characterization of the structural dynamics of proteins. Science, 330(6002):341–346, 2010. Kresten Lindorff-Larsen, Stefano Piana, Ron O Dror, and David E Shaw. How fast-folding proteins fold. Science, 334(6055):517–520, 2011. Diego Del Alamo, Davide Sala, Hassane S Mchaourab, and Jens Meiler. Sampling alternative conformational states of transporters and receptors with alphafold2. Elife, 11:e75751, 2022. Hannah K Wayment-Steele, Adedolapo Ojoawo, Renee Otten, Julia M Apitz, Warintra Pitsawong, Marc Hömberger, Sergey Ovchinnikov, Lucy Colwell, and Dorothee Kern. Predicting multiple conformations via sequence clustering and alphafold2. Nature, 625(7996):832–839, 2024.

24

Kaihui Cheng, Ce Liu, Qingkun Su, Jun Wang, Liwei Zhang, Yining Tang, Yao Yao, Siyu Zhu, and Yuan Qi. 4d diffusion for dynamic protein structure prediction with reference and motion guidance. Proceedings of the AAAI Conference on Artificial Intelligence, 39(1):93–101, 4 2025. Jianmin Wang, Xun Wang, Yanyi Chu, Chunyan Li, Xue Li, Xiangyu Meng, Yitian Fang, Kyoung Tai No, Jiashun Mao, and Xiangxiang Zeng. Exploring the conformational ensembles of protein–protein complex with transformer-based generative model. Journal of Chemical Theory and Computation, 20(11):4469– 4480, 2024c. Yusong Wang, Youjun Xu, Wentao Li, Haoyu Yu, Wenjuan Tan, Shaoning Li, Qiaojing Huang, Nanjun Chen, Xuan Wu, Qilong Wu, et al. Learning the all-atom equilibrium distribution of biomolecular interactions at scale. bioRxiv, pages 2026–03, 2026. Giacomo Janson, Alexander Jussupow, and Michael Feig. dependent structural ensembles of proteins. bioRxiv, 2025.

Deep generative modeling of temperature-

Liang Shi, Jiarui Lu, Junqi Liu, Chence Shi, Zhi Yang, and Jian Tang. Atomic trajectory modeling with https://arxiv.org/abs/2603.17633. state space models for biomolecular dynamics, 2026. URL https://arxiv.org/abs/2603.17633 Sarah Lewis, Tim Hempel, José Jiménez-Luna, Michael Gastegger, Yu Xie, Andrew YK Foong, Victor García Satorras, Osama Abdin, Bastiaan S Veeling, Iryna Zaporozhets, et al. Scalable emulation of protein equilibrium ensembles with generative deep learning. Science, 389(6761):eadv9817, 2025. Bin Feng, Jiying Zhang, Xinni Zhang, Ming Zhang, Patrick Barth, Zijing Liu, and Yu Li. Physically grounded generative modeling of all-atom biomolecular dynamics. bioRxiv, pages 2026–02, 2026b. Saro Passaro, Gabriele Corso, Jeremy Wohlwend, Mateo Reveiz, Stephan Thaler, Vignesh Ram Somnath, Noah Getz, Tally Portnoi, Julien Roy, Hannes Stark, et al. Boltz-2: Towards accurate and efficient binding affinity prediction. BioRxiv, 2025. Allan dos Santos Costa, Manvitha Ponnapati, Dana Rubin, Tess Smidt, and Joseph Jacobson. Accelerating protein molecular dynamics simulation with deepjump. arXiv preprint arXiv:2509.13294, 2025. Shuxin Zheng, Jiyan He, Chang Liu, Yu Shi, Ziheng Lu, Weitao Feng, Fusong Ju, Jiaxi Wang, Jianwei Zhu, Yaosen Min, et al. Predicting equilibrium distributions for molecular systems with deep learning. Nature Machine Intelligence, 6(5):558–567, 2024. Zirui Fan, Junjie Zhu, and Hai-Feng Chen. Dynafold: A latent diffusion based generative framework for protein dynamic trajectory. bioRxiv, 2025a. Wei Lu, Jixian Zhang, Weifeng Huang, Ziqiao Zhang, Xiangyu Jia, Zhenyu Wang, Leilei Shi, Chengtao Li, Peter G Wolynes, and Shuangjia Zheng. Dynamicbind: predicting ligand-specific protein-ligand complex structure with a deep equivariant generative model. Nature Communications, 15(1):1071, 2024b. Allan Dos Santos Costa, Ilan Mitnikov, Franco Pellegrini, Ameya Daigavane, Mario Geiger, Zhonglin Cao, Karsten Kreis, Tess Smidt, Emine Kucukbenli, and Joseph Jacobson. Equijump: Protein dynamics simulation via so (3)-equivariant stochastic interpolants. arXiv preprint arXiv:2410.09667, 2024. Shaoning Li, Yusong Wang, Mingyu Li, Jian Zhang, Bin Shao, Nanning Zheng, and Jian Tang. F3 low: Frame-to-frame coarse-grained molecular dynamics with se (3) guided flow matching. arXiv preprint arXiv:2405.00751, 2024. Aditya Sengar, Jiying Zhang, Pierre Vandergheynst, and PATRICK BARTH. Beyond ensembles: Simulating all-atom protein dynamics in a learned latent space. In The Fourteenth International Conference on Learning Representations, 2026. URL https://openreview.net/forum?id=AwowReRWXI https://openreview.net/forum?id=AwowReRWXI.

25

Junjie Zhu, Zhengxin Li, Zhuoqi Zheng, Bo Zhang, Bozitao Zhong, Jie Bai, Xiaokun Hong, Taifeng Wang, Ting Wei, Jianyi Yang, et al. Accurate generation of conformational ensembles for intrinsically disordered proteins with idpfold. Advanced Science, page e11636, 2025. Giacomo Janson, Gilberto Valdes-Garcia, Lim Heo, and Michael Feig. Direct generation of protein conformational ensembles via machine learning. Nature Communications, 14(1):774, 2023. Mathias Schreiner, Ole Winther, and Simon Olsson. Implicit transfer operator learning: Multiple timeresolution models for molecular dynamics. Advances in Neural Information Processing Systems, 36:36449– 36462, 2023. Kacper Kapuśniak, Cristian Gabellini, Michael M. Bronstein, Prudencio Tossou, and Francesco Di Giovanni. Mars-FM: Generative modeling of molecular dynamics via markov state models. In The Fourteenth International Conference on Learning Representations, 2026. URL https://openreview.net/forum?id=jP3HnYXoIp. https://openreview.net/forum?id=jP3HnYXoIp Mhd Hussein Murtada, Z. Faidon Brotzakis, and Michele Vendruscolo. Md-llm-1: A large language model for molecular dynamics, 2025. URL https://arxiv.org/abs/2508.03709 https://arxiv.org/abs/2508.03709. Yaowei Jin, Qi Huang, Ziyang Song, Mingyue Zheng, Dan Teng, and Qian Shi. P2dflow: A protein ensemble generative model with se(3) flow matching. Journal of Chemical Theory and Computation, 21(6):3288– 3296, 2025. Ivan Anishchenko, Yakov Kipnis, Indrek Kalvet, Guangfeng Zhou, Rohith Krishna, Samuel J. Pellock, Anna Lauko, Gyu Rie Lee, Linna An, Justas Dauparas, Frank DiMaio, and David Baker. Modeling protein–small molecule conformational ensembles with placer. Proceedings of the National Academy of Sciences, 122(45): e2427161122, 2025. Kai Xu, Jianmin Wang, Mingquan Liu, Kewei Zhou, Shaolong Lin, Weihong Li, Lin Shi, Peng Zhou, Huanxiang Liu, and Xiaojun Yao. Efficient generation of protein and protein–protein complex dynamics via se(3)-parameterized diffusion models. Journal of Chemical Information and Modeling, 65(22):12366–12376, 2025b. Ziyang Yu, Wenbing Huang, and Yang Liu. Unified biomolecular trajectory generation via pretrained variational bridge. In The Fourteenth International Conference on Learning Representations, 2026. URL https://openreview.net/forum?id=8HH9dBOxwu. https://openreview.net/forum?id=8HH9dBOxwu Yinuo Zhang, Sophia Tang, and Pranam Chatterjee. ScooBDoob: Schrödinger bridge with doob’s htransform for molecular dynamics. In 2nd edition of Frontiers in Probabilistic Inference: Learning meets https://openreview.net/forum?id=PtRLxAAKzK. Sampling, 2025. URL https://openreview.net/forum?id=PtRLxAAKzK Yuyang Wang, Jiarui Lu, Navdeep Jaitly, Josh Susskind, and Miguel Angel Bautista. Simplefold: Folding proteins is simpler than you think. arXiv preprint arXiv:2509.18480, 2025a. Nima Shoghi, Yuxuan Liu, Yuning Shen, Rob Brekelmans, Pan Li, and Quanquan Gu. Scalable spatiotemporal SE(3) diffusion for long-horizon protein dynamics. In The Fourteenth International Conference on Learning Representations, 2026. URL https://openreview.net/forum?id=Q1JpRZkR3S https://openreview.net/forum?id=Q1JpRZkR3S. Leon Klein, Andrew Foong, Tor Fjelde, Bruno Mlodozeniec, Marc Brockschmidt, Sebastian Nowozin, Frank Noé, and Ryota Tomioka. Timewarp: Transferable acceleration of molecular dynamics by learning timecoarsened dynamics. Advances in Neural Information Processing Systems, 36:52863–52883, 2023a. Jiahao Fan, Ziyao Li, Eric Alcaide, Guolin Ke, Huaqing Huang, and Weinan E. Accurate conformation sampling via protein structural diffusion. Journal of Chemical Information and Modeling, 64(22):8414– 8426, 2024.

26

Ziyang Yu, Wenbing Huang, and Yang Liu. Unisim: A unified simulator for time-coarsened dynamics of biomolecules. In Forty-second International Conference on Machine Learning, 2025. URL https://openreview.net/forum?id=W598uW37zt. https://openreview.net/forum?id=W598uW37zt Thomas Hayes, Roshan Rao, Halil Akin, Nicholas J Sofroniew, Deniz Oktay, Zeming Lin, Robert Verkuil, Vincent Q Tran, Jonathan Deaton, Marius Wiggert, et al. Simulating 500 million years of evolution with a language model. Science, 387(6736):850–858, 2025. Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of naacL-HLT, volume 1. Minneapolis, Minnesota, 2019. Leon Klein, Andreas Krämer, and Frank Noe. Equivariant flow matching. In Thirty-seventh Conference on https://openreview.net/forum?id=eLH2NFOO1B. Neural Information Processing Systems, 2023b. URL https://openreview.net/forum?id=eLH2NFOO1B Jonas Köhler, Andreas Krämer, and Frank Noé. Smooth normalizing flows. Advances in Neural Information Processing Systems, 34:2796–2809, 2021. Laurence Illing Midgley, Vincent Stimper, Gregor N. C. Simm, Bernhard Schölkopf, and José Miguel Hernández-Lobato. Flow annealed importance sampling bootstrap. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=XCTVFJwS9LJ https://openreview.net/forum?id=XCTVFJwS9LJ. Joseph C Kim, David Bloore, Karan Kapoor, Jun Feng, Ming-Hong Hao, and Mengdi Wang. Scalable normalizing flows enable boltzmann generators for macromolecules. arXiv preprint arXiv:2401.04246, 2024. Tara Akhound-Sadegh, Jarrid Rector-Brooks, Avishek Joey Bose, Sarthak Mittal, Pablo Lemos, ChengHao Liu, Marcin Sendera, Siamak Ravanbakhsh, Gauthier Gidel, Yoshua Bengio, Nikolay Malkin, and Alexander Tong. Iterated denoising energy matching for sampling from boltzmann densities. In Proceedings of the 41st International Conference on Machine Learning, ICML’24. JMLR.org, 2024. RuiKang OuYang, Bo Qiang, Zixing Song, and José Miguel Hernández-Lobato. Bnem: A boltzmann sampler based on bootstrapped noised energy matching, 2025. URL https://arxiv.org/abs/2409.09787 https://arxiv.org/abs/2409.09787. Tara Akhound-Sadegh, Jungyoon Lee, Joey Bose, Valentin De Bortoli, Arnaud Doucet, Michael M. Bronstein, Dominique Beaini, Siamak Ravanbakhsh, Kirill Neklyudov, and Alexander Tong. Progressive inference-time annealing of diffusion models for sampling from boltzmann densities. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. URL https://openreview.net/forum?id=vf2GHcxzMV. https://openreview.net/forum?id=vf2GHcxzMV Henrik Schopmans and Pascal Friederich. Temperature-annealed boltzmann tors. In Forty-second International Conference on Machine Learning, 2025. https://openreview.net/forum?id=RqtRSrCbNu. https://openreview.net/forum?id=RqtRSrCbNu

generaURL

Niclas Dern, Lennart Redl, Sebastian Pfister, Marcel Kollovieh, David Lüdke, and Stephan Günnemann. Energy-weighted flow matching: Unlocking continuous normalizing flows for efficient and scalable boltzmann sampling, 2025. URL https://arxiv.org/abs/2509.03726 https://arxiv.org/abs/2509.03726. Christopher von Klitzing, Denis Blessing, Henrik Schopmans, Pascal Friederich, and Gerhard Neumann. Learning boltzmann generators via constrained mass transport. In The Fourteenth International Conferhttps://openreview.net/forum?id=MQmrcX5jnk. ence on Learning Representations, 2026. URL https://openreview.net/forum?id=MQmrcX5jnk Danyal Rehman, Oscar Davis, Jiarui Lu, Jian Tang, Michael Bronstein, Yoshua Bengio, Alexander Tong, and Avishek Joey Bose. Efficient regression-based training of normalizing flows for boltzmann generators. In The Fourteenth International Conference on Learning Representations, 2026. URL https://arxiv.org/abs/2506.01158. https://arxiv.org/abs/2506.01158 27

Henrik Schopmans, Christopher von Klitzing, and Pascal Friederich. Efficient training of boltzmann generhttps://arxiv.org/abs/2602.03729. ators using off-policy log-dispersion regularization, 2026. URL https://arxiv.org/abs/2602.03729 Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio. Density estimation using real nvp. arXiv preprint arXiv:1605.08803, 2016. Shuangfei Zhai, Ruixiang ZHANG, Preetum Nakkiran, David Berthelot, Jiatao Gu, Huangjie Zheng, Tianrong Chen, Miguel Ángel Bautista, Navdeep Jaitly, and Joshua M. Susskind. Normalizing flows are capable generative models. In Forty-second International Conference on Machine Learning, 2025. URL https://openreview.net/forum?id=2uheUFcFsM. https://openreview.net/forum?id=2uheUFcFsM Siyuan Chen, Minghao Guo, Caoliwen Wang, Anka He Chen, Yikun Zhang, Jingjing Chai, Yin Yang, Wojciech Matusik, and Peter Yichen Chen. Physically valid biomolecular interaction modeling with gaussseidel projection. In The Fourteenth International Conference on Learning Representations, 2026. URL https://openreview.net/forum?id=sJABnBEYeh. https://openreview.net/forum?id=sJABnBEYeh Juno Nam, Bálint Máté, Artur P. Toshev, Manasa Kaniselvan, Rafael Gomez-Bombarelli, Ricky T. Q. Chen, Brandon M. Wood, Guan-Horng Liu, and Benjamin Kurt Miller. Enhancing diffusion-based sampling with molecular collective variables. In The Fourteenth International Conference on Learning Representations, https://openreview.net/forum?id=1bJN1EQByS. 2026. URL https://openreview.net/forum?id=1bJN1EQByS Adil Kabylda, J Thorben Frank, Sergio Suárez-Dou, Almaz Khabibrakhmanov, Leonardo Medrano Sandonas, Oliver T Unke, Stefan Chmiela, Klaus-Robert Müller, and Alexandre Tkatchenko. Molecular simulations with a pretrained neural network and universal pairwise force fields. Journal of the American Chemical Society, 147(37):33723–33734, 2025. Maciej Majewski, Adrià Pérez, Philipp Thölke, Stefan Doerr, Nicholas E. Charron, Toni Giorgino, Brooke E. Husic, Cecilia Clementi, Frank Noé, and Gianni De Fabritiis. Machine learning coarse-grained potentials of protein thermodynamics. Nature Communications, 14(1), 9 2023. Jonas Köhler, Yaoyi Chen, Andreas Krämer, Cecilia Clementi, and Frank Noé. Flow-matching: Efficient coarse-graining of molecular dynamics without forces. Journal of Chemical Theory and Computation, 19 (3):942–952, 2023. Shriram Chennakesavalu, David J Toomer, and Grant M Rotskoff. Ensuring thermodynamic consistency with invertible coarse-graining. The Journal of chemical physics, 158(12):124126, 3 2023. Marloes Arts, Victor Garcia Satorras, Chin-Wei Huang, Daniel Zügner, Marco Federici, Cecilia Clementi, Frank Noé, Robert Pinsler, and Rianne van den Berg. Two for one: Diffusion models and force fields for coarse-grained molecular dynamics. Journal of Chemical Theory and Computation, 19(18):6151–6159, 2023. Klara Bonneau, Jonas Lederer, Clark Templeton, David Rosenberger, Lorenzo Giambagli, Klaus-Robert Müller, and Cecilia Clementi. Peering inside the black box by learning the relevance of many-body functions in neural network potentials. Nature Communications, 16(1), 11 2025. Meital Bojan, Sanketh Vedula, Sai Advaith Maddipatla, Nadav Bojan, Anar Rzayev, Federico Napoli, Paul Schanda, and Alexander Bronstein. Representing local protein environments with machine learning force fields. In The Fourteenth International Conference on Learning Representations, 2026. URL https://openreview.net/forum?id=9ZogcRkhoG. https://openreview.net/forum?id=9ZogcRkhoG Justin S Smith, Olexandr Isayev, and Adrian E Roitberg. Ani-1: an extensible neural network potential with dft accuracy at force field computational cost. Chemical science, 8(4):3192–3203, 2017. Philipp Thölke and Gianni De Fabritiis. Equivariant transformers for neural network based molecular potentials. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=zNHzqZ9wrRB. https://openreview.net/forum?id=zNHzqZ9wrRB 28

Simon Batzner, Albert Musaelian, Lixin Sun, Mario Geiger, Jonathan P Mailoa, Mordechai Kornbluth, Nicola Molinari, Tess E Smidt, and Boris Kozinsky. E (3)-equivariant graph neural networks for dataefficient and accurate interatomic potentials. Nature communications, 13(1):2453, 2022. Yuanqing Wang, Josh Fass, Benjamin Kaminow, John E. Herr, Dominic Rufa, Ivy Zhang, Iván Pulido, Mike Henry, Hannah E. Bruce Macdonald, Kenichiro Takaba, and John D. Chodera. End-to-end differentiable construction of molecular mechanics force fields. Chemical Science, 13(41):12016–12033, 2022. Abdul Raafik Arattu Thodika, Xiaoliang Pan, Yihan Shao, and Kwangho Nam. Machine learning quantum mechanical/molecular mechanical potentials: Evaluating transferability in dihydrofolate reductasecatalyzed reactions. Journal of Chemical Theory and Computation, 21(2):817–832, 2025. Xinhu Sha, Zhuo Chen, Daiqian Xie, and Yanzi Zhou. Modeling enzyme reaction and mutation by direct machine learning/molecular mechanics simulations. Journal of Chemical Theory and Computation, 21(9): 4335–4346, 2025. Valentin Gradisteanu, Elliot W. Chan, Lester Hedges, Meritxell Malagarriga, Rolf David, Miguel de la Puente, Damien Laage, Iñaki Tuñón, Marc W. van der Kamp, and Kirill Zinovjev. Simulating enzyme catalysis with electrostatically embedded machine learn10.26434/chemrxiv-2025-nw9lt. URL ing potentials. ChemRxiv, 2025(0707), 2025. doi:10.26434/chemrxiv-2025-nw9lt https://chemrxiv.org/doi/abs/10.26434/chemrxiv-2025-nw9lt. https://chemrxiv.org/doi/abs/10.26434/chemrxiv-2025-nw9lt Xujian Wang, Haocheng Tang, Xiongwu Wu, Bernard Brooks, Junmei Wang, and WanLu Li. Redefining computational enzymology with multiscale machine learning/molecular mechanics metadynamics: Deciphering catalytic mechanism and stereoselectivity in diels–alderases. ChemRxiv, 2025(1002), 2025b. doi:10.26434/chemrxiv-2025-x2k6j 10.26434/chemrxiv-2025-x2k6j. URL https://chemrxiv.org/doi/abs/10.26434/chemrxiv-2025-x2k6j. https://chemrxiv.org/doi/abs/10.26434/chemrxiv-2025-x2k6j Peter Eastman, Pavan Kumar Behara, David L. Dotson, Raimondas Galvelis, John E. Herr, Josh T. Horton, Yuezhi Mao, John D. Chodera, Benjamin P. Pritchard, Yuanqing Wang, Gianni De Fabritiis, and Thomas E. Markland. Spice, a dataset of drug-like molecules and peptides for training machine learning potentials. Scientific Data, 10(1), 1 2023. ISSN 2052-4463. Han Yang, Chenxi Hu, Yichi Zhou, Xixian Liu, Yu Shi, Jielan Li, Guanzhi Li, Zekun Chen, Shuizhou Chen, Claudio Zeni, et al. Mattersim: A deep learning atomistic model across elements, temperatures and pressures. arXiv preprint arXiv:2405.04967, 2024a. Haokai Hong, Wanyu Lin, Zhang Chusong, and KC Tan. Geometric graph neural diffusion for stable molecular dynamics. In The Fourteenth International Conference on Learning Representations, 2026. URL https://openreview.net/forum?id=T8VcTykTf1 https://openreview.net/forum?id=T8VcTykTf1. Paulo C. T. Souza, Riccardo Alessandri, Jonathan Barnoud, Sebastian Thallmair, Ignacio Faustino, Fabian Grünewald, Ilias Patmanidis, Haleh Abdizadeh, Bart M. H. Bruininks, Tsjerk A. Wassenaar, Peter C. Kroon, Josef Melcr, Vincent Nieto, Valentina Corradi, Hanif M. Khan, Jan Domański, Matti Javanainen, Hector Martinez-Seara, Nathalie Reuter, Robert B. Best, Ilpo Vattulainen, Luca Monticelli, Xavier Periole, D. Peter Tieleman, Alex H. de Vries, and Siewert J. Marrink. Martini 3: a general purpose force field for coarse-grained molecular dynamics. Nature Methods, 18(4):382–388, 3 2021. ISSN 1548-7091. Brooke E. Husic, Nicholas E. Charron, Dominik Lemm, Jiang Wang, Adrià Pérez, Maciej Majewski, Andreas Krämer, Yaoyi Chen, Simon Olsson, Gianni de Fabritiis, Frank Noé, and Cecilia Clementi. Coarse graining molecular dynamics with graph neural networks. The Journal of Chemical Physics, 153(19), 11 2020. Winfried Ripken, Michael Plainer, Gregor Lied, Thorben Frank, Oliver T. Unke, Stefan Chmiela, Frank Noé, and Klaus-Robert Müller. Learning hamiltonian flow maps: Mean flow consistency for large-timestep molecular dynamics, 2026. URL https://arxiv.org/abs/2601.22123 https://arxiv.org/abs/2601.22123.

29

Mohammad M. Sultan and Vijay S. Pande. Automated design of collective variables using supervised machine learning. The Journal of Chemical Physics, 149(9), 9 2018. Wei Chen, Aik Rui Tan, and Andrew L. Ferguson. Collective variable discovery and enhanced sampling using autoencoders: Innovations in network architecture and error function design. The Journal of Chemical Physics, 149(7), 5 2018. Emmanuel Oluwatobi Salawu. Desp: Deep enhanced sampling of proteins’ conformation spaces using aiinspired biasing forces. Frontiers in molecular biosciences, 8:587151, 5 2021. Zineb Belkacemi, Paraskevi Gkeka, Tony Lelièvre, and Gabriel Stoltz. Chasing collective variables using autoencoders and biased trajectories. Journal of Chemical Theory and Computation, 18(1):59–78, 2022. Diego E. Kleiman and Diwakar Shukla. Active learning of the conformational ensemble of proteins using maximum entropy vampnets. Journal of Chemical Theory and Computation, 19(14):4377–4388, 2023. Seonghyun Park, Kiyoung Seong, Soojung Yang, Rafael Gomez-Bombarelli, and Sungsoo Ahn. Learning collective variables from time-lagged generation. In ICML 2025 Generative AI and Biology (GenBio) https://openreview.net/forum?id=Ki24kOIH6l. Workshop, 2025. URL https://openreview.net/forum?id=Ki24kOIH6l Subarna Sasmal, Martin McCullagh, and Glen M. Hocky. Reaction coordinates for conformational transitions using linear discriminant analysis on positions. Journal of Chemical Theory and Computation, 19(14): 4427–4435, 2023. Enrico Trizio and Michele Parrinello. From enhanced sampling to reaction profiles. The Journal of Physical Chemistry Letters, 12(35):8621–8626, 2021. Mihir Bafna, Bowen Jing, and Bonnie Berger. Learning residue level protein dynamics with multiscale gaussians. In The Fourteenth International Conference on Learning Representations, 2026. URL https://openreview.net/forum?id=uKn9PdREBA. https://openreview.net/forum?id=uKn9PdREBA Zahra Shamsi, Kevin J. Cheng, and Diwakar Shukla. Reinforcement learning based adaptive sampling: Reaping rewards by exploring protein conformational landscapes. The Journal of Physical Chemistry B, 122(35):8386–8395, 8 2018. Diego E. Kleiman and Diwakar Shukla. Multiagent reinforcement learning-based adaptive sampling for conformational dynamics of proteins. Journal of Chemical Theory and Computation, 18(9):5422–5434, 2022. Hyungro Lee, Matteo Turilli, Shantenu Jha, Debsindhu Bhowmik, Heng Ma, and Arvind Ramanathan. Deepdrivemd: Deep-learning driven adaptive molecular simulations for protein folding. In 2019 IEEE/ACM Third Workshop on Deep Learning on Supercomputers (DLS), pages 12–19, 2019. doi:10.1109/DLS49591.2019.00007 10.1109/DLS49591.2019.00007. Adrià Pérez, Pablo Herrera-Nieto, Stefan Doerr, and Gianni De Fabritiis. Adaptivebandit: A multi-armed bandit framework for adaptive sampling in molecular simulations. Journal of Chemical Theory and Computation, 16(7):4685–4693, 2020. Wenhui Shen, Kaiwei Wan, Dechang Li, Huajian Gao, and Xinghua Shi. Adaptive cvgen: Leveraging reinforcement learning for advanced sampling in protein folding and chemical reactions. Proceedings of the National Academy of Sciences, 121(45):e2414205121, 2024. Nicholas Ho, John Kevin Cava, John Vant, Ankita Shukla, Jake Miratsky, Pavan Turaga, Ross Maciejewski, and Abhishek Singharoy. Learning free energy pathways through reinforcement learning of adaptive steered molecular dynamics. bioRxiv, 10 2022.

30

Jiahao Fan, Yanze Wang, Dongdong Wang, and Linfeng Zhang. Rid-kit: software package designed to do enhanced sampling using reinforced dynamics. BMC Methods, 2(1), 6 2025b. Thorben Fröhlking, Valerio Rizzi, Simone Aureli, and Francesco Luigi Gervasio. Deeplne++ leveraging knowledge distillation for accelerated multi-state path-like collective variables. The Journal of Chemical Physics, 161(11), 9 2024. Soojung Yang, Juno Nam, Johannes CB Dietschreit, and Rafael Gómez-Bombarelli. Learning collective variables with synthetic data augmentation through physics-inspired geodesic interpolation. Journal of Chemical Theory and Computation, 20(15):6559–6568, 2024b. Bodhi P Vani, Akashnathan Aranganathan, Dedi Wang, and Pratyush Tiwary. Alphafold2-rave: From sequence to boltzmann ranking. Journal of chemical theory and computation, 19(14):4351–4354, 2023. Z. Faidon Brotzakis, Shengyu Zhang, Mhd Hussein Murtada, and Michele Vendruscolo. Alphafold prediction of structural ensembles of disordered proteins. Nature Communications, 16(1), 2 2025. Seonghyun Park, Kiyoung Seong, Soojung Yang, Rafael Gomez-Bombarelli, and Sungsoo Ahn. Learning collective variables from bioemu with time-lagged generation. In The Fourteenth International Conference https://openreview.net/forum?id=1PYj4fMeLe. on Learning Representations, 2026. URL https://openreview.net/forum?id=1PYj4fMeLe Zakarya Benayad and Guillaume Stirnemann. Hamiltonian replica exchange augmented with diffusion-based generative models and importance sampling to assess biomolecular conformational basins and barriers. Journal of Chemical Theory and Computation, 21(21):10692–10704, 2025. Hao Tian, Xi Jiang, Sian Xiao, Hunter La Force, Eric C. Larson, and Peng Tao. Last: Latent space-assisted adaptive sampling for protein trajectories. Journal of Chemical Information and Modeling, 63(1):67–75, 2023. Sergio Contreras Arredondo, Chenyu Tang, Radu A. Talmazan, Alberto Megías, Cheng Giuseppe Chen, and Christophe Chipot. From atoms to dynamics: Learning the committor without collective variables, 2025. URL https://arxiv.org/abs/2507.17700 https://arxiv.org/abs/2507.17700. Haochuan Chen, Benoît Roux, and Christophe Chipot. Discovering reaction pathways, slow variables, and committor probabilities with machine learning. Journal of Chemical Theory and Computation, 19(14): 4414–4426, 2023. Alberto Megías, Sergio Contreras Arredondo, Cheng Giuseppe Chen, Chenyu Tang, Benoît Roux, and Christophe Chipot. Iterative variational learning of committor-consistent transition pathways using artificial neural networks. Nature computational science, 5(7):592–602, 7 2025. ISSN 2662-8457. doi:10.1038/s43588-025-00828-3 10.1038/s43588-025-00828-3. Bojun Liu, Jordan G Boysen, Ilona Christy Unarta, Xuefeng Du, Yixuan Li, and Xuhui Huang. Exploring transition states of protein conformational changes via out-of-distribution detection in the hyperspherical latent space. Nature communications, 16(1):349, 1 2025. ISSN 2041-1723. doi:10.1038/s41467-024-55228-4 10.1038/s41467-024-55228-4. Stephan Thaler, Zhiyi Wu, William G. Glass, Richard T. Bradshaw, Gail Bartlett, Prudencio Tossou, and Geoffrey P. F. Wood. Boltz-abfe: Free energy perturbation without crystal structures. Journal of Chemical Theory and Computation, 22(4):1823–1833, 2026. doi:10.1021/acs.jctc.5c01451 10.1021/acs.jctc.5c01451. Sam Giannakoulias, John Ferrie, and Andrew Apicello. Fep ω: The end of parameter tuning. ChemRxiv, 2025(1023), 2025. doi:10.26434/chemrxiv-2025-bg1t9 10.26434/chemrxiv-2025-bg1t9. Donald J. M. van P and Willem Jespers. Integrating machine learning into free energy perturbation workflows. Journal of Chemical Information and Modeling, 65(19):9856–9864, 2025. doi:10.1021/acs.jcim.5c01449 10.1021/acs.jcim.5c01449.

31

Helen M Berman, John Westbrook, Zukang Feng, Gary Gilliland, Talapady N Bhat, Helge Weissig, Ilya N Shindyalov, and Philip E Bourne. The protein data bank. Nucleic acids research, 28(1):235–242, 2000. Mihaly Varadi, Damian Bertoni, Paulyna Magana, Urmila Paramval, Ivanna Pidruchna, Malarvizhi Radhakrishnan, Maxim Tsenkov, Sreenath Nair, Milot Mirdita, Jingi Yeo, et al. Alphafold protein structure database in 2024: providing structure coverage for over 214 million protein sequences. Nucleic acids research, 52(D1):D368–D375, 2024. Jingi Yeo, Yewon Han, Nicola Bordin, Andy M Lau, Shaun M Kandathil, Hyunbin Kim, Eli Levy Karin, Milot Mirdita, David T Jones, Christine Orengo, et al. Metagenomic-scale analysis of the predicted protein structure universe. bioRxiv, pages 2025–04, 2025. Danny Reidenbach, Zhonglin Cao, Zuobai Zhang, Kieran Didi, Tomas Geffner, Guoqing Zhou, Jian Tang, Christian Dallago, Arash Vahdat, Emine Kucukbenli, et al. Consistent synthetic sequences unlock structural diversity in fully atomistic de novo protein design. arXiv preprint arXiv:2512.01976, 2025. Kieran Didi, Zuobai Zhang, Guoqing Zhou, Danny Reidenbach, Zhonglin Cao, Sooyoung Cha, Tomas Geffner, Christian Dallago, Jian Tang, Michael M. Bronstein, Martin Steinegger, Emine Kucukbenli, Arash Vahdat, and Karsten Kreis. Scaling atomistic protein binder design with generative pretraining and test-time compute. In The Fourteenth International Conference on Learning Representations (ICLR), 2026. Yann Vander Meersche, Gabriel Cretin, Aria Gheeraert, Jean-Christophe Gelly, and Tatiana Galochkina. Atlas: protein flexibility description from atomistic molecular dynamics simulations. Nucleic acids research, 52(D1):D384–D392, 2024. Ce Liu, Jun Wang, Zhiqiang Cai, Yingxu Wang, Huizhen Kuang, Kaihui Cheng, Liwei Zhang, Qingkun Su, Yining Tang, Fenglei Cao, et al. Dynamic pdb: A new dataset and a se (3) model extension by integrating dynamic behaviors and physical properties in protein structures. arXiv preprint arXiv:2408.12413, 2024. Omid Mokhtari, Emmanuelle Bignon, Hamed Khakzad, and Yasaman Karami. Dynarepo: the repository of macromolecular conformational dynamics. Nucleic Acids Research, page gkaf1130, 11 2025. Antonio Mirarchi, Toni Giorgino, and Gianni De Fabritiis. mdcath: A large-scale md dataset for data-driven computational biophysics. Scientific Data, 11(1):1299, 2024. Yihang Zhou, Chen Wei, Minghao Sun, Jin Song, Yang Li, Lin Wang, and Yang Zhang. Proteinconformers: Benchmark dataset for simulating protein conformational landscape diversity and plausibility. In The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks https://openreview.net/forum?id=GClrNUTqly. Track, 2025. URL https://openreview.net/forum?id=GClrNUTqly Amitava Roy, Ethan Ward, Illyoung Choi, Michele Cosi, Tony Edgin, Travis S Hughes, Md Shafayet Islam, Asif M Khan, Aakash Kolekar, Mariah Rayl, Isaac Robinson, Paul Sarando, Edwin Skidmore, Tyson L Swetnam, Mariah Wall, Zhuoyun Xu, Michelle L Yung, Nirav Merchant, and Travis J Wheeler. Mdrepo—an open data warehouse for community-contributed molecular dynamics simulations of proteins. Nucleic Acids Research, 53(D1):D477–D486, 11 2024. ISSN 1362-4962. Ismael Rodríguez-Espigares, Mariona Torrens-Fontanals, Johanna K. S. Tiemann, David Aranda-García, Juan Manuel Ramírez-Anguita, Tomasz Maciej Stepniewski, Nathalie Worp, Alejandro Varela-Rial, Adrián Morales-Pastor, Brian Medel-Lacruz, Gáspár Pándy-Szekeres, Eduardo Mayol, Toni Giorgino, Jens Carlsson, Xavier Deupi, Slawomir Filipek, Marta Filizola, José Carlos Gómez-Tamayo, Angel Gonzalez, Hugo Gutiérrez-de Terán, Mireia Jiménez-Rosés, Willem Jespers, Jon Kapla, George Khelashvili, Peter Kolb, Dorota Latek, Maria Marti-Solano, Pierre Matricon, Minos-Timotheos Matsoukas, Przemyslaw Miszta, Mireia Olivella, Laura Perez-Benito, Davide Provasi, Santiago Ríos, Iván R. Torrecillas, Jessica Sallander, Agnieszka Sztyler, Silvana Vasile, Harel Weinstein, Ulrich Zachariae, Peter W. Hildebrand, Gianni De Fabritiis, Ferran Sanz, David E. Gloriam, Arnau Cordomi, Ramon Guixà-González, and Jana Selent. Gpcrmd uncovers the dynamics of the 3d-gpcrome. Nature Methods, 17(8):777–787, 7 2020. 32

Phillip J. Stansfeld, Joseph E. Goose, Martin Caffrey, Elisabeth P. Carpenter, Joanne L. Parker, Simon Newstead, and Mark S.P. Sansom. Memprotmd: Automated insertion of membrane protein structures into explicit lipid membranes. Structure, 23(7):1350–1361, 7 2015. Sören von Bülow, Kristoffer E Johansson, and Kresten Lindorff-Larsen. Af-calvados: Alphafold-guided simulations of multi-domain proteins at the proteome level. bioRxiv, pages 2025–10, 2025. Till Siebenmorgen, Filipe Menezes, Sabrina Benassou, Erinc Merdivan, Kieran Didi, André Santos Dias Mourão, Radosław Kitel, Pietro Liò, Stefan Kesselheim, Marie Piraud, et al. Misato: machine learning dataset of protein–ligand complexes for structure-based drug discovery. Nature computational science, 4 (5):367–378, 2024. Maodong Li, Jiying Zhang, Bin Feng, Wenqi Zeng, Dechin Chen, Zhijun Pan, Yu Li, Zijing Liu, and Yi Isaac Yang. Enhanced sampling, public dataset and generative model for drug-protein dissociation dynamics. arXiv preprint arXiv:2504.18367, 2025. Jeffrey C Hoch, Kumaran Baskaran, Harrison Burr, John Chin, Hamid R Eghbalnia, Toshimichi Fujiwara, Michael R Gryk, Takeshi Iwata, Chojiro Kojima, Genji Kurisu, et al. Biological magnetic resonance data bank. Nucleic acids research, 51(D1):D368–D376, 2023. Alexey G Kikhney, Clemente R Borges, Dmitry S Molodenskiy, Cy M Jeffries, and Dmitri I Svergun. Sasbdb: Towards an automatically curated and validated repository for biological scattering data. Protein science, 29(1):66–75, 2020. Damiano Piovesan, Alessio Del Conte, Mahta Mehdiabadi, Maria Cristina Aspromonte, Matthias Blum, Giulio Tesei, Sören von Bülow, Kresten Lindorff-Larsen, and Silvio CE Tosatto. Mobidb in 2025: integrating ensemble properties and function annotations for intrinsically disordered proteins. Nucleic Acids Research, 53(D1):D495–D503, 2025. Maria Victoria Nugnes, Kamel Eddine Adel Bouhraoua, Mehdi Zoubiri, Rita Pancsa, Erzsébet Fichó, Peter Tompa, Damiano Piovesan, Silvio CE Tosatto, and Maria Cristina Aspromonte. Disprot in 2026: enhancing intrinsically disordered proteins accessibility, deposition, and annotation. Nucleic Acids Research, page gkaf1175, 2025. John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, et al. Highly accurate protein structure prediction with alphafold. nature, 596(7873):583–589, 2021. Alex Bateman, Maria-Jesus Martin, Sandra Orchard, Michele Magrane, Aduragbemi Adesina, Shadab Ahmad, Emily H Bowler-Barnett, Hema Bye-A-Jee, David Carpentier, Paul Denny, Jun Fan, Penelope Garmiri, Leonardo Jose da Costa Gonzales, Abdulrahman Hussein, Alexandr Ignatchenko, Giuseppe Insana, Rizwan Ishtiaq, Vishal Joshi, Dushyanth Jyothi, Swaathi Kandasaamy, Antonia Lock, Aurelien Luciani, Jie Luo, Yvonne Lussi, Juan Sebastian Martinez Marin, Pedro Raposo, Daniel L Rice, Rafael Santos, Elena Speretta, James Stephenson, Prabhat Totoo, Nidhi Tyagi, Nadya Urakova, Preethi Vasudev, Kate Warner, Supun Wijerathne, Conny Wing-Heng Yu, Rossana Zaru, Alan J Bridge, Lucila Aimo, Ghislaine Argoud-Puy, Andrea H Auchincloss, Kristian B Axelsen, Parit Bansal, Delphine Baratin, Teresa M Batista Neto, Marie-Claude Blatter, Jerven T Bolleman, Emmanuel Boutet, Lionel Breuza, Blanca Cabrera Gil, Cristina Casals-Casas, Kamal Chikh Echioukh, Elisabeth Coudert, Beatrice Cuche, Edouard de Castro, Anne Estreicher, Maria L Famiglietti, Marc Feuermann, Elisabeth Gasteiger, Pascale Gaudet, Sebastien Gehant, Vivienne Gerritsen, Arnaud Gos, Nadine Gruaz, Chantal Hulo, Nevila Hyka-Nouspikel, Florence Jungo, Arnaud Kerhornou, Philippe Le Mercier, Damien Lieberherr, Patrick Masson, Anne Morgat, Salvo Paesano, Ivo Pedruzzi, Sandrine Pilbout, Lucille Pourcel, Sylvain Poux, Monica Pozzato, Manuela Pruess, Nicole Redaschi, Catherine Rivoire, Christian J A Sigrist, Karin Sonesson, Shyamala Sundaram, Anastasia Sveshnikova, Cathy H Wu, Cecilia N Arighi, Chuming Chen, Yongxing Chen, Hongzhan Huang, Kati Laiho, Minna Lehvaslaiho, Peter McGarvey, Darren A Natale, Karen Ross, 33

C R Vinayaka, Yuqi Wang, and Jian Zhang. Uniprot: the universal protein knowledgebase in 2025. Nucleic 10.1093/nar/gkae1010. URL Acids Research, 53(D1):D609–D617, November 2024. ISSN 1362-4962. doi:10.1093/nar/gkae1010 http://dx.doi.org/10.1093/nar/gkae1010. http://dx.doi.org/10.1093/nar/gkae1010 Rommie E. Amaro, Johan Åqvist, Ivet Bahar, Federica Battistini, Adam Bellaiche, Daniel Beltran, Philip C. Biggin, Massimiliano Bonomi, Gregory R. Bowman, Richard A. Bryce, Giovanni Bussi, Paolo Carloni, David A. Case, Andrea Cavalli, Chia-En A. Chang, Thomas E. Cheatham, Margaret S. Cheung, Christophe Chipot, Lillian T. Chong, Preeti Choudhary, G. Andres Cisneros, Cecilia Clementi, Rosana CollepardoGuevara, Peter Coveney, Roberto Covino, T. Daniel Crawford, Matteo Dal Peraro, Bert L. de Groot, Lucie Delemotte, Marco De Vivo, Jonathan W. Essex, Franca Fraternali, Jiali Gao, Josep Ll. Gelpí, Francesco L. Gervasio, Fernando D. González-Nilo, Helmut Grubmüller, Marina G. Guenza, Horacio V. Guzman, Sarah Harris, Teresa Head-Gordon, Rigoberto Hernandez, Adam Hospital, Niu Huang, Xuhui Huang, Gerhard Hummer, Javier Iglesias-Fernández, Jan H. Jensen, Shantenu Jha, Wanting Jiao, William L. Jorgensen, Shina C. L. Kamerlin, Syma Khalid, Charles Laughton, Michael Levitt, Vittorio Limongelli, Erik Lindahl, Kresten Lindorff-Larsen, Sharon Loverde, Magnus Lundborg, Yun L. Luo, F. Javier Luque, Charlotte I. Lynch, Alexander D. MacKerell, Alessandra Magistrato, Siewert J. Marrink, Hugh Martin, J. Andrew McCammon, Kenneth Merz, Vicent Moliner, Adrian J. Mulholland, Sohail Murad, Athi N. Naganathan, Shikha Nangia, Frank Noe, Agnes Noy, Julianna Oláh, Megan L. O’Mara, Mary Jo Ondrechen, Jose N. Onuchic, Alexey Onufriev, Sílvia Osuna, Giulia Palermo, Anna R. Panchenko, Sergio Pantano, Carol Parish, Michele Parrinello, Alberto Perez, Tomas Perez-Acle, Juan R. Perilla, B. Montgomery Pettitt, Adriana Pietropaolo, Jean-Philip Piquemal, Adolfo B. Poma, Matej Praprotnik, Maria J. Ramos, Pengyu Ren, Nathalie Reuter, Adrian Roitberg, Edina Rosta, Carme Rovira, Benoit Roux, Ursula Rothlisberger, Karissa Y. Sanbonmatsu, Tamar Schlick, Alexey K. Shaytan, Carlos Simmerling, Jeremy C. Smith, Yuji Sugita, Katarzyna Świderek, Makoto Taiji, Peng Tao, D. Peter Tieleman, Irina G. Tikhonova, Julian Tirado-Rives, Iñaki Tuñón, Marc W. van der Kamp, David van der Spoel, Sameer Velankar, Gregory A. Voth, Rebecca Wade, Ariel Warshel, Valerie Vaissier Welborn, Stacey D. Wetmore, Travis J. Wheeler, Chung F. Wong, Lee-Wei Yang, Martin Zacharias, and Modesto Orozco. The need to implement fair principles in biomolecular simulations. Nature Methods, 22(4):641–645, 4 2025. Damiano Piovesan, Alessio Del Conte, Damiano Clementel, Alexander Miguel Monzon, Martina Bevilacqua, Maria Cristina Aspromonte, Javier A Iserte, Fernando E Orti, Cristina Marino-Buslje, and Silvio CE Tosatto. Mobidb: 10 years of intrinsically disordered proteins. Nucleic acids research, 51(D1):D438–D444, 2023. Giulio Imbalzano, Yongbin Zhuang, Venkat Kapil, Kevin Rossi, Edgar A. Engel, Federico Grasselli, and Michele Ceriotti. Uncertainty estimation for molecular dynamics and sampling. The Jour10.1063/5.0036522. URL nal of Chemical Physics, 154(7), February 2021. ISSN 1089-7690. doi:10.1063/5.0036522 http://dx.doi.org/10.1063/5.0036522. http://dx.doi.org/10.1063/5.0036522

34

Record · ID 141512 · SHA-256 5616bb64f4b3e4a9
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.