Fractal basins trap latent reasoning Jeffrey Lai,1, ∗ Anthony Bao,2, ∗ John Quinn,1 and William Gilpin3, † 1 The Oden Institute, The University of Texas at Austin, Austin, Texas 78712, USA Department of Electrical Engineering, The University of Texas at Austin, Austin, Texas 78712, USA 3 Department of Physics, The University of Texas at Austin, Austin, Texas 78712, USA (Dated: September 7, 2026)
Reasoning allows artificial intelligence models to revisit and correct their mistakes, enabling recent frontier advances in mathematical theorem solving, software engineering, and autonomous task planning. Reasoning models are widely observed to reason for longer on harder tasks, but the general mechanism responsible for these slowdowns is unknown. Here, we show that reasoning models exhibit transient chaos, a physical consequence of the computational complexity of difficult tasks. As a consequence, we show that diverse leading reasoning models are dynamical systems with fractal basins, with fractality increasing with task difficulty across diverse tasks like Sudoku and maze solving, visual puzzles, and mathematical logic. We show that transient chaos emerges due to reasoning becoming trapped for extended durations near saddle points, which we show correspond to nearly-correct attempted solutions of the underlying problem. Our results show that reasoning slowdowns are an inevitable consequence of problem hardness in modern artificial intelligence models, and establish reasoning traces as a rich new class of dynamical system.
Here, we seek a physical understanding of reasoning using the framework of dynamical systems theory. Many reasoning models represent tasks as optimization problems: the problem statement and trained model weights parametrize a dynamical system, and the solution represents a stable fixed point [5, 6, 15, 16]. The dynamics either evolve explicitly in the form of intermediate outputs (as in chain-of-thought models) or in the form of a latent encoding of the problem and partial solutions as a hidden state. In this sense, modern reasoning models ex-
tend efforts to model task-solving using recurrent neural networks, which identify low-dimensional attractors that support computation [17–19]. A
Inject Problem
Decode Solution
B
16
Latent 2
Recent advances in frontier artificial intelligence are driven not by scale alone, but rather by the ability to reason. Reasoning models revisit and correct earlier mistakes before producing outputs, allowing them to achieve state-of-the-art performance on logical tasks such as visual reasoning, mathematical derivation, and puzzle solving [1–3]. Reasoning capabilities recently allowed small recurrent models (7 million parameters) to outperform large language models (exceeding 10 billion parameters) on the Abstract Reasoning Corpus (ARC-AGI), a set of complex visual riddles used to benchmark frontier models’ fluid intelligence and generalization ability [4–6]. Yet despite their efficiency, reasoning models are widely reported to suffer from overthinking, in which the models become trapped reasoning for extended periods without converging to a solution [7]. Overthinking can nearly double the inference cost of reasoning without improving its accuracy, and adversarially-chosen prompts can induce denial-of-service attacks on frontier reasoning models by consuming 10× more computing resources than nearly-identical benign prompts [8]. Overthinking is often attributed informally to the need to explore larger or more uncertain solution spaces [9–11]. Yet a mechanistic understanding of reasoning dynamics is currently missing, hindering empirical efforts to consistently identify and mitigate reasoning slowdowns [12–14].
8
Loop until converged Initial Latent State
Latent 1
Loops to Converge (L)
arXiv:2609.04963v1 [cs.LG] 4 Sep 2026
2
0
Figure 1. Recurrent depth reasoning models produce fractal basins. (A) A looped reasoning model encodes a task as a dynamical system evolving in a latent space, and the problem’s solution as a fixed point. (B) Convergence time basins on a difficult Sudoku puzzle not seen during training.
The latent state for modern reasoning models is usually initialized arbitrarily, such as by sampling a random vector. We thus introduce a new probe of reasoning dynamics by choosing a model and problem instance, and continuously varying the initial latent state to test how initialization affects convergence across replicates (Fig. 1A). Across diverse reasoning models, we observe complex basins of attraction when we label initial conditions by their convergence time, and we find that these basins become self-similar fractals on harder tasks, like Sudoku puzzles (Fig. 1B). As a consequence, even an infinitesimal change to the model’s initialization increases convergence time by orders of magnitude. Though reasoning models are trained to solve discrete puzzles, their fractal basins are smooth and naturalistic, resembling optical caustics or transient lifetimes in turbulence—both phenomena arising from spatial variation in the path-
2
B
16
Loops to Converge (L)
Sudoku
C Basin Entropy
A
Mazes
Loops to Converge (L)
Basin Entropy
0 73
0
2,3,5,7,9 76 9×7+5×3−2 = 76
Basin Entropy
Loops to Converge (L)
Mathematical Logic
55
Visual Puzzles
Loops to Converge (L)
Easy
Medium
Hard
0
Basin Entropy
0 10
Mean Loops to Converge
Figure 2. Hard problems produce fractal basins across diverse models and tasks. (A) Examples of Sudoku-Extreme, Maze-Hard, Countdown, and ARC-AGI-1 reasoning tasks, each of which we probe using a different architecture (Equilibrium Reasoners, Fixed Point Reasoning Models, Parcae, and Tiny Recursive Model) [4–6, 15, 16, 20] (B) Example basins for varying problem difficulties across the four tasks and architectures. (C) Scaling of basin entropy with the number of reasoning iterations required to converge, each across randomly-sampled task instances of varying underlying difficulty.
lengths required for a physical system to reach fixed points [21, 22]. In isolated previous cases, fractals have been observed when continuous dynamics are used to represent discrete optimization problems, such as sequence alignment, travelling salesperson routing, and spin glass ground state estimation [23–25]. In particular, a previous work derives a continuous-time dynamical system to solve Boolean satisfiability problems (which include Sudoku puzzles) and finds that fractal basins emerge as constraints increase [26, 27]. However, here we show that such hardness-induced fractal basins are ubiquitous in modern artificial reasoning models: even in the absence of solvers hand-designed for a particular task, reasoning models produce fractal basins whenever they encounter difficult problems, across highly-distinct model architectures and tasks (Figure 2A). We determine the prevalence of fractal basins using basin entropy, a recently-introduced metric that distin-
guishes self-similar fractal basins from random or smooth basins [28]. For each reasoning model and task, we randomly sample initial latent states along randomlyoriented two-dimensional slices through the full initial latent space. For each slice, we measure the average number of loops required for the model to converge across all initial conditions, and compare it to the basin entropy of the entire slice (Fig. 2C). We find that the basin entropy strongly correlates with the number of iterations, a measure of overall task difficulty. These findings persist across alternative architectures (Appendix C) and basin complexity metrics (Appendix D). Fractal basins form due to transient chaos, the phenomenon responsible for extended transients in highdimensional dynamical systems [29–31]. Successfully solving a hard problem requires that all trajectories of a reasoning model eventually converge to a single fixed point—and thus that no sustained chaos occurs. How-
3
A
Closely-spaced initial positions
B
D
1
3
1 Initial Puzzle
pca-2
λF
pca-1
Final solutions converge
2
C Residual
Weakly-unstable saddle (reasoning trap)
3 Solution
2 Nearly-correct trap
Iterations
Figure 3. Scattering from incorrect solutions produces reasoning transients. (A) The “Plinko” model of highdimensional dynamics. Trajectories that start and end at nearly the same point can take very different routes if one encounters weakly-unstable saddle points. The fast Lyapunov indicator (λF ) measures their maximum separation. (B) The latent state dynamics projected into two dimensions using PCA, showing two representative initial conditions that start out closely-spaced, but which take direct and indirect routes scattered by a saddle point. (C) The error between the latent state and converged latent state for trajectories originating from the initial conditions in (B). (D) Decoding the latent state near the saddle point produces a nearly-correct solution attempt.
We thus consider the underlying cause of saddle formation in reasoning tasks. We project the sequence of latent states encountered during high-dimensional reasoning traces into fewer dimensions using principal components analysis. We find that faster-converging trajectories take direct paths toward the true solution, while slower trajectories take longer, indirect routes that visit saddle points before reaching the true solution (Fig. 3B). Decoding these meandering trajectories into problem states reveals that reasoning becomes trapped near answers that are nearly, but not quite, correct. In mazes, saddle regions correspond to dead ends; in Sudoku puzzles, they represent grids with repeated digits (Fig. 3D). The amount of time that a dynamical system spends near a saddle depends on the number of outgoing escape directions [33]. In gradient systems, this corresponds to the number of downhill directions.
A
B Fast Lyapunov Indicator λF
ever, trajectories can take widely-varying routes to equilibrium, due to the presence of saddle points (weaklyunstable solutions) scattered throughout phase space, that redirect trajectories for extended durations and thus delay their convergence to equilibrium. High-dimensional dynamics thus resemble a “Plinko” toy: initially closelyspaced initial conditions scatter off of weakly-unstable saddle sets, which pull trajectories apart and create diverging routes to equilibrium (Fig. 3A) [32]. The scrambling of trajectories caused by saddles produces the complex structure of basins.
0.0
0.5
Fast Lyapunov Indicator λF
1.0
Frequency of Answer Changes
Figure 4. Saddle sets indicate uncertain reasoning trajectories. (A) The fast Lyapunov Indicator λF versus initial condition for a Sudoku-solving model, with two representative initial conditions highlighted (magenta and blue crosses). (B) λF versus the frequency at which reasoning traces switch solutions. We measure the frequency by decoding the latent reasoning dynamics to a candidate solution at each timestep, and measuring the number of changes divided by the total number of iterations before convergence. Each point corresponds to a different puzzle instance.
We directly measure the saddle sets responsible for transient chaos using the fast Lyapunov indicator λF , a quantity originally developed to measure the stability of asteroid orbits [34]. Two initially-close trajectories that take different routes to equilibrium have larger λF , implying that they straddle opposite sides of at least
4 Step 156642
Step 156644
Step 156645
Convergence Time
Latent Direction 2
Saddle Escape Directions
Latent Direction 2
B
Latent Direction 1
Latent Direction 1
C
Fraction Correct
Basin Entropy
Saddle Escape Directions
Training Step
D Lyapunov Exp. (FTLE)
Training Step 156641
Correct Incorrect
A
Direct Substitution Variables Multistep Elimination Variables
Training Step
Figure 5. Fractal basins emerge as bifurcations during training. (A) A looped transformer is trained to solve integer linear systems. Panels correspond to a series of training steps around the bifurcation that occurs 150k steps into training, when accurate solving first emerges. The top row colors initial conditions based on whether they converge to the correct solution, and the bottom row colors trajectories by their convergence time. (B) Latent reasoning trajectories before and after the bifurcation, projected into two latent components. Open circles denote starting points, and crosses denote final fixed points. (C) The fraction of correct solutions, basin entropy, and number of saddle escape directions along trajectories across the full training procedure. (D) The finite-time Lyapunov exponents separately-calculated for the directly-substitutable variables, and more difficult core variables that require multistep Gaussian elimination to solve.
one saddle. On single puzzles, λF reveals boundaries between different solution routes in reasoning traces (Fig. 4A). From an algorithmic perspective, these represent maximally-uncertain solution routes. Across many different puzzles for the three different models considered here, we find that the fast Lyapunov indicator strongly correlates with the number of distinct solutions that a given reasoning trajectory passes through before converging. Thus, saddle points shape the reasoning landscape by encoding the solution structure of the underlying problem (Fig. 4B). Generally, when high-dimensional dynamical systems transition from a complex phase to a monostable phase, their local minima disappear, leaving a single stable equilibrium but many weakly-unstable saddle points [35, 36]. In reasoning tasks, this transition occurs when an underconstrained puzzle with multiple solutions (such as a maze with two routes) loses a solution due to the addition of a new constraint (such as an additional wall blocking an exit route). The former solution still affects the convergence time of solvers, which may traverse all the way through the previous solution before reaching the barrier, requiring backtracking [37]. Generally, constraint-solving
problems are often hardest at this constrainedness transition: the problem is still solvable, but with few distinct routes to the correct solution—thus increasing the amount of required trial-and-error [38, 39]. We find that saddles encode these dead end reasoning routes in reasoning dynamical systems. To study the emergence of saddle-mediated fractal basins, we train a miniaturized reasoning model to solve integer linear systems: For prime p, given A ∈ ×N N FM , b ∈ FM p p , solve Ax = b for x ∈ Fp . There N are p candidate solutions corresponding to all possible N -dimensional integer-valued vectors over Fp = {0, 1, . . . , p − 1}. Typical discrete algorithms for this problem test sequences of integers until a constraint is found violated. We train an encoder-only looped transformer on many random instances with p = 3, M = N = 8, and at each stage of training we analyze the reasoning dynamics on newly sampled problem instances. In early stages of training, the model exhibits a basin landscape resembling that of a random dynamical system [36], with basins surrounding distinct stable fixed points, most of which represent incorrect solutions (Fig. 5A). However, in this regime, the convergence times do not show fractal
5 basins. As training proceeds, a bifurcation occurs, with the model’s weights acting as a control parameter: the incorrect solutions lose stability and become saddles, and reasoning trajectories begin converging to the fixed point of the true solution (Fig. 5A). In projections of the latent space, trajectories transition from smoothly converging to their nearest (incorrect) solution to taking more circuitous routes that pass near saddle points (Fig. 5B). As a result, at the bifurcation we observe an abrupt increase in solvability and basin entropy, and the appearance of unstable saddle directions, as measured by the number of unstable eigenvalues in the Jacobian of the reasoning model (Fig. 5C). These unstable directions confirm that local minima associated with incorrect solutions lose stability. However, saddles still weakly attract trajectories, producing basins in the convergence time field that resemble the former basins of incorrect solutions. Such saddle-mediated bifurcations produce fractal basins in many classical nonlinear systems [29]. We calculate the spectrum of finite-time Lyapunov exponents (FTLE), a set of rates that quantify how small perturbations to the latent state are initially amplified or suppressed along independent directions [40]. Pre-bifurcation, all FTLE are negative and so no transient chaos occurs: trajectories smoothly approach nearby local minima representing incorrect solutions. After the bifurcation, unstable directions appear, and transient chaos emerges due to scattering from saddle points. To interpret the unstable directions, we partition the N variables of the original problem into a subset of simple variables solvable with direct substitution, and core variables which require multi-step Gaussian elimination to resolve. This partitioning is analogous to constraint satisfaction problems with leaf variables (resolvable by constraint propagation) and the core (the sub-problem remaining after constraint propagation) [41]. Only the core variables produce positive FTLE, implying that the algorithm that the model uses to resolve these variables produces transient chaos (Fig. 5D). We thus find that the bifurcation to solvability during training occurs when the model gains multi-step reasoning capability. This, in turn, allows the model to escape from incorrect solutions in the core, but it leads to transient chaos and fractal basins. We have discovered that diverse reasoning models navigate fractal basins when they solve hard problems. Our observations hold across diverse model architectures and tasks, underscoring that fractality is a general phenomenon in learning models, arising from the algorithms reasoning models learn to apply during inference. Practically, our findings introduce a new probe of hidden reasoning processes, and show that models’ capabilities are related to their ability to escape saddle points near incorrect solutions. Our findings support a general interpretation of fractal basins and transient chaos as the generic manner in which computational complexity manifests in analogue systems [42], thus connecting the classi-
cal physics of nonlinear dynamics to the discrete symbolic reasoning tasks encountered by modern artificial intelligence models. Our results thus place modern reasoning models within a longstanding tradition framing computation as a physical process [43, 44]—supporting a view of complexity as a physical quantity, with consequences for inference as artificial intelligence methods continue to evolve.
ACKNOWLEDGMENTS
We thank Unconventional AI for supporting our group. This work was performed in part at the Aspen Center for Physics, which is supported by NSF grant PHY2210452. W.G. is supported by NSF DMS 2436233 and NSF CMMI 2440490, Grant No. DAF2023-329596 from the Chan Zuckerberg Initiative DAF (an advised fund of Silicon Valley Community Foundation), and grant CSCSA-2026-075 from Research Corporation for Science Advancement. Computational analyses were performed using the Biomedical Research Computing Facility at UT Austin, Center for Biomedical Research Support (RRID: SCR 021979) and the Texas Advanced Computing Center (TACC).
∗
These authors contributed equally to this work. [email protected] [1] J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al., Advances in neural information processing systems 35, 24824 (2022). [2] D. Guo, D. Yang, H. Zhang, J. Song, P. Wang, Q. Zhu, R. Xu, R. Zhang, S. Ma, X. Bi, et al., Nature 645, 633 (2025). [3] C. V. Snell, J. Lee, K. Xu, and A. Kumar, in The Thirteenth International Conference on Learning Representations (2025). [4] F. Chollet, arXiv preprint arXiv:1911.01547 (2019). [5] G. Wang, J. Li, Y. Sun, X. Chen, C. Liu, Y. Wu, M. Lu, S. Song, and Y. A. Yadkori, Hierarchical reasoning model (2025), arXiv:2506.21734 [cs.AI]. [6] A. Jolicoeur-Martineau, Less is more: Recursive reasoning with tiny networks (2025), arXiv:2510.04871 [cs.LG]. [7] X. Chen, J. Xu, T. Liang, Z. He, J. Pang, D. Yu, L. Song, Q. Liu, M. Zhou, Z. Zhang, et al., in Forty-second International Conference on Machine Learning (2025). [8] X. Liu, X. Wang, Y. Zhang, S. Kariyappa, C. Xiang, M. Chen, G. E. Suh, and C. Xiao, arXiv preprint arXiv:2602.00154 (2026). [9] A. Schwarzschild, E. Borgnia, A. Gupta, F. Huang, U. Vishkin, M. Goldblum, and T. Goldstein, Advances in Neural Information Processing Systems 34, 6695 (2021). [10] J. Geiping, S. M. McLeish, N. Jain, J. Kirchenbauer, S. Singh, B. R. Bartoldson, B. Kailkhura, A. Bhatele, and T. Goldstein, in The Thirty-ninth Annual Conference on Neural Information Processing Systems (2025). †
6 [11] X. Wang, J. Wei, D. Schuurmans, Q. V. Le, E. H. Chi, S. Narang, A. Chowdhery, and D. Zhou, in The Eleventh International Conference on Learning Representations (2023). [12] P. Shojaee, I. Mirzadeh, M. Horton, S. Bengio, M. Farajtabar, et al., Advances in Neural Information Processing Systems 38, 108018 (2026). [13] Center for AI Safety, Scale AI, and HLE Contributors Consortium, Nature 649, 1139 (2026). [14] W. Yang, S. Ma, Y. Lin, and F. Wei, Advances in Neural Information Processing Systems 38, 43605 (2026). [15] S. Movahedi, V. Milovanović, S. L. Feigin, A. Theus, T. Hofmann, V. Boeva, T. K. Rusch, and A. Orvieto, Fixed-point reasoners: Stable and adaptive deep looped transformers (2026), arXiv:2606.18206 [cs.AI]. [16] B. Huang, Z. Geng, and J. Z. Kolter, in Forty-third International Conference on Machine Learning (2026). [17] D. Sussillo and O. Barak, Neural computation 25, 626 (2013). [18] M. Khona and I. R. Fiete, Nature Reviews Neuroscience 23, 744 (2022). [19] D. Durstewitz, G. Koppe, and M. I. Thurm, Nature Reviews Neuroscience 24, 693 (2023). [20] H. Prairie, Z. Novack, T. Berg-Kirkpatrick, and D. Y. Fu, Parcae: Scaling laws for stable looped language models (2026), arXiv:2604.12946 [cs.LG]. [21] J. Moehlis, H. Faisst, and B. Eckhardt, New Journal of Physics 6, 56 (2004). [22] E. G. Altmann, J. S. E. Portela, and T. Tél, Rev. Mod. Phys. 85, 869 (2013). [23] J. J. Hopfield and D. W. Tank, Biological cybernetics 52, 141 (1985). [24] L. Chen and K. Aihara, Neural networks 8, 915 (1995). [25] V. Elser, I. Rankenburg, and P. Thibault, Proceedings of the National Academy of Sciences 104, 418 (2007). [26] M. Ercsey-Ravasz and Z. Toroczkai, Nature Physics 7, 966 (2011). [27] M. Ercsey-Ravasz and Z. Toroczkai, Scientific reports 2, 725 (2012). [28] A. Daza, A. Wagemakers, B. Georgeot, D. Guéry-Odelin, and M. A. Sanjuán, Scientific reports 6, 31416 (2016). [29] C. Grebogi, E. Ott, and J. A. Yorke, Physical Review Letters 50, 935 (1983). [30] T. Tél and Y.-C. Lai, Physics Reports 460, 245 (2008). [31] A. E. Motter, M. Gruiz, G. Károlyi, and T. Tél, Physical review letters 111, 194101 (2013). [32] Y. N. Dauphin, R. Pascanu, C. Gulcehre, K. Cho, S. Ganguli, and Y. Bengio, Advances in neural information processing systems 27 (2014). [33] H. Kantz and P. Grassberger, Physica D: Nonlinear Phenomena 17, 75 (1985). [34] C. Froeschlé, R. Gonczi, and E. Lega, Planetary and space science 45, 881 (1997). [35] G. Ben Arous, Y. V. Fyodorov, and B. A. Khoruzhenko, Proceedings of the National Academy of Sciences 118, e2023719118 (2021). [36] G. Wainrib and J. Touboul, Physical review letters 110, 118101 (2013). [37] R. Dechter and I. Meiri, Artificial Intelligence 68, 211 (1994). [38] P. C. Cheeseman, B. Kanefsky, W. M. Taylor, et al., in Ijcai, Vol. 91 (1991) pp. 331–337. [39] R. Monasson, R. Zecchina, S. Kirkpatrick, B. Selman, and L. Troyansky, Nature 400, 133 (1999).
[40] Y.-C. Lai and T. Tél, Transient chaos: complex dynamics on finite time scales (Springer Science & Business Media, 2011). [41] A. Braunstein, M. Leone, F. Ricci-Tersenghi, and R. Zecchina, Journal of Physics A: Mathematical and General 35, 7559 (2002). [42] C. Moore, Physical Review Letters 64, 2354 (1990). [43] R. Shaw, Zeitschrift für Naturforschung A 36, 80 (1981). [44] J. P. Crutchfield and K. Young, Physical Review Letters 63, 105 (1989).
1 CONTENTS
Acknowledgments
5
References
5
Appendix A. Code Availability
1
Appendix B. Supplementary Methods
1
Appendix C. Replicate basin experiments with alternative reasoning models
3
Appendix D. Metrics for basin complexity
3
Appendix E. Measuring fractality across scales
3
References
11
Appendix Appendix A: Code Availability
Our basin probing code is available at: https://github.com/GilpinLab/loopscape
Appendix Appendix B: Supplementary Methods
All experiments probe publicly-released recurrent reasoning models for which the model’s latent state and its decoded output can be extracted at every reasoning loop. In all cases the reasoning model and problem instance jointly define an autonomous discrete-time dynamical system acting on the latent state, and we decoded the latent state after each loop as an intermediate representation of the solution. Equilibrium Reasoners (EqR) [1] iterate paired latents (zH , zL ) over a sequence of 97 tokens (for Sudoku, for which the hidden dimension is 512) or 916 tokens (for mazes, for which the hidden dimension is 128). We use the authors’ released checkpoints trained on Sudoku-Extreme and Maze-Unique. Because EqR is trained with stochastic inference, we disable its injected noise and fix the loop budget at 24 iterations, and so the inference dynamics correspond to a deterministic map of the initial latent state. Fixed-Point Reasoning Models (FPRM) [2] iterate a single latent state toward a fixed point, halting once the relative residual falls below a fixed threshold τ . We use the final released checkpoints and evaluation setting. For Sudoku, we set τ = 0.02 with a 1000-iteration maximum budget. For mazes we set τ = 0.1 with a 100-iteration cap. Parcae [3] is a 140M parameter looped language model that we finetuned on the Countdown arithmetic task. We treat early decoding of text describing the solution as the puzzle output. The original model released by the authors was pretrained with an effective timescale of ∆t = 1.0 associated with each latent iteration. We fine-tune this model with a smaller timestep of ∆t = 0.3. Basin maps of initial-condition space. For each model and problem instance, we hold the model, prompt, and loop schedule fixed and vary only the initial latent state along a random two-dimensional slice of the full latent space. We sample two orthonormal directions from the flattened latent space (e.g. 97 × 512 dimensions for Sudoku), using QR decomposition of a random Gaussian matrix initialized in the space. We initialize the reasoning model at z(a, b) = z0 + au + bv on a uniform grid (a, b) ∈ [−h, h]2 where h is a half-width chosen for each model, and we use a resolution of 200 × 200 when computing population-level statistics. The base point z0 is a fixed random Gaussian sample in the space, and we scale (a, b) based on the scale of each reasoning model’s typical latent dynamics. For FPRM mazes, the random slice is confined to the fifteen “register” tokens that receive no input injection. For every initial condition, we record the decoded output (solution state) at each loop. We define the convergence (settling) time as the number of loops until the decoded output of a given trajectory stops changing. Initial conditions still changing at the loop cap are marked as errors. The resulting settling time field over the two-dimensional slice constitutes the basin map. Task instances and difficulty. Sudoku instances are drawn uniformly from the Sudoku-Extreme test set, using the benchmark’s annotated per-puzzle difficulty ratings, which are based on backtracking encountered by a classical solver [4]. FPRM maze instances are drawn from the Maze-Unique and Maze-Hard 30 × 30 test splits, with difficulty
2
B
16
Loops to Converge (L)
Mazes
C Basin Entropy
A
0
Mean Loops to Converge
2,3,5,7,9 76 9×7+5×3−2 = 76
Easy
Medium
Hard
0
Basin Entropy
Loops to Converge (L)
Mathematical Logic
40
Mean Loops to Converge
Figure S1. Hard problems produce fractal basins across alternative models. (A) Examples of Maze-Hard and Countdown reasoning tasks, each of which we probe using a different architecture (Equilibrium Reasoners and a looped transformer) [1, 4] (B) Example basins for varying problem difficulties across the two tasks and architectures. (C) Scaling of basin entropy with the number of reasoning iterations required to converge, each across randomly-sampled task instances of varying underlying difficulty.
annotated by shortest-path length. For EqR maze experiments we use a set of 300 generated perfect mazes, which correspond to random depth-first-search spanning trees on a 30 × 30 board. The wall fraction of these generated mazes is ≈ 0.57, and the solution path lengths vary between 101–169 steps. A maze with a unique shortest path can optionally be degenerated by opening k pairwise-independent corner walls, each exactly doubling the number of optimal routes (2k shortest paths). Each sweep covers 100–300 instances per random seed, with independent slice orientations across seeds. Quantifying basin fractality. For each sampled two-dimensional slice of settling times we compute the basin P entropy [5]: we partition the slice into non-overlapping boxes, and compute the Shannon entropy − i pi ln pi of the settling times within each box. The basin entropy Sb is the mean of this quantity over all boxes. A related quantity, the boundary entropy Sbb , averages only boxes containing more than one value. We also estimate the uncertainty exponent [6] from the fraction f (ε) of pixel pairs at separation ε = 1, . . . , 5 pixels whose settling times differ. The basinboundary dimension is 2 − α. Sampled slices are excluded from population statistics when fewer than 90% of initial conditions reach the model’s final solution (corresponding to cases in which the model had multistable or inconsistent solves). We also exclude slices where more than 1% of pixels reach the maximum number of loops (corresponding to cases where simulating the dynamics to convergence is not computationally feasible). This produces at least 300 valid slices for problems within each pair of reasoning models and tasks, from which we calculate correlations between the mean settling time across a slice, and Sb , Sbb , and α. Fast Lyapunov Indicator. We localize the boundaries separating distinct routes to equilibrium by calculating the fast Lyapunov indicator [7] directly on the solution trajectories associated with two-dimensional slices of initial conditions. At each loop, we compare every pixel’s decoded state under the dynamics at that time with the decoded state of the pixel’s initial nearest neighbors. The maximum Euclidean norm of this time-dependent local deviation estimates λF . Large λF therefore marks initial conditions with trajectories that separate maximally from their immediate neighbors before reconverging at the solution, thus corresponding to points straddling a saddle-like boundary between solution routes. Trajectory geometry and decoding near saddles. To study the geometry of trajectories, we analyze the full latent dynamics of a set of initial conditions (and not just the dynamics in the solution space, as in previous analyses). At each iteration, we compute the residual between the latent timepoint and the final latent solution. For a given puzzle, we perform principal components analysis on all latent trajectories, pooled across both initial conditions and timepoints.
3 Appendix Appendix C: Replicate basin experiments with alternative reasoning models
In order to rule out fractal basins only arising from certain task-model pairings, we replicate our findings using alternative architectures to solve the Maze and Countdown tasks (Fig. S1). We consider an Equilibrium Reasoning model checkpoint trained by the original authors to solve maze tasks, and we train an encoder-only looped transformer to solve Countdown [1]. We find that fractal basins emerge on difficult problems, and we observe a relationship between fractality and task difficulty. As we observed with other models, the qualitative appearance and structure of the basins change, likely because they are a property of the particular optimization algorithm that each model uses to solve a task. Appendix Appendix D: Metrics for basin complexity
Mean Loops to Converge
Uncertainty Exponent
TRM ARC-AGI
Mean Loops to Converge
Mean Loops to Converge
Mean Loops to Converge
Mean Loops to Converge
Mean Loops to Converge
Mean Loops to Converge
Mean Loops to Converge
Parcae Countdown Uncertainty Exponent
Basin Entropy
Basin Entropy
Uncertainty Exponent
Boundary Basin Entropy
Encoder-only Countdown
Mean Loops to Converge
Mean Loops to Converge
Boundary Basin Entropy
Basin Entropy Mean Loops to Converge
Boundary Basin Entropy
Mean Loops to Converge
Mean Loops to Converge
Uncertainty Exponent
Uncertainty Exponent
FPRM Maze Basin Entropy
Boundary Basin Entropy
Mean Loops to Converge
Basin Entropy
Mean Loops to Converge
EqR Sudoku
Boundary Basin Entropy
Basin Entropy
Uncertainty Exponent
EqR Maze
Boundary Basin Entropy
Across all tasks and architectures, we compute three measures of basin complexity (Fig. S2). The basin entropy measures the average variation in the convergence time within finite-resolution neighborhoods of different initial conditions, while the boundary basin entropy computes this quantity only over a subset of initial conditions straddling multiple basins [5]. We also compute the uncertainty exponent, which measures how the fraction f (ε) of initial conditions whose convergence time changes under a perturbation of size ε scales as f (ε) ∼ εα [8]. Equivalently, α = D − db , where D is the dimension of the sampled initial-condition space (D = 2 in our experiments) and db is the boundary’s box-counting dimension. Due to this relationship, we broadly observe that the different metrics are all strongly correlated across different puzzle instances and tasks.
Mean Loops to Converge
Mean Loops to Converge
Mean Loops to Converge
Figure S2. Alternative metrics of basin complexity. We repeat the basin quantification procedure using three alternative metrics. Each point corresponds to metrics calculated over a random two-dimensional slice through a random problem instance.
Appendix Appendix E: Measuring fractality across scales
We show the fractal structure of different reasoning models and tasks by repeatedly zooming into regions of the space of initial conditions. We show plots of the convergence time and the fast Lyapunov indicator at each scale,
4 across different model architectures [7]. In order to quantify the scale dependence of fractality, for each model and task we measure the uncertainty exponent [8], which measures how variation in outcomes within a fixed-radius-ϵ ball of initial conditions varies with ϵ. For true fractals, uncertainty and radius should exhibit a fixed power-law relationship, supporting the standard view of fractals as scale-free objects [9]. We find that some reasoning models produce true fractals, and thus have similar degrees of qualitative complexity when one zooms in. Other reasoning models, however, produce recently-characterized “slim fractals,” which correspond to the fractal dimension decreasing with the level of resolution [10]. Slim fractals are scale-dependent, and the true basin boundary has codimension 1. They typically arise in physical systems with dissipation, where the saddle-like set that normally scatters trajectories (and causes transient chaos) decays over time [11, 12]. For example, a double pendulum has a sustained chaotic attractor in the absence of friction. However, when friction is introduced, energy gradually leaves the system, and the only long-term attractor is a fixed point (in which the pendulum is at rest) [13]. At any given instant (and energy level) the pendulum has a saddle-like set, but this set decays over time. This general phenomenon, doubly-transient chaos, differs from singly-transient chaos (such as scattering fields in the three-body problem) in which the saddle-like set is constant, and transience arises purely from trajectories approaching non-chaotic steady states at long times [14, 15]. Reasoning models are not typically designed to directly map onto physical systems, and so we generally expect that their dynamics can produce types of basins and transient chaos beyond singly- or doubly-transient. However, many properties built into these models, such as a tendency to converge to fixed answers, likely represent a form of dissipation, and so we expect that most reasoning models will exhibit some form of doubly transient chaos.
5
10x
100x
1000x
10x
100x
1000x
Figure S3. Progressive magnifications of a fractal produced by the Equilibrium Reasoning model on a difficult Sudoku task. Top Left: settle-time field (loops to converge); Top Right: fast Lyapunov indicator (FLI) on the settle-time field. Middle Row: zoom sequence on the settle-time field. Bottom Row: zoom sequence on the FLI field.
6
f(ε) = Fraction of Pairs with Different Labels
100
α=
10−1
α=
α=
6 0.21 3 0.24 0.24
9
10−2
EqR Sudoku α = 0.249 ± 6e-3 10x zoom α = 0.243 ± 9e-3 100x zoom α = 0.216 ± 4e-3
10−14
10−13
10−12
10−11 10−10 Separation ε
10−9
10−8
Figure S4. Scaling of the uncertainty exponent for the fractal in Fig. S3.
10−7
7
10x
100x
Figure S5. Progressive magnification of a fractal produced by FPRM on the Maze task.
8 100
f(ε) = Fraction of Pairs with Different Labels
10−1
α
=
85 0.
3
10−2
10−3
10−4
f(0) floor
< 8.0e-05
FPRM Maze α = 0.853 ± 3e-2
10−5
10−4
10−3 Separation ε
10−2
Figure S6. Scaling of the uncertainty exponent for the fractal in Fig. S5.
9
10x
100x
1000x
10x
100x
1000x
Figure S7. Progressive magnification of a fractal produced by the fine-tuned Parcae model on the Countdown mathematical logic task.
10
f(ε) = Fraction of Pairs with Different Labels
100
10−1
10−2 α
=
88 0.
7
α
=
1 86 0.
α
=
84 0.
6
α
=
91 0.
2
10−3
10−4
f(0) floor
< 2.0e-05
1x α = 0.912 ± 8e-3 10x α = 0.846 ± 6e-3 100x α = 0.861 ± 3e-2 1000x α = 0.887 ± 2e-2
10−7
10−6
10−5
10−4
10−3 Separation ε
10−2
10−1
100
Figure S8. Scaling of the uncertainty exponent for the fractal in Fig. S7.
101
11
∗
These authors contributed equally to this work. [email protected] [1] B. Huang, Z. Geng, and J. Z. Kolter, in Forty-third International Conference on Machine Learning (2026). [2] S. Movahedi, V. Milovanović, S. L. Feigin, A. Theus, T. Hofmann, V. Boeva, T. K. Rusch, and A. Orvieto, Fixed-point reasoners: Stable and adaptive deep looped transformers (2026), arXiv:2606.18206 [cs.AI]. [3] H. Prairie, Z. Novack, T. Berg-Kirkpatrick, and D. Y. Fu, Parcae: Scaling laws for stable looped language models (2026), arXiv:2604.12946 [cs.LG]. [4] G. Wang, J. Li, Y. Sun, X. Chen, C. Liu, Y. Wu, M. Lu, S. Song, and Y. A. Yadkori, Hierarchical reasoning model (2025), arXiv:2506.21734 [cs.AI]. [5] A. Daza, A. Wagemakers, B. Georgeot, D. Guéry-Odelin, and M. A. Sanjuán, Scientific reports 6, 31416 (2016). [6] S. W. McDonald, C. Grebogi, E. Ott, and J. A. Yorke, Physica D: Nonlinear Phenomena 17, 125 (1985). [7] C. Froeschlé, R. Gonczi, and E. Lega, Planetary and space science 45, 881 (1997). [8] C. Grebogi, S. W. McDonald, E. Ott, and J. A. Yorke, Physics Letters A 99, 415 (1983). [9] B. Mandelbrot, science 156, 636 (1967). [10] X. Chen, T. Nishikawa, and A. E. Motter, Physical Review X 7, 021040 (2017). [11] A. E. Motter, M. Gruiz, G. Károlyi, and T. Tél, Physical review letters 111, 194101 (2013). [12] O. E. Omel’chenko and T. Tel, Journal of Physics: Complexity 3, 010201 (2022). [13] G. Károlyi and T. Tél, Journal of Physics: Complexity 2, 035001 (2021). [14] N. C. Stone and N. W. Leigh, Nature 576, 406 (2019). [15] S. Wiggins, Normally hyperbolic invariant manifolds in dynamical systems, Vol. 105 (Springer Science & Business Media, 1994). †