Transformers as In-Context Samplers: From Closed-Form Diffusion to Estimation-Free Sampling
Arman Adibi1
arXiv:2609.08981v1 [cs.LG] 8 Sep 2026
1
Alireza Jafari2 Mohammad Ghavamzadeh3 Hadi Daneshmand2 School of Computer and Cyber Sciences, Augusta University 2 Department of Computer Science, University of Virginia 3 Qualcomm AI Research [email protected]
Abstract A growing body of work establishes that large language models are not mere statistical memorizers, but are capable of in-context learning: performing inference at test time using only examples provided in the prompt, without any parameter updates. Prior theoretical work has shown that this capability extends to supervised learning tasks such as linear regression. We prove that in-context learning extends further to data generation: frozen transformers can simulate iterative generative samplers from in-context samples. We first show that transformers can realize closed-form and smoothed closed-form diffusion samplers. The construction identifies a concrete generative role for softmax attention: it computes responsibility weights and weighted empirical averages, while feedforward layers implement Euler updates. To empirically relate these constructions to pretrained language models, we study semantic-topic sampling: prompts consisting of words drawn from a common semantic category, such as animals, foods, or cities. Across transformer layers, the normalized hidden states exhibit a two-stage geometry: they move toward a uniform spherical reference in intermediate layers and then return to structured, topic-dependent representations near the output. We further measure an interactingparticle energy on these hidden-state clouds and observe the same U-shaped pattern. We then prove that transformers can approximate an energy-based sampler, constructing the same U-shaped energy across the layers.
1
Introduction
"In-context learning" [1] has changed how we think about transformers. A transformer can adapt its behavior at inference time using only information provided in a prompt, without updating its parameters. This phenomenon was first made prominent by large language models, where a prompt containing a few demonstrations can induce the model to perform a new task [1–5]. It has since become a central paradigm for studying how models use demonstrations, instructions, and contextual information at inference time [6]. Existing studies have mostly focused on conditional data generation, specifically for supervised learning: the prompt contains examples from an unknown task, and the model predicts the output for a new query. A standard formalization studies prompts of the form (x1 , f (x1 ), . . . , xk , f (xk ), xquery ) | {z } prompt
7−→
f (xquery ) . | {z } completion
For example, Garg et al. [7] ask whether transformers can learn simple function classes in context, such as linear functions, sparse linear functions, decision trees, and shallow neural networks. Subsequent Preprint.
work studies whether transformers implement particular learning algorithms in context, including ridge regression, gradient descent, higher-order optimization, and algorithm selection [8–11]. These results suggest that transformers can do more than match surface patterns in the prompt: they can execute nontrivial computations inside their forward pass. Transformers were originally proposed for conditional text generation [12], but have since become the backbone of powerful unconditional generative models, including latent diffusion models [13]. While recent studies have extensively investigated the mechanistic interpretation of in-context learning in conditional text generation, it remains an open question whether such capabilities extend to unconditional data generation and broader generative modeling applications. In this paper, we ask whether in-context learning extends to data generation: Can a transformer implement a sampler in context? Given i.i.d. samples x1 , . . . , xn ∈ Rd from an unknown distribution P, a sampling algorithm aims to generate a fresh sample xn+1 ̸= x1 , . . . , xn from P. Akin to the standard settings of generative models, we assume that the density function of the data distribution is unknown, and we only have access to i.i.d. samples from the target distribution. Inspired by prior in-context learning studies [14, 7], we encode the sampling problem as input:
x1 , . . . , xn ∼i.i.d. P
7−→
output:
xn+1 ∼ P,
where x1 , . . . , xn are encoded in the word embeddings of a large language model and xn+1 is an output embedding vector. We refer to this prompting strategy as in-context sampling: the transformer parameters remain fixed while the distribution of P changes at test-time. A model can leverage two sources of information for inference: statistics from training data and information from in-context samples. In-context learning emerges when the model relies on the latter, adapting its response to the in-context input at test time. Figure 1 illustrates the phenomenon that motivates our study. The model is trained on samples from a face distribution containing the face boundary and eyes, but no smiling mouth, where samples x1 , . . . , xn are encoded in word embeddings. When the in-context samples come from components seen during training, such as the boundary or the eyes, the generated samples follow the corresponding prompt distribution. The key test is the smiling mouth: although this component is absent from training, providing mouth-shaped samples in context causes the transformer to generate new samples along the same unseen curve. Thus, the model is not simply reproducing the global training distribution or selecting among memorized components. Rather, the inference at test time is based on in-context samples. Our goal is to explain the mechanism underlying this capability. Training samples
(a)
In-context samples
(b)
Predicted IID tokens
(c)
(d)
Figure 1: In-context sampling. We train a transformer to generate points from the blue training distribution in panel (a), which contains the face boundary and eyes but excludes the smiling mouth. At test time, the model receives in-context samples shown in orange and generates new samples shown in green. When the context samples are drawn from components present during training, the generated samples follow the face boundary in panel (b) and the eyes in panel (c). Panel (d) gives the out-of-distribution test: the context samples form a smiling mouth, a component entirely absent from training. The model nevertheless generates new samples along the same unseen curve. This indicates that the transformer is using in-context samples for inference, rather than relying on statistics of the training distribution. See the Experiments section for training and data generation details. Contributions We make the following contributions to demonstrate the generative power of transformers. In-context simulation of diffusion models. We prove that transformers can simulate closed-form diffusion models [15] from in-context samples. Closed-form diffusion models leverage the explicit 2
score function to simulate the backward diffusion process. We show that softmax attention heads are particularly well-suited to compute the empirical score function directly from in-context samples. Mechanistic analysis of pretrained LLMs. We examine the data generation mechanism in pretrained language models, including GPT-2 [16], Llama-3.3-70B-Instruct [17], OpenLM-Llama13B/OpenLLaMA [18], Cerebras-GPT-13B [19], Qwen2.5 [20], Falcon-40B [21], OPT-66B [22], and BLOOM-7B1 [23]. The models are prompted by words drawn from the same semantic topic. This prompt construction mirrors the i.i.d. assumption in the in-context sampling setting. Tracking the empirical embedding distributions across layers, we observe that the distribution of normalized embeddings progressively approaches a uniform distribution over a sphere in the middle layers, before becoming non-uniform again towards the output layers. The same U-shaped pattern is observed for the interacting energy of word embedding. Estimation-free sampling. We identify a sampler whose layer-wise energy profile qualitatively matches the observed U-shaped. We prove that transformers can implement estimation-free sampling (EFS), an energy-based generative framework comprising two steps: (1) transporting the empirical data distribution toward a uniform distribution, and (2) inverting this process to push the empirical distribution toward a target distribution. This two-step mechanism resembles the internal distribution shift observed in pretrained transformers. Connections to prior work. Our work builds on the algorithmic view of in-context learning, where transformers are studied as models that can execute non-trivial computations from prompts rather than only predict labels [7, 8, 14, 9, 10]. It is also related to diffusion and score-based generative modeling [24–26], as well as Estimation-Free Sampling [27]. The closest work connecting diffusion and in-context learning is Prompt Diffusion [28], which trains a diffusion model to perform visual tasks from in-context examples. Our direction is different: we ask whether a frozen transformer can itself implement diffusion-style and particle-based sampling algorithms from the prompt. A more detailed discussion of related work is deferred to Appendix A.
2
Preliminaries
2.1
Closed-form diffusion process
Score-based diffusion models [25] are among the most powerful generative models for transporting a simple reference distribution, such as a Gaussian distribution, to a target data distribution. In these models, the transport dynamics are driven by the score function, i.e., the gradient of the log-density of a smoothed data distribution. In practice, this score is typically approximated by a neural network trained through denoising. In continuous time, the reverse-time generative dynamics can be described through Anderson’s reverse-time diffusion framework [29]. In an abstract form, the reverse time process is written as dzt = (−zt − ∇z log pt (zt )) dt + dWt ,
(1)
where pt denotes the density of the target distribution after Gaussian smoothing at time t, and Wt is a Wiener process. The term ∇z log pt (zt ) drives the sample toward regions of high probability under the smoothed data distribution. 1 The Gaussian smoothing and stochastic noise injection play an annealing role: at large noise levels, the density is smoother and easier to explore, while at smaller noise levels, the dynamics refine samples toward the data distribution. Closed-form diffusion models use the observation that, for a finite empirical distribution, the Gaussiand smoothed density can be written explicitly as a Gaussian mixture [15]. Let {xi }N i=1 ⊂ R be the empirical dataset. For t ∈ (0, 1), define ρ⋆t (z) =
N 1 X ϕ z; txi , (1 − t)2 Id , N i=1
z ∈ Rd ,
where ϕ(·; µ, Σ) is the Gaussian density with mean µ and covariance Σ. Thus, ρ⋆t is a mixture of isotropic Gaussians centered at the scaled samples txi , with covariance (1 − t)2 Id . 1 The negative sign of ∇ log p (z ) arises because the process runs in reverse time, from T back to 0. t t
3
For a current point z, define the responsibility weight of sample xi by ϕ z; txi , (1 − t)2 Id , i = 1, . . . , N. (2) wi (t, z) = PN 2 j=1 ϕ(z; txj , (1 − t) Id ) PN These weights satisfy wi (t, z) ≥ 0 and i=1 wi (t, z) = 1. They measure how much the Gaussian component centered at txi contributes to the density at z. Define the corresponding weighted mean PN kt (z) = i=1 wi (t, z)(txi ). Since the distribution of p∗t (z) is a Gaussian mixture with explicit density function, its score can be computed in closed-form as [15] 1 kt (z) − z . (3) ∇z log ρ⋆t (z) = 2 (1 − t) Thus, the score tells the sampler how to move z: it points from the current point z toward the weighted average kt (z) of the data samples. The sampler converts this score into an update direction, which we write it as vt (z) =
1 (z + (1 − t)∇z log ρ⋆t (z)) . t
(4)
Here vt (z) ∈ Rd is the direction in which the sampler moves the current point z at time t. Substituting (3) gives 1 1 1 1 vt (z) = z+ (kt (z) − z) = − z+ kt (z). t 1−t 1−t t(1 − t) Therefore, the main data-dependent computation is the weighted mean kt (z). Once kt (z) is known, the update direction vt (z) is a fixed linear combination of z and kt (z). Given a time grid τ = {(ts , hs )}S−1 s=0 , where 0 < ts < 1 and hs > 0, the closed-form diffusion sampler relies on the following recurrence for sampling: zs+1 = zs + hs vts (zs ),
s = 0, . . . , S − 1.
(5)
This explicit form is central to our construction. The sampler does not require a trained score network: the score and update direction are computed directly from the in-context samples x1 , . . . , xN . In later sections, we show that a frozen transformer can implement this computation using attention to compute the weights wi (t, z) and the weighted mean kt (z), and using feedforward layers to perform the Euler update. Smoothing. Since the closed-form diffusion model may memorize data, [15] proposes smoothing to avoid data memorization. Smoothing is obtained by averaging the closed-form mean over fixed d perturbations of the current state. Let σ ≥ 0 and E = {ϵm }M m=1 ⊂ R be fixed perturbation vectors. PM −1 Define kσ,t (z) = M m=1 kt (z + σϵm ), where kt (·) is the closed-form responsibility-weighted mean defined above. The smoothed velocity field is vσ,t (z) = −
1 1 z+ kσ,t (z). 1−t t(1 − t)
(6)
Thus, the smoothed sampler replaces kt (z) by the averaged mean kσ,t (z). Setting σ = 0 recovers the closed-form diffusion method. [15] has examined the data generation capability of the above method in comparison to original denoising diffusion models. This modification preserves the same transformer-realization structure: attention computes kt (z + σϵm ) for each fixed perturbation, the heads are averaged to obtain kσ,t (z), and the feedforward layer applies the affine update (6). 2.2
Embedding encoding
We encode a closed-form diffusion instance in the input word embedding matrix of a transformer, denoted by (N +1)×p X = Enc({xi }N , p = 3d + 2, i=1 , z0 ) ∈ R 4
where z0 ∈ Rd is a Gaussian random vector which is equivalent to starting state of the closed-form diffusion model defined in (6). The encoded embedding consists of N data tokens and one state token. We write X ∈ R(N +1)×p , with p = 3d + 2, and define its rows by 2 ⊤ ⊤ ⊤ ⊤ Xi = x⊤ i = 1, . . . , N, XN +1 = 0⊤ i , ∥xi ∥ , 0d , 0d , 1 , d , 0, z0 , 0d , 1 . Here, the first N rows store the in-context samples, while the final row stores the initial sampler (x) (r) (z) (k) (1) state z0 . Equivalently, each row is partitioned as Xi = Xi | Xi | Xi | Xi | Xi , where (x) (z) (k) (r) (1) Xi , Xi , Xi ∈ Rd , Xi ∈ R, and Xi ∈ R. These four blocks serve distinct purposes in the design: the x and r blocks encode training samples for in-context learning; the z block acts as memory for newly generated samples; and the k block functions as a scratchpad for intermediate computation, inspired by prior work [30]. The state readout map Πstate : Rm×p → Rd extracts (z) the z-block of the final row: Πstate (X) = XN +1 = XN +1, d+2: 2d+1 . We will later use the above notation to reconstruct the output of a transformer. 2.3
Transformer
We use a standard transformer architecture with softmax self-attention [12]. Let Tθ : Rm×p → Rm×p denote a depth-L transformer, where m is the number of tokens, p is the token dimension, L ∈ N is the number of layers, and θ denotes the collection of attention and feedforward weight matrices. Thus, the transformer maps an input embedding matrix to an output embedding matrix of the same size. We generate a sample by applying the state readout map Πstate : Rm×p → Rd , defined in the N previous section, to the final embedding matrix. For a sampling instance {xi }i=1 with initial state z0 , N the transformer output is zbL = Πstate Tθ Enc({xi }i=1 , z0 ) . Our goal is to construct parameter choices θ such that zbL matches the output of a specified iterative generative sampler.
3
Simulation of Closed-form Diffusion with Transformers
We now show that the transformer architecture in Section 2.3 can simulate the (smoothed) closed-form diffusion sampler by a suitable choice of its parameters. The prompt contains the empirical samples d d {xi }N i=1 ⊂ R and the initial state z0 ∈ R . Theorem 1 (Transformer simulation of closed-form diffusion). There exists a choice of transformer ⋆ d d parameters θCFD such that, for every empirical dataset {xi }N i=1 ⊂ R and every initial state z0 ∈ R , if zL is generated by the closed-form diffusion recursion, then ⋆ Πstate TθCFD Enc({xi }N = zL , i=1 , z0 ) where L is the number of attention blocks. Moreover, fix σ ≥ 0 and E = {ϵm }M m=1 . There exists ⋆ another choice of transformer parameters θσ-CFD , using the same prompt encoding and state readout, such that, for every empirical dataset and initial state, if zL is generated by the smoothed recursion, then ⋆ Πstate Tθσ-CFD Enc({xi }N = zL . i=1 , z0 ) The parameter choices may depend on N, d, τ , and, in the smoothed case, on σ and E, but they do not depend on the realized dataset {xi }N i=1 or on z0 . The theorem states that the standard transformer architecture can be parameterized to execute the closed-form diffusion update rule in context. Notably, the results hold for all possible inputs while the parameters are frozen, showing that the transformer is capable of out-of-distribution generalization. We observe such generalization in Figure 1, where the transformer generates samples from a distribution distinct from the training data. In addition to this out-of-distribution generation, the theorem establishes uniform in-context generalization over the data values: for any fixed context length N , dimension d, and sampling schedule, the same transformer parameters work for every realized dataset {xi }N i=1 and every initial state z0 . The proof of Theorem 1 builds on the well-established computational capabilities of softmax attention layers [12], demonstrating that they can implement the iterative closed-form diffusion method. Prior work on in-context learning has similarly shown that transformers can implement various iterative 5
algorithms, including gradient descent for regression [31, 14] and temporal difference learning for reinforcement learning [32]. However, these studies rely on linear attention layers to simplify the analysis. What distinguishes our work from these prior studies is the use of standard softmax attention in place of linear attention [14]. In fact, the softmax normalization is essential to our construction: the closed-form diffusion update requires responsibility weights, which are precisely the normalized weights produced by softmax attention. Such normalization is omitted in linear attention due to the challenges it poses for theoretical analysis. Remark 1 (Relation to Rosu et al. [33]). Rosu et al. [33] show that transformers can implement in-context denoising steps using a modified attention operator, denoted Attn, which incorporates an RBF-type modification. In contrast, our construction uses the standard softmax self-attention mechanism employed by transformers in practice. The norm-dependent quantities required to compute the Gaussian-mixture responsibilities are included explicitly in our input encoding, rather than recovered through a modification of the attention mechanism. Thus, our closed-form diffusion result should be viewed as a warm-up showing that the original transformer architecture, without the RBF modification of [33], can also implement closed-form diffusion models from in-context samples. Together, these results show that transformers are expressive enough to implement several sampling algorithms. Expressivity alone, however, does not determine which algorithm best explains the computations learned by a trained transformer. Our main objective is to identify a generative mechanism that captures the layer-wise evolution of representations in pretrained language models. Closed-form diffusion does not reproduce the observed U-shaped dynamics, in which the empirical embedding distribution first approaches a uniform distribution and subsequently becomes nonuniform. Moreover, the unsmoothed empirical closed-form diffusion model targets the empirical distribution and can therefore regenerate in-context samples. These observations motivate our analysis of Estimation-Free Sampling (EFS), whose forward and backward dynamics closely resemble the observed U-shaped mechanism. Under the corresponding mean-field assumptions, EFS is proved to generate from the underlying data distribution without the data memorization exhibited by the empirical closed-form diffusion model [27].
4
Mechanistic Analysis of Pretrained Transformers
We established the expressivity of transformers to implement closed-form diffusion models. In this section, we investigate the internal mechanism of sampling in pretrained transformers. Our goal is not to claim that pretrained language models exactly implement a sampling algorithm. Instead, we study how they shape word embedding distribution for sampling. In-context sampling can be formulated as generating new words from a semantic topic, such as animals, cities, or foods, given random words from the topic. We construct prompts consisting of random words from a common semantic category to ensure that the i.i.d. assumption holds for sampling: since words are chosen independently at random, they are independent, and drawing samples from the same topic ensures they are identically distributed. Then, we analyze how the distribution of word embeddings changes across the layers. Specifically, we normalize the embeddings and measure their distance to the uniform distribution on the sphere. We use squared maximum mean discrepancy, MMD2 , as the distance metric; MMD is a standard kernel-based discrepancy between probability distributions [34]. We use a Gaussian kernel to compute MMD. A smaller MMD2 means that the layer-wise embeddings are closer to a uniform spherical reference distribution. We observe a two-stage pattern: the distribution of word embeddings spreads toward a uniform distribution over the sphere and then returns toward a non-uniform distribution for three pretrained foundation models. This observation holds for two variants of LLaMA [17, 18] and CerebrasGPT [19] in Figure 2. To illustrate this dynamic, we present a scatter plot of word embeddings for a small trained model across the layers in Figure 3, where each point represents one word. We can clearly observe the dynamics of convergence to a uniform distribution in the middle layers. We observe a similar pattern when pretrained language models are prompted with natural sentences rather than i.i.d. words. Specifically, we prompt a pretrained model with natural sentences from the CBT dataset [35]. Figure 4 plots the MMD distance between word embeddings and samples drawn uniformly from the unit sphere. We use only one word embedding for repeated words when computing the MMD distance. 6
0.55
0.600 0.5
0.50
Animals Cities Foods
Animals Cities Foods
0.575
0.550
0.45 0.4
0.525
0.40
0.35
0.500
0.3
0.475
0.30 0.2
0.25 Animals Cities Foods
0.20 shaded band: 95% CI = ±1.96 SEM
10
20
30
0.425 0.1
40
50
60
70
80
Llama-3.3-70B-Instruct [17]
0.450
shaded band: 95% CI = ±1.96 SEM
5
10
15
shaded band: 95% CI = ±1.96 SEM
20
25
30
35
40
Openlm-Llama-13b [18]
5
10
15
20
25
30
35
40
Cerebras-GPT-13B [19]
Figure 2: Mechanism of sampling in pretrained language models. y-axis: Layer-wise MMD2 distance between normalized token embeddings and the uniform distribution on the sphere. x-axis: the layer index of word embeddings. The models are prompted with i.i.d. words drawn from categories animal (blue), city (orange), and food (green). Several models exhibit a U-shaped profile: intermediate layers move closer to the uniform reference distribution.
in-context samples
(a) Layer 1
(b) Layer 3
predicted samples
(c) Layer 5
predicted iid sample
(d) Layer 10
(e) Layer 15
(f) Output
Figure 3: Uniform bias for in-context sampling. A small GPT-2-style model is trained to reproduce samples from the two-moons dataset (see Appendix D for details). Orange points represent intermediate embeddings extracted at layers 1, 3, 5, 10, 15, and output layers, illustrating the progressive data evolution across transformer layers at test time. Cbt-Ne Stories
0.18
Cbt-Ne Stories
Cbt-Ne Stories
0.60 0.25 0.55
0.16
0.20
0.50
0.14
0.45
0.15
0.12 0.40 0.10
0.10 0.35
0.08 0.30 shaded band: 95% CI = ±1.96 SEM
10
20
30
0.05
shaded band: 95% CI = ±1.96 SEM
40
50
60
70
80
Llama-3.3-70B-Instruct [17]
5
10
15
20
25
30
35
40
Cerebras-GPT-13B [19]
shaded band: 95% CI = ±1.96 SEM
5
10
15
20
25
30
35
40
Openlm-Llama-13b [18]
Figure 4: Uniform intermediate representations in pretrained language models. y-axis: layerwise MMD2 between normalized token embeddings and the uniform distribution on the unit sphere. x-axis: transformer layer index. Models are prompted with sentences from the CBT dataset, where each repeated word is represented by a single embedding vector.
5
Physics of In-context Sampling
5.1
Observation: Embeddings as Interacting Particles
While proving the sampling capability, Theorem 1 cannot explain the U-shaped mechanism of sampling in pretrained models discussed in the last section. This is due to the fact that closed-form diffusion does not change the empirical data distribution. We use tools from interacting particles to provide insights into the shaping of the embedding distribution for sampling. Consider the following interacting energy defined over word embeddings E(x1 , . . . , xN ) =
N X N X 1 W (s) (xi − xj ). N (N − 1) i=1 j=1 ϵ
(7)
j̸=i (s)
where x1 , . . . , xN ∈ Rd are word embeddings and Wϵ Wϵ(s) (r) =
is an interacting potentials defined as
1 1 ∥r∥2 + , 2 s(∥r∥2 + ϵ)s/2
r ∈ Rd .
The above interacting potential has been used to model the interactive forces between particles in physics. For example, Smale’s 7th open problem aims at characterizing the solution for s → 0 [36]. 7
The mathematical physics community has made significant advances in studying the distribution of points achieving the minimum energy level. When s → 0, it is known that the distribution of minimizers converges to a uniform distribution over a sphere as N → ∞ [37, 38]. We use this energy function to explain the evolution of the data distribution for sampling observed in the last section. We compute the energy of word embeddings where x1 , . . . , xN are word embeddings across layers of pretrained transformers. When the input prompt is generated i.i.d. from the same topic, Figure 5 shows that the energy function produces a U-shaped plot across the layers, similar to the MMD distance plot in Figure 2. Remarkably, optimizing the energy function constructs a sampling method proven by [27]. We prove that a transformer can approximately simulate this method with softmax attention layers and explain how it can explain the U-shape energy dynamics. 1.2
Animals Cities Foods minimum energy = 0.5
1.1
1.1
1.6
Animals Cities Foods minimum energy = 0.5
Animals Cities Foods minimum energy = 0.5
1.4
1.0
1.0
1.2
0.9
0.9
0.8
0.8
0.7
0.7
0.6
0.6
1.0
0.8
0.6 0.5
0.5
0.4
shaded band: 95% CI = ±1.96 SEM
10
20
30
40
50
60
70
80
Llama-3.3-70B-Instruct [17]
0.4
shaded band: 95% CI = ±1.96 SEM
5
10
15
20
25
30
35
40
Openlm-Llama-13b [18]
0.4
shaded band: 95% CI = ±1.96 SEM
5
10
15
20
25
30
35
40
Cerebras-GPT-13B [19]
Figure 5: Repulsive-interactive energy of word embeddings. y-axis: the energy function E for word embeddings across the layers of a pretrained transformer. x-axis: the layer index. We compute the energy in the logarithmic limit s → 0, which ensures the minimum energy is 0.5[37], marked with a red dashed line, and normalized the word embeddings to norm 0.78 to compute the energy.
5.2
Theory: Energy-based sampling with particle gradient descent
Estimation-Free Sampling (EFS) [27] is an energy-based sampling method that proceeds in two steps: (1) minimizing the energy with gradient descent and then (2) maximizing it with inverse gradient descent. The forward step (1) minimizes the energy function with gradient descent as (j+1) (j) (j) (j) xi = xi − γ∇xi EN,ϵ x1 , . . . , xN , j = 0, . . . , K − 1. (k)
(k)
With gradient descent, the energy function decreases and the empirical distribution of x1 , . . . , xN becomes close to a uniform distribution over a sphere or a ball. When the distribution is uniform, the algorithm generates an additional sample uniformly from the sphere/ball, denoted by z (K) ∈ Rd . In the backward phase, an initial reference point z (K) ∈ Rd is transported back toward the empirical distribution, through the following recurrence (j) (j) z (j−1) = z (j) + γ∇z EN,ϵ z (j) , x1 , . . . , xN . The final generated sample is z (0) . We collect the fixed EFS parameters in τ = (K, γ, s, ϵ), together with any fixed schedule choices. The mechanism of EFS is similar to the U-shaped mechanism in pretrained transformers: it first minimizes the energy function to make the empirical data distribution uniform and then maximizes the energy to generate new samples from the target distribution. We prove that transformers can indeed implement EFS for in-context sampling under a weak assumption. Theorem 2 (Transformer simulation of EFS). Suppose the EFS iterates remain in a compact set ⋆ K. Then, for every ε > 0, there exists a choice of transformer parameters θEFS such that, for every (0) (0) d (K) d (0) admissible in-context samples x1 , . . . , xN ∈ R and z ∈ R , if z is generated by the EFS forward–backward recursion above, then (0) (0) ⋆ Πstate TθEFS EncEFS (x1 , . . . , xN , z (K) ) − z (0) ≤ ε. ⋆ The parameter choice θEFS may depend on N, d, K, τ, ε, and the compact domain K, but not on the (0) (0) realized particle cloud x1 , . . . , xN or reference point z (K) .
8
The theorem states that a standard transformer can be parameterized to execute the EFS update rule in context. Since EFS first makes the data distribution uniform and then non-uniform, a transformer implementing EFS follows the same mechanism across the layers. In other words, transformers are computationally expressive to implement a sampler with a similar U-shaped dynamics of energy observed in Figure 5. This result relies on a particular encoding of inputs explained in the Appendix C.1.
6
Experiments
The theory above shows that transformer depth can implement iterative sampling algorithms in context. We now test whether this perspective is visible in trained and pretrained transformers. Our experiments are designed around two questions: (i) can a transformer generate samples from a distribution specified only by the prompt, and (ii) do hidden states exhibit an intermediate transportlike phase consistent with the EFS picture? In-context sampling in a controlled setting. We train a small GPT-2-style decoder-only transformer [16] on two-dimensional i.i.d. point-cloud sequences using a causal next-token objective. Figure 1 shows a held-out compositional test: the training distribution contains the face boundary and eyes, but no smile component. When smile-shaped points are supplied in context, the model generates new samples along the same crescent geometry. This suggests that the prompt acts as an empirical distribution, rather than merely selecting among memorized training components. Layerwise transport in a learned sampler. We train the same architecture on two-moons data and project intermediate hidden states back to the data space. The resulting geometry evolves in a transport-like manner: early layers concentrate the generated points, middle layers spread them toward a more uniform configuration, and later layers recover the structured two-moons geometry specified by the context. A full layer-by-layer visualization is provided in Appendix D.2. Uniformization in pretrained language models. We next examine pretrained language models using prompts of i.i.d. semantic-category words, such as animals, foods, and cities. At each layer, we normalize token embeddings and measure their unbiased RBF MMD2 distance to the uniform distribution on the sphere. Figure 2 shows a U-shaped profile across several models: intermediate layers move closer to the uniform reference, while later layers move away from it toward structured, topic-dependent representations. The same qualitative behavior appears for natural text prompts from the CBT dataset [35] in Figure 4. Additional results and experimental details are deferred to Appendix D.3. Energy-based evidence. Finally, we evaluate the EFS-style interaction energy on the same hiddenstate clouds. Figure 5 shows that the energy follows the same qualitative middle-layer regularization pattern as the MMD curves. This supports the interacting-particle interpretation: intermediate layers move representations toward a lower-energy, more uniform configuration, while later layers recover structured, topic-dependent geometry.
7
Discussion and Limitations
Failure mode across model scales. The uniformization effect is not equally pronounced across all models. In the Qwen2.5 family [20], smaller models exhibit only a short or weak movement toward the uniform reference, whereas larger variants show a clearer two-stage profile as shown in Figure 6. This provides a concrete failure mode for the proposed mechanism: when the model has insufficient effective depth or capacity, the intermediate transport phase may be incomplete. This observation is consistent with our theoretical construction, where transformer depth controls the number of iterative computation steps available to the model. Additional failure-mode experiments for Qwen2.5 and also GPT-2 families are provided in Appendix D.4. Gap between expressivity and mechanism. Our theoretical results are expressivity statements: they show that some transformer parameters can implement closed-form diffusion and EFS from in-context samples, but they do not prove that pretrained language models execute these algorithms 9
0.6
0.6 0.6
Animals Cities Foods
Animals Cities Foods
0.6
Animals Cities Foods
0.5
0.5
0.5
0.4
0.4
Animals Cities Foods
Animals Cities Foods
0.5
0.5
0.4
0.4
0.4
0.3
0.3
0.3
0.3
0.3
0.2
0.2
0.2
shaded band: 95% CI = ±1.96 SEM
5
10
shaded band: 95% CI = ±1.96 SEM
15
20
(a) 1.5B
25
0.2
0.2
5
10
15
20
(b) 3B
25
30
35
0.1
shaded band: 95% CI = ±1.96 SEM
shaded band: 95% CI = ±1.96 SEM
5
10
15
(c) 7B
20
25
10
20
shaded band: 95% CI = ±1.96 SEM
30
(d) 14B
40
50
10
20
30
40
50
60
(e) 32B
Figure 6: Qwen2.5 family [20] on i.i.d. semantic-category prompts. Unlike the clearer U-shaped profiles observed in larger pretrained models, the Qwen2.5 family exhibits a weaker and less pronounced U-shaped pattern. The curves still suggest a middle-layer movement toward a uniform spherical reference distribution followed by a later movement away from it, but the effect is shorter and less robust across model scales. internally. Our experiments provide compatible mechanistic evidence, including middle-layer movement toward a uniform reference distribution and a decrease in EFS-style interaction energy. Still, these observations do not uniquely identify the internal algorithm. This gap is common in in-context learning: prior work shows that transformers can learn regression functions [7], implement gradientdescent-like computations [14, 31], or realize temporal-difference updates [32], but such results mainly characterize representational capacity or behavior in controlled settings. Similarly, our work identifies sampling algorithms that transformers can implement and shows compatible layerwise dynamics, while leaving a full characterization of the mechanism to future work. Generative AI statement. The authors used generative AI tools to assist with code implementation, proof development and verification, and language editing. All mathematical arguments, experimental results, and generated text were independently reviewed and validated by the authors, who take full responsibility for the content of this paper.
10
References [1] Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020. [2] Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240): 1–113, 2023. [3] Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023. [4] Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023. [5] Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023. [6] Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, Heming Xia, Jingjing Xu, Zhiyong Wu, Baobao Chang, et al. A survey on in-context learning. In Proceedings of the 2024 conference on empirical methods in natural language processing, pages 1107–1128, 2024. [7] Shivam Garg, Dimitris Tsipras, Percy S Liang, and Gregory Valiant. What can transformers learn in-context? a case study of simple function classes. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems, volume 35, pages 30583–30598. Curran Associates, Inc., 2022. URL https://proceedings.neurips.cc/paper_files/paper/2022/file/ c529dba08a146ea8d6cf715ae8930cbe-Paper-Conference.pdf. [8] Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou. What learning algorithm is in-context learning? investigations with linear models. In International Conference on Learning Representations, 2023. [9] Yingcong Li, Muhammed Emrullah Ildiz, Dimitris Papailiopoulos, and Samet Oymak. Transformers as algorithms: Generalization and stability in in-context learning. In International Conference on Machine Learning, pages 19565–19594, 2023. [10] Yu Bai, Fan Chen, Huan Wang, Caiming Xiong, and Song Mei. Transformers as statisticians: Provable in-context learning with in-context algorithm selection. In Advances in Neural Information Processing Systems, volume 36, 2023. [11] Deqing Fu, Tian-Qi Chen, Robin Jia, and Vatsal Sharan. Transformers learn higher-order optimization methods for in-context learning: A study with linear models. arXiv preprint arXiv:2310.17086, 2023. [12] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. [13] Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. Highresolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. [14] Kwangjun Ahn, Xiang Cheng, Hadi Daneshmand, and Suvrit Sra. Transformers learn to implement preconditioned gradient descent for in-context learning. Advances in Neural Information Processing Systems, 36:45614–45650, 2023. 11
[15] Christopher Scarvelis, Haitz Sáez de Ocáriz Borde, and Justin Solomon. Closed-form diffusion models. Transactions on Machine Learning Research, 2025. ISSN 2835-8856. URL https: //openreview.net/forum?id=JkMifr17wc. [16] Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019. [17] Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024. [18] Xinyang Geng and Hao Liu. Openllama: An open reproduction of llama, May 2023. URL https://github.com/openlm-research/open_llama. [19] Nolan Dey, Gurpreet Gosal, Hemant Khachane, William Marshall, Ribhu Pathria, Marvin Tom, Joel Hestness, et al. Cerebras-gpt: Open compute-optimal language models trained on the cerebras wafer-scale cluster. arXiv preprint arXiv:2304.03208, 2023. [20] A Yang Qwen, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengpeng Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al. Qwen2. 5 technical report. arXiv preprint, 2024. [21] Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Mérouane Debbah, Étienne Goffinet, Daniel Hesslow, Julien Launay, Quentin Malartic, et al. The falcon series of open language models. arXiv preprint arXiv:2311.16867, 2023. [22] Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022. [23] BigScience Workshop, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, et al. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100, 2022. [24] Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020. [25] Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020. [26] Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020. [27] Hadi Daneshmand and Ashkan Soleymani. Data generation without function estimation. NeurIPS workshop on optimization for machine learning, 2025. [28] Zhendong Wang, Yifan Jiang, Yadong Lu, Yelong Shen, Pengcheng He, Weizhu Chen, Zhangyang Wang, and Mingyuan Zhou. In-context learning unlocked for diffusion models. In Advances in Neural Information Processing Systems, volume 36, 2023. [29] Brian DO Anderson. Reverse-time diffusion equation models. Stochastic Processes and their Applications, 12(3):313–326, 1982. [30] Hadi Daneshmand. In-context learning for discrete optimal transport: Can transformers sort? In The 29th International Conference on Artificial Intelligence and Statistics. [31] Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov. Transformers learn in-context by gradient descent. In International Conference on Machine Learning, pages 35151–35174, 2023. 12
[32] Jiuqi Wang, Ethan H Blaser, Hadi Daneshmand, and Shangtong Zhang. Transformers learn temporal difference methods for in-context reinforcement learning. In ICML 2024 Workshop on In-Context Learning. [33] Paul Rosu, Lawrence Carin, and Xiang Cheng. From softmax to score: Transformers can effectively implement in-context denoising steps. Advances in Neural Information Processing Systems, 38:142994–143017, 2026. [34] Arthur Gretton, Karsten M Borgwardt, Malte J Rasch, Bernhard Schölkopf, and Alexander Smola. A kernel two-sample test. The journal of machine learning research, 13(1):723–773, 2012. [35] Felix Hill, Antoine Bordes, Sumit Chopra, and Jason Weston. The goldilocks principle: Reading children’s books with explicit memory representations. arXiv preprint arXiv:1511.02301, 2015. [36] Steve Smale. Mathematical problems for the next century. The mathematical intelligencer, 20 (2):7–15, 1998. [37] Rupert L Frank and Ryan W Matzke. Minimizers for an aggregation model with attractive– repulsive interaction. Archive for Rational Mechanics and Analysis, 249(2):15, 2025. [38] Sylvia Serfaty. Mean field limit for coulomb-type flows. Duke Mathematical Journal, 169(15): 2887–2935, 2020. Appendix by Mitia Duerinckx. [39] Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen. What makes good in-context examples for GPT-3? In Proceedings of Deep Learning Inside Out: The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures, pages 100–114, 2022. [40] Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, pages 8086–8098, 2022. [41] Zhiyong Wu, Yaoxiang Wang, Jiacheng Ye, and Lingpeng Kong. Self-adaptive in-context learning: An information compression perspective for in-context example selection and ordering. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, pages 1423–1436, 2023. [42] Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al. In-context learning and induction heads. arXiv preprint arXiv:2209.11895, 2022. [43] Lean Wang, Lei Li, Damai Dai, Deli Chen, Hao Zhou, Fandong Meng, Jie Zhou, and Xu Sun. Label words are anchors: An information flow perspective for understanding in-context learning. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 9840–9855, 2023. [44] Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma. An explanation of in-context learning as implicit Bayesian inference. In International Conference on Learning Representations, 2022. [45] Yufeng Zhang, Fengzhuo Zhang, Zhuoran Yang, and Zhaoran Wang. What and how does in-context learning learn? Bayesian model averaging, parameterization, and generalization. arXiv preprint arXiv:2305.19420, 2023. [46] Damai Dai, Yutao Sun, Li Dong, Yaru Hao, Shuming Ma, Zhifang Sui, and Furu Wei. Why can GPT learn in-context? language models secretly perform gradient descent as meta-optimizers. In Findings of the Association for Computational Linguistics: ACL 2023, pages 4005–4019, 2023. [47] Arvind Mahankali, Tatsunori B. Hashimoto, and Tengyu Ma. One step of gradient descent is provably the optimal in-context learner with one layer of linear self-attention. arXiv preprint arXiv:2307.03576, 2023. 13
[48] Hadi Daneshmand. In-context learning for discrete optimal transport: Can transformers sort? AISTATS, 2026. [49] Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. Qwen technical report. arXiv preprint arXiv:2309.16609, 2023.
Technical appendices and supplementary material A
Related Work
In-context learning. In-context learning refers to the ability of a model to adapt its behavior at inference time using only information supplied in the prompt, without updating its parameters. This phenomenon was popularized by large language models [1–5, 17] and has since become a central topic in the study of transformer-based models [6]. Much of the empirical literature studies how performance depends on prompt construction, including the choice, ordering, and formatting of demonstrations [39–41]. Complementary work studies mechanisms behind in-context learning, including induction heads, attention-based information flow, Bayesian inference, and gradient-descentlike computation inside transformers [42–46, 31, 14, 47]. Most theoretical studies of in-context learning focus on prediction. In the standard formulation, the prompt contains input–output examples from an unknown task, and the model predicts the output for a new query. Garg et al. [7] showed that transformers trained from scratch can learn simple function classes from examples in the prompt, including linear functions, sparse linear functions, decision trees, and shallow neural networks. Related work studies whether transformers implement particular learning algorithms in context, including ridge regression, gradient descent, higher-order optimization, and algorithm selection [8–11]. Our work follows this algorithmic view, but changes the target from prediction to generation: instead of asking whether a transformer can predict a label or function value, we ask whether it can execute a sampler in context. Transformers as algorithm executors. A growing line of work studies transformers as models capable of executing structured computations. Prior work has shown that transformer depth and attention can implement iterative optimization procedures, temporal-difference style updates, and discrete optimal transport or sorting-type computations [14, 32, 48, 30]. These results suggest that attention and depth can act as computational resources: attention aggregates information from the prompt, while depth implements repeated updates. Our results extend this perspective to generative algorithms. Closed-form diffusion and EstimationFree Sampling are not prediction rules; they are iterative procedures that transform an initial state into a sample. Showing that frozen transformers can implement these procedures in context strengthens the interpretation of transformers as general-purpose algorithm executors. The main difference from prior algorithmic ICL work is that the output of the computation is not a predicted label or regression value, but the terminal state of a sampler. Diffusion and score-based generative models. Diffusion models generate samples by gradually transforming noise into data through an iterative reverse process [24]. Score-based generative modeling provides a continuous-time formulation in which a forward stochastic process maps data to noise and a reverse-time process generates data from noise using estimated score functions [25]. Related samplers, including DDIM, modify the reverse process through deterministic or implicit updates [26]. These methods have made iterative sampling a central computational paradigm in modern generative modeling. Our work does not propose a new diffusion model or train a new score network. Instead, we isolate the computational structure of diffusion-style sampling. In the closed-form setting, the score and sampling dynamics are explicit functions of the empirical samples and the current state. This allows us to ask a representational question: can a frozen transformer implement the corresponding sampler when the empirical distribution and initial state are supplied in the prompt? Our answer is affirmative, showing that a core diffusion-style computation can be realized as in-context computation by a transformer. 14
In-context generation with diffusion models. There is prior work on enabling in-context learning inside diffusion models. Prompt Diffusion trains a diffusion-based model to perform vision tasks from example input–output pairs and a query image, thereby bringing in-context learning to diffusion-based generation [28]. The paper is the closest prior work connecting diffusion models and in-context learning, but its direction is different: it equips a diffusion model with in-context learning ability. Our question reverses this direction. Rather than asking whether a diffusion model can perform in-context learning, we ask whether a transformer can implement diffusion-style and particle-based samplers in context. Thus, the central object in our paper is not a diffusion architecture conditioned on examples, but a frozen transformer that executes a sampling algorithm specified by the prompt. Estimation-Free Sampling. Estimation-Free Sampling (EFS) proposes a generative mechanism that avoids explicit score estimation and neural generative training [27]. Instead of learning a score or density-dependent function, EFS uses deterministic particle optimization. Its forward step transports empirical particles toward a simple reference distribution, and its backward step maps a newly sampled reference point back toward the data distribution. The EFS paper explicitly frames this as a generation without score/function estimation, neural-network training, or noise injection in the mean-field regime. We include EFS because it tests whether in-context sampling is limited to diffusion-style score dynamics. Our result shows that it is not. A frozen transformer can also implement a particleoptimization sampler in context. This broadens the scope of the theory: transformers can realize not only closed-form diffusion updates, but also deterministic particle-based generative algorithms. Positioning. Taken together, prior work shows that transformers can learn from context, diffusion models can generate through iterative sampling, and particle methods can generate without explicit function estimation. Our contribution connects these threads. We introduce in-context sampling as a framework for studying generative algorithms executed by frozen transformers. In this framework, the prompt specifies the sampling instance, attention computes interactions among samples or particles, and transformer depth performs the iterative updates. This perspective broadens in-context learning from prediction to generation and places closed-form diffusion and Estimation-Free Sampling under a common computational view.
B
Proof of Theorem 1
We use the prompt encoding from Section 2.2. Each token has a block form (x) (r) (z) (k) (1) Xi = Xi , Xi , Xi , Xi , Xi , (x)
(z)
(k)
(x)
= xi ,
where Xi , Xi , Xi Xi
(r)
(1)
∈ Rd , and Xi , Xi (r)
Xi
∈ R. For data tokens i ≤ N , (z)
= ∥xi ∥2 ,
Xi
(k)
= Xi
= 0d ,
(1)
Xi
= 1.
The state token is i = N + 1; its z-block stores the current sampler state and its k-block is scratch space. Let L be the number of transformer layers. We take L = S, one layer per Euler step, and write the schedule as τ = {(tℓ , hℓ )}L−1 ℓ=0 . We prove the result by induction over layers. The invariant is that after ℓ blocks, the state token stores zℓ in its z-block and 0d in its k-block, while the data tokens still store xi and ∥xi ∥2 . The base case follows from the prompt encoding. It remains to construct one block Bt,h that maps z 7→ z + hvt (z) and resets the scratch block. B.1
Attention layer convention
A single-head attention layer is specified by fixed matrices WQ , WK ∈ Rp×dq , 15
WV ∈ Rp×p .
For X ∈ R(N +1)×p , define qi = Xi WQ ,
ki = Xi WK ,
vi = Xi WV .
The state token attends only to data tokens 1, . . . , N . Hence exp(⟨qN +1 , ki ⟩) , aN +1,i (X) = PN j=1 exp(⟨qN +1 , kj ⟩) and Attn(X)N +1 =
N X
i = 1, . . . , N,
aN +1,i (X)vi .
i=1
p The usual 1/ dq scaling can be absorbed into WQ or WK , so we omit it. B.2
One closed-form diffusion step
Fix t ∈ (0, 1) and h > 0. By (5), vt (z) = −
1 1 z+ kt (z). 1−t t(1 − t)
Thus it suffices for attention to compute kt (z). Query, key, and value projection weights. at =
Let
t , (1 − t)2
bt = −
t2 . 2(1 − t)2
Choose dq = d + 1. With block ordering (x, r, z, k, 1), define 0d×d 0d×1 Id 01×d 01×d 0 (t) (t) WQ = at Id 0d×1 , WK = 0d×d 0 0 d×d 0d×1 d×d 01×d 1 01×d
0d×1 bt 0d×1 . 0d×1 0
Therefore (t) (z) (1) qi = Xi WQ = at Xi , Xi ,
(x) (t) (r) ki = Xi WK = Xi , bt Xi .
For the state token N + 1 and a data token i ≤ N , t2 t ⟨z, xi ⟩ − ∥xi ∥2 . ⟨qN +1 , ki ⟩ = 2 (1 − t) 2(1 − t)2 Moreover, ∥z − txi ∥2 t t2 ∥z∥2 2 − = ⟨z, x ⟩ − ∥x ∥ − . i i 2(1 − t)2 (1 − t)2 2(1 − t)2 2(1 − t)2 The last term is independent of i, so it cancels in the softmax over data tokens. Hence aN +1,i (X) = wi (t, z),
i = 1, . . . , N,
where wi (t, z) are the responsibility weights defined in (2). The value projection is the block matrix 0d×d 01×d (t) WV = 0d×d 0 d×d 01×d
0d×1 0 0d×1 0d×1 0
0d×d 01×d 0d×d 0d×d 01×d
tId 01×d 0d×d 0d×d 01×d
0d×1 0 0d×1 . 0d×1 0
Thus, for a data token i, (t)
vi = Xi WV = [0d , 0, 0d , txi , 0], and the attention output at the state token satisfies (k) Attn(X)N +1 =
N X
aN +1,i (X)txi =
i=1
N X i=1
16
wi (t, z)(txi ) = kt (z).
Feedforward update.
Let
Y = X + Attn(X). The feedforward layer is the row-wise linear map (t,h)
FFt,h (Yi ) = Yi WFF , where, with the same block ordering (x, r, z, k, 1), 0d×d 0d×1 0d×d 01×d 0 01×d h (t,h) 0 0 − d×1 WFF = d×d 1−t Id h 0d×d 0d×1 t(1−t) Id 01×d 0 01×d
0d×d 01×d 0d×d −Id 01×d
0d×1 0 0d×1 . 0d×1 0
Equivalently, (z)
FFt,h (Yi ) = −
h h (z) (k) Yi + Y , 1−t t(1 − t) i
(k)
(k)
FFt,h (Yi ) = −Yi
,
and all other output blocks are zero. For the state token, (z)
(k)
YN +1 = z,
YN +1 = kt (z).
Therefore the residual update gives (z)
(z)
(z)
Bt,h (X)N +1 = YN +1 + FFt,h (YN +1 ) = z − and
(k)
(k)
h h z+ kt (z) = z + hvt (z), 1−t t(1 − t) (k)
Bt,h (X)N +1 = YN +1 + FFt,h (YN +1 ) = 0d . Thus Bt,h implements one closed-form diffusion Euler step, leaves the data tokens unchanged, and resets the scratch block. B.3
Completion of the induction
For each layer ℓ = 0, . . . , L − 1, choose Bℓ = Btℓ ,hℓ . ⋆ The parameter choice θCFD is obtained by stacking ⋆ TθCFD = BL−1 ◦ · · · ◦ B0 .
The base case ℓ = 0 follows from the prompt encoding. If the induction invariant holds at layer ℓ, then Bℓ updates zℓ 7→ zℓ + hℓ vtℓ (zℓ ) = zℓ+1 , resets the k-block to zero, and leaves the data tokens unchanged. Hence the invariant holds at layer ℓ + 1. By induction, after L layers the state token stores zL . Therefore ⋆ Πstate TθCFD Enc({xi }N = zL , i=1 , z0 ) where zL is generated by the closed-form diffusion recursion. B.4
Smoothed closed-form diffusion
The smoothed case uses the same induction argument, with one modification: the one-step block computes kσ,t (z) instead of kt (z). Fix t ∈ (0, 1), h > 0, smoothing level σ ≥ 0, and perturbations ϵ1 , . . . , ϵM ∈ Rd . Recall that
M
1 X kσ,t (z) = kt (z + σϵm ). M m=1 17
Thus we compute kt (z + σϵm ) for each m and average the results. (t)
(t)
(t)
Use M attention heads. For head m, keep WK and WV as above and replace WQ by 0d×d 0d×1 0 01×d (σ,t) WQ,m = at Id 0d×1 . 0d×d 0d×1 at σϵ⊤ 1 m Then the state-token query in head m is (m) qN +1 = at (z + σϵm ), 1 . For a data token i ≤ N , (m)
⟨qN +1 , ki ⟩ =
t t2 ⟨z + σϵm , xi ⟩ − ∥xi ∥2 . 2 (1 − t) 2(1 − t)2
This equals the Gaussian responsibility logit for the perturbed state z + σϵm , up to an additive term independent of i. Hence (m) aN +1,i (X) = wi (t, z + σϵm ). (t)
Using the same value projection WV , head m outputs N X
wi (t, z + σϵm )(txi ) = kt (z + σϵm )
i=1
in its scratch output. The multi-head output projection averages these heads: M
1 X kt (z + σϵm ) = kσ,t (z). M m=1 (t,h)
The same feedforward matrix WFF , with kt (z) replaced by kσ,t (z), updates the state token to z + hvσ,t (z) = z −
h h z+ kσ,t (z), 1−t t(1 − t)
and resets the scratch block. Applying this construction at every layer ℓ = 0, . . . , L−1 with t = tℓ and h = hℓ , the same induction ⋆ gives a parameter choice θσ-CFD such that ⋆ Πstate Tθσ-CFD Enc({xi }N = zL , i=1 , z0 ) where zL is generated by the smoothed closed-form diffusion recursion. This completes the proof.
C
Proof of Theorem 2
C.1
Prompt encoding for EFS
We now describe the prompt used in Theorem 2. The EFS parameters τ = (K, γ, s, ϵ) are fixed and compiled into the transformer weights. The prompt stores only the instance-specific quantities: the (0) (0) initial particles x1 , . . . , xN and the reference point z (K) . We use m = N + 2 tokens. The first N rows are particle tokens, row N + 1 stores the reference point z (K) , and row N + 2 is the state token. Each token is partitioned into six blocks: [0] [1] (x , x , . . . , x[K] ) (r[0] , r[1] , . . . , r[K] ) z rz g 1 . Here x[j] ∈ Rd stores the j-th forward particle iterate, r[j] ∈ R stores its squared norm, z ∈ Rd stores the moving point during the backward phase, rz ∈ R stores ∥z∥2 , g ∈ Rd is scratch space for interaction terms, and the final coordinate is a constant feature. Therefore, p = (K + 1)d + (K + 1) + d + 1 + d + 1 = (K + 3)d + K + 3. 18
The initial prompt is (0) ⊤ ⊤ (x1 ) , 0d , . . . , 0⊤ d .. . XEFS = (0) ⊤ ⊤ ⊤ (xN ) , 0d , . . . , 0d 0⊤ , 0⊤ , . . . , 0⊤ d
d
(0)
0⊤ d .. .
∥x1 ∥2 , 0, . . . , 0 .. . (0)
∥xN ∥2 , 0, . . . , 0 0⊤ d (K) ⊤ 0, 0, . . . , 0 (z ) 0, 0, . . . , 0 0⊤ d
d
⊤ ⊤ 0⊤ d , 0 d , . . . , 0d
(0)
∥z
0 .. .
0⊤ d .. .
0
0⊤ d 0⊤ d 0⊤ d
(K) 2
∥
0
1 .. . ∈ R(N +2)×p . 1 1 1
(0)
Initially, particle row i ≤ N stores xi and ∥xi ∥2 , while the later blocks x[1] , . . . , x[K] and r[1] , . . . , r[K] are zero. During the forward phase, the transformer fills these blocks with approxima(j) tions of the forward EFS iterates xi and their squared norms. The reference row stores z (K) and (K) 2 ∥z ∥ . The final row is the state token used during the backward phase. The readout extracts the z-block of this final row: (z) Πstate (X) = XN +2 . After the transformer computation, this readout approximates the EFS output z (0) . We use the notation of Section C.1, with one auxiliary state-token selector coordinate. Each token is partitioned as [0] [K] [0] [K] Xi = xi , . . . , xi | ri , . . . , ri | zi | rz,i | gi | ci | ηi , [j]
[j]
where xi , zi , gi ∈ Rd , ri , rz,i , ci , ηi ∈ R, ci = 1, and ηi identifies the state token. For particle tokens i ≤ N and the reference token i = N + 1, set ηi = 0. For the state token i = N + 2, set ηN +2 = 1. The state token has zero x[j] - and r[j] -blocks. For an attention head h, queries, keys, and values are qi = Xi WQ,h ,
ki = Xi WK,h ,
vi = Xi WV,h .
We use standard softmax attention over the N + 2 tokens: exp(⟨qi , kℓ ⟩) (h) αiℓ = PN +2 , r=1 exp(⟨qi , kr ⟩) and
N +2 X
Attn(h) (X)i =
ℓ = 1, . . . , N + 2,
(h)
αiℓ vℓ .
ℓ=1
C.2
Explicit kernel head
Fix a stage j ∈ {0, . . . , K} and a kernel parameter λ > 0. We first consider the forward-stage query token i ≤ N . The goal is to recover N X
[j] [j] 2 [j] [j] e−λ∥xi −xℓ ∥ xi − xℓ .
ℓ=1 [j]
Let the query/key dimension be d + 1. Define WQ,λ ∈ Rp×(d+1) by [j]
WQ,λ :
x[j] 7→ 2λx[j] ,
c 7→ 1,
with all other input blocks mapped to zero. Equivalently, [j]
qi = [2λxi | 1]. [j]
Define WK,λ ∈ Rp×(d+1) by [j]
WK,λ :
x[j] 7→ x[j] ,
r[j] 7→ −λr[j] ,
with all other input blocks mapped to zero. Thus [j]
[j]
kℓ = [xℓ | −λrℓ ]. 19
[j]
[j]
For a particle token ℓ ≤ N , since rℓ = ∥xℓ ∥2 , [j]
[j]
[j]
⟨qi , kℓ ⟩ = 2λ⟨xi , xℓ ⟩ − λ∥xℓ ∥2 . [j]
[j]
For the state token N + 2, xN +2 = 0 and rN +2 = 0, so kN +2 = 0 and the corresponding logit is zero. Let αiℓ be the attention weight on token ℓ, and write αi0 ≡ αi,N +2 for the attention weight on the state token. Then αiℓ [j] [j] [j] = exp 2λ⟨xi , xℓ ⟩ − λrℓ , ℓ ≤ N. αi0 [j]
Multiplying by e−λri gives [j]
e−λri
[j] 2 [j] αiℓ = e−λ∥xi −xℓ ∥ . αi0
[j]
Now define the value projection WV,λ ∈ Rp×p by [j]
x[j] 7→ g,
WV,λ :
η 7→ rz ,
with all other output blocks zero. Hence [j]
(vℓ )(g) = xℓ ,
(vℓ )(rz ) = ηℓ .
Therefore (g)
Attn(X)i
=
N +2 X
[j]
αiℓ xℓ =
N X
[j]
αiℓ xℓ ,
ℓ=1
ℓ=1 [j]
because the reference and state tokens have zero x -blocks, and (r )
Attn(X)i z =
N +2 X
αiℓ ηℓ = αi,N +2 = αi0 .
ℓ=1
Assuming that EFS iterates lies in a compact set, all logits are uniformly bounded on the compact domain, so αi0 is uniformly bounded away from zero. Therefore the tokenwise feedforward layer [j] can uniformly approximate division by αi0 , multiplication by e−λri , and the map ! N N X X [j] [j] 2 [j] [j] [j] [j] [j] xi , ri , αiℓ xℓ , αi0 7→ e−λ∥xi −xℓ ∥ xi − xℓ . ℓ=1
ℓ=1
Thus this head, followed by a tokenwise feedforward map, computes the Gaussian-kernel interaction sum uniformly on the compact domain. C.3
Backward-stage kernel head
For the backward stage, the query token is the state token N + 2, whose z-block stores the current [j] [j] backward iterate z (j) . The key and value projections remain WK,λ and WV,λ . The query projection fQ,λ , defined by is replaced by W fQ,λ : W
z 7→ 2λz,
c 7→ 1,
with all other input blocks mapped to zero. Thus [j]
qN +2 = [2λzN +2 | 1],
The same argument yields, up to arbitrary uniform approximation, N X
[j]
kℓ = [xℓ | −λrℓ ].
[j] 2 [j] e−λ∥zN +2 −xℓ ∥ zN +2 − xℓ .
ℓ=1
20
C.4
Approximating the EFS interaction field
Let a = s/2 + 1. The Laplace identity gives Z ∞ 2 1 v = ta−1 e−tϵ e−t∥v∥ v dt. (∥v∥2 + ϵ)a Γ(a) 0 On the compact domain, all relevant differences satisfy ∥v∥ ≤ R. Hence, for every δ > 0, there exist H ∈ N, nodes λ1 , . . . , λH > 0, and coefficients ω1 , . . . , ωH ∈ R such that H
sup ∥v∥≤R
X 2 v − ωh e−λh ∥v∥ v ≤ δ. 2 a (∥v∥ + ϵ) h=1
Using one explicit kernel head for each λh , and combining the head outputs by the output projection with coefficients ωh , the transformer computes N X
[j]
y − xℓ
[j] 2 s/2+1 ℓ=1 (∥y − xℓ ∥ + ϵ)
up to uniform error. The attractive term is affine after summation: N N X X [j] [j] (y − xℓ ) = N y − xℓ . ℓ=1
ℓ=1
It is computed by a uniform-attention head. Explicitly, choose WQavg = 0,
avg WK = 0,
WVavg : x[j] 7→ g,
with all other output blocks zero. Then all logits are zero, so attention returns (g)
Attnavg (X)i
=
N +2
N
ℓ=1
ℓ=1
1 X [j] 1 X [j] xℓ = xℓ , N +2 N +2 [j]
because the reference and state tokens have zero x -blocks. A row-wise affine feedforward map PN [j] multiplies this by N + 2 and combines it with y to obtain N y − ℓ=1 xℓ . Combining the attractive and repulsive parts yields a uniform approximation of the EFS gradient field. C.5
Forward and backward simulation [j]
[j]
At forward stage j = 0, . . . , K − 1, particle token i ≤ N stores xi and ri . The layer uses the forward-stage interaction construction to approximate (j) (j) ∇xi EN,ϵ x1 , . . . , xN and writes
[j+1] [j] [j+1] [j+1] 2 b (j) , xi = xi − γ G ri ≈ ∥xi ∥ . i After the forward phase, a copy layer writes the reference row into the state token: (z)
(r )
XN +2 = z (K) ,
z XN +2 = ∥z (K) ∥2 .
At backward stage j = K, K − 1, . . . , 1, the state token uses the backward-stage interaction construction to approximate (j) (j) ∇z EN,ϵ z (j) , x1 , . . . , xN , and the feedforward update writes (z) (z) b (j) , XN +2 ← XN +2 + γ G
(r )
(z)
z XN +2 ≈ ∥XN +2 ∥2 .
Since the number of stages is finite and all maps are continuous on the compact domain, the accumulated error satisfies a finite stability recursion. Choosing the per-layer approximation error sufficiently small yields (0) (0) ⋆ Πstate TθEFS EncEFS (x1 , . . . , xN , z (K) ) − z (0) ≤ ε. This proves Theorem 2. 21
D
Experiments Details
D.1
Smile Experiment
Figure 1 provided a controlled visual example of the sampling behavior studied in this work. We construct a two-dimensional smiling emoji distribution with four components: a noisy circular face boundary, two Gaussian eye clusters, and a crescent-shaped smile. The smile component is removed from the training distribution. We use a small GPT-2-style decoder-only transformer trained from scratch with 16 layers and 32-dimensional hidden embeddings. Each training example contains 64 input points, and the model is trained on 1,000 IID training sequences for 2,000 optimization steps. Because the architecture is decoder-only and causal, we formulate the task in the same way as next-token prediction. For each IID sequence (x1 , . . . , xT +1 ), the model receives (x1 , . . . , xT ) as input and is trained to predict the shifted sequence (x2 , . . . , xT +1 ). The final target point xT +1 is not included in the input context, so it acts as a fresh IID sample from the same distribution. Since the samples are unordered, we train with a full-sequence optimal-transport matching loss rather than an index-wise regression loss. This encourages the predicted point cloud to match the target distribution, while remaining compatible with the next-token structure of the transformer. For visualization, we plot only the prediction at the last position. It is produced after the model has attended to the entire in-context sequence. Earlier positions are also trained, but each of them is conditioned only on a shorter prefix because of the causal mask. The held-out smile gives the central qualitative test. Although smile-shaped samples are never included in the training distribution, the model can generate new IID samples along the same smile geometry. D.2
Synthetic layerwise geometry
Figure 3 used the two-moons distribution to show how the learned sampler transforms an in-context point cloud across depth. We generate a pool of 5,000 2-dimensional points and split it at the point level. Using the same causal next-token sampling setup as in Subsection D.1, we train the model for 2,000 optimization steps on 1,000 IID training sequences, where each input sequence contains 64 points. To inspect the internal computation, we project the hidden states after selected transformer layers back to the two-dimensional output space using the model’s output head. In early layers, the predicted points are concentrated in a small region, then expand in the middle layers toward a more uniform geometry, and finally contract back onto the structured two-moons distribution specified by the in-context samples. This layerwise trajectory visually supports the view of transformer depth as an iterative transport process. Full layer-by-layer visualization is provided in Fig 7.
in-context samples
predicted samples
predicted iid sample
(a) Input emb.
(b) Layer 1
(c) Layer 2
(d) Layer 3
(e) Layer 4
(f) Layer 5
(g) Layer 6
(h) Layer 7
(i) Layer 8
(j) Layer 9
(k) Layer 10
(l) Layer 11
(m) Layer 12
(n) Layer 13
(o) Layer 14
(p) Layer 15
(q) Layer 16
(r) Output
Figure 7: Geometry evolution of one test sample across the transformer layers. 22
D.3
Layerwise Uniformization in Large Language Models
Figure 2 shows that pretrained autoregressive language models reshape in-context semantic distributions across depth. We construct prompts from three topic vocabularies: animals, foods, and cities. For each model, we keep only tokenizer-eligible single-token words and use unique tokens, so each prompt contains up to 256 distinct words from one topic. We run 5 independent prompts per topic and record the raw hidden states after each transformer block, excluding the embedding layer, final layer normalization, and vocabulary logits. The complete semantic-token experiment was conducted across nine pretrained autoregressive language models; the corresponding layerwise MMD results are reported in Figure 8. 0.55
0.6
Animals Cities Foods
0.50 0.5
0.5
0.45
0.40
Animals Cities Foods
0.4
0.4
0.35 0.3
0.30
0.3
0.25 Animals Cities Foods
0.20 shaded band: 95% CI = ±1.96 SEM
10
20
30
40
50
60
70
0.2
0.2 shaded band: 95% CI = ±1.96 SEM
80
10
(a) Llama-3.3-70B-Instruct [17] Animals Cities Foods
shaded band: 95% CI = ±1.96 SEM
20
30
40
50
10
(b) GPT-2 XL [16] 0.60
20
30
40
50
(c) Qwen2.5-14B [49]
Animals Cities Foods
Animals Cities Foods
0.60
0.5 0.55 0.55 0.50 0.4 0.45
0.50
0.40
0.3 0.45
0.35
0.2 0.30 0.40
shaded band: 95% CI = ±1.96 SEM
10
20
30
40
50
shaded band: 95% CI = ±1.96 SEM
60
10
(d) Qwen2.5-32B [49] 0.55
0.50
shaded band: 95% CI = ±1.96 SEM
20
30
40
50
60
10
(e) Falcon-40B [21]
20
30
40
50
60
(f) OPT-66B [22]
0.600 Animals Cities Foods
Animals Cities Foods
0.575
0.45
0.550
0.40
0.525
0.35
0.500
0.30
0.475
0.25
0.450
0.5
Animals Cities Foods
0.4
0.20
0.3
0.2
0.425 shaded band: 95% CI = ±1.96 SEM
5
10
shaded band: 95% CI = ±1.96 SEM
15
20
(g) Bloom-7b1 [23]
25
30
5
10
15
0.1 20
25
30
35
(h) Cerebras-GPT-13B [19]
40
shaded band: 95% CI = ±1.96 SEM
5
10
15
20
25
30
35
40
(i) Openlm-Llama-13b [18]
Figure 8: Additional large-model results for IID semantic-category prompts. The curves show the layerwise MMD2 distance to the uniform spherical reference distribution. In Figure 4, we also repeat the same analysis on natural text prompts from the CBT dataset [35]: for each story, we feed the first 256 tokens to the model, but compute the layerwise metrics only on the first occurrence of each unique token. At each layer, we treat the hidden states as an empirical particle cloud. We apply row-center normalization by subtracting the coordinate mean of each hidden vector and projecting it to the unit sphere. We then compare the resulting cloud to samples from the uniform spherical distribution using unbiased RBF MMD2 . Smaller MMD2 indicates that the hidden-state distribution is closer to the uniform reference. The curves report the mean over trials, with shaded bands denoting 95% confidence intervals. 23
Across several large language models, the curves show a U-shaped profile: early layers preserve topic-specific structure, middle layers move the token cloud closer to the uniform sphere, and later layers move away from uniformity toward a structured, topic-dependent representation. In Figure 5, we further measure the corresponding EFS-style logarithmic interaction energy. For this metric, the normalized hidden states are rescaled to radius 0.78, and the energy is computed directly on the layerwise particle cloud. The energy curves support the same interpretation as the MMD plots: transformer layers first regularize the in-context distribution toward a uniform-like configuration and later recover structured semantic geometry. D.4
Failure mode
The layerwise uniformization pattern is not equally strong across all models. This effect is visible not only within the Qwen2.5 family [20], but also within the GPT-2 family [16], and it appears in both the i.i.d. semantic-topic setting and the CBT story setting; see Figure 9, Figure 10, Figure 6, and Figure 11. Figure 6 shows this effect within the Qwen2.5 model family [20]. Smaller models, especially the 1.5B, 3B, and 7B variants, show only a short or weak movement toward the uniform spherical reference before the MMD2 distance increases again. In these models, the middle-layer uniformization phase is compressed, suggesting that the model may not have enough depth to fully realize the two-stage sampler-like computation. The larger 14B and 32B models show a clearer separation between the early movement toward uniformity and the later movement away from it. This behavior is consistent with the theoretical construction. In Theorem 1, the transformer simulates an iterative sampler by assigning transformer blocks to sampler steps. Therefore, depth is not only an architectural detail; it controls how many test-time computational steps the model can perform. If the model has too few layers, the transport process may be incomplete: the representation starts moving toward a reference geometry, but the model does not have enough intermediate computation to form a stable uniform-like phase before reconstructing semantic structure. We therefore interpret shallow or smaller models as a failure mode of in-context sampling. 0.55
Animals Cities Foods
0.60
Animals Cities Foods
Animals Cities Foods
Animals Cities Foods
0.5 0.50
0.5
0.55 0.4
0.4
0.45 0.50 0.40
0.3 0.3 0.45
0.35 0.2 0.2 0.30
0.40 shaded band: 95% CI = ±1.96 SEM
2
4
shaded band: 95% CI = ±1.96 SEM
6
8
10
12
5
(a) GPT-2
10
shaded band: 95% CI = ±1.96 SEM
15
20
25
5
(b) GPT-2 Medium
10
15
shaded band: 95% CI = ±1.96 SEM
20
25
30
35
10
(c) GPT-2 Large
20
30
40
50
(d) GPT-2 XL
Figure 9: GPT-2 family [16] on IID semantic-category prompts. Across model sizes, the layerwise distance to the uniform spherical reference distribution decreases in the middle layers and increases near the output layers. Cbt-Ne Stories
0.500
0.40
Cbt-Ne Stories
0.40 Cbt-Ne Stories
Cbt-Ne Stories
0.50 0.475
0.35
0.35
0.45 0.450 0.40
0.30
0.30
0.25
0.25
0.425 0.35
0.400
0.30
0.375 0.20
0.15
0.325 0.20
0.20
0.350
0.25
shaded band: 95% CI = ±1.96 SEM
2
4
shaded band: 95% CI = ±1.96 SEM
6
8
(a) GPT-2
10
12
5
10
0.15
shaded band: 95% CI = ±1.96 SEM
15
20
25
(b) GPT-2 Medium
5
10
15
20
25
30
(c) GPT-2 Large
35
shaded band: 95% CI = ±1.96 SEM
10
20
30
40
50
(d) GPT-2 XL
Figure 10: GPT-2 family [16] on CBT story-token prompts. The same middle-layer movement toward a uniform spherical reference distribution appears for natural text prompts.
24
0.40
0.35 0.40
Cbt-Ne Stories
Cbt-Ne Stories
Cbt-Ne Stories
0.35
Cbt-Ne Stories
Cbt-Ne Stories
0.225
0.35
0.35
0.30
0.30 0.200
0.30
0.30
0.25
0.25
0.175 0.25
0.25 0.20
0.150 0.20
0.20
0.20 0.125
0.15
0.15
0.10
0.10
0.15 0.15
0.100 0.10 0.075 0.05
shaded band: 95% CI = ±1.96 SEM
5
10
0.05 15
20
(a) 1.5B
25
shaded band: 95% CI = ±1.96 SEM
5
10
15
0.10 shaded band: 95% CI = ±1.96 SEM
20
25
30
(b) 3B
35
5
10
shaded band: 95% CI = ±1.96 SEM
15
(c) 7B
20
25
10
20
0.05 30
(d) 14B
40
50
shaded band: 95% CI = ±1.96 SEM
10
20
30
40
50
60
(e) 32B
Figure 11: Qwen2.5 family [20] on CBT story-token prompts. The U-shaped profile becomes visible across model scales, indicating a middle-layer movement toward a uniform spherical reference distribution followed by a later movement away from it. D.5
Computational Resources
All experiments were run on a single NVIDIA H100 GPU machine with 128 GB of system memory. The controlled synthetic experiments in Appendix D.1 and D.2 are computationally lightweight: they use a small GPT-2-style decoder-only transformer. If training is mentioned, the model is trained for 2,000 optimization steps on 1,000 IID training sequences. For the pretrained-language-model analyses in Appendix D.3, no model fine-tuning is performed. The experiments require only forward passes through publicly available autoregressive language models in order to extract hidden states after each transformer block. D.6
Existing Assets and Licenses
We use publicly available pretrained language models only for evaluation and do not redistribute their weights. For each model family, we cite the original paper, technical report, or official release, and follow the license or model-card terms associated with the official release. The evaluated models include GPT-2 [16], Llama/Llama-3.3 [4, 17], OpenLLaMA/OpenLM-Llama, Cerebras-GPT [19], Qwen2.5 [20], Falcon [21], OPT [22], and BLOOM [23]. The CBT dataset is cited as [35]. License and access terms are taken from the corresponding official repositories or model cards.
25
Table 1: Existing pretrained models and datasets used in the experiments. We use these assets only for evaluation and do not redistribute model weights or dataset copies. License information is taken from the corresponding official release pages or model cards. Asset
Use in this paper
Source / citation
GPT-2
Pretrained autoregressive language model for layerwise hidden-state analysis
OpenAI release; as [16]
Llama / Llama-3.370B-Instruct
Pretrained autoregressive language model for semantic-topic and CBT experiments
Meta Llama releases; cited as [4, 17]
OpenLLaMA / OpenLM-Llama-13B
Pretrained autoregressive language model for semantic-topic experiments
OpenLLaMA/OpenLM Research official release or model card; cited as [18]
Cerebras-GPT-13B
Pretrained autoregressive language model for semantic-topic and CBT experiments Pretrained autoregressive language models for model-scale comparison
Cerebras-GPT cited as [19]
Falcon-40B
Pretrained autoregressive language model for additional model-family experiments
Falcon as [21]
OPT-66B
Pretrained autoregressive language model for additional model-family experiments
OPT release; cited as [22]
BLOOM-7B1
Pretrained autoregressive language model for additional model-family experiments Natural-text prompts for layerwise MMD experiments
BigScience BLOOM release; cited as [23]
Qwen2.5 family
CBT dataset
26
License or terms cited
release;
Qwen release; cited as [20]
release;
cited
Children’s Book Test; cited as [35]
MIT / Modified MIT license, according to the official model release or model card. Meta Llama community/model license; access and use governed by the official Llama license terms. Apache-2.0 or permissive open-source license, according to the official model card. Apache-2.0 license. Apache-2.0 for most open-source Qwen2.5 variants; some variants, including 3B and 72B, use modelspecific Qwen license terms. Apache-2.0 license, according to the official release/model card. Meta OPT license terms, according to the official model card or license file. BigScience RAIL License v1.0. Dataset release terms associated with the original CBT release; the dataset is used only for evaluation.