Unifying ICL, SFT, KL-Regularized RL Through a Bayesian Lens Junxin Fan Fudan University [email protected]
arXiv:2609.05111v1 [cs.AI] 4 Sep 2026
September 7, 2026
Abstract Large language models are now trained and evaluated under a diverse set of paradigms: supervised fine-tuning (SFT), few-shot in-context learning (ICL), KL-regularized RLHF/RLVR, on-policy distillation (OPD), and test-time reasoning with search and chain-of-thought. These methods are often discussed as fundamentally different, and recent empirical results—such as the mixed impact of few-shot prompting on RL-tuned reasoning models—can appear puzzling. This note develops a unified Bayesian perspective that puts these procedures on the same footing. At the core is a two-step template: (i) construct a (generalized) Bayes or Gibbs posterior q ∗ over outputs or actions given a context, using a prior/reference model and a utility signal (log-likelihood, reward, or advantage); and (ii) approximate q ∗ by a forward-KL projection onto a parametric family, either in-weights (SFT/RL) or in-context (ICL). Part I formalizes fewshot ICL and SFT as amortized and in-weights projections onto the Bayes posterior predictive. Parts II–IV show that KL-regularized RLHF/RLVR, reward- weighted SFT, reward-weighted ICL (RW-ICL), and advantage-weighted SFT (AWSFT) are all instances of forward-KL projection onto Gibbs posteriors induced by rewards or advantages. We disentangle where these equivalences hold (objectives and first-order updates) and where they do not (source and granularity of the learning signal). Part V sketches implications for modern reasoning pipelines. We view RLHF/RLVR recipes as “posterior design + projection”, explain why some form of cold-start or supervised warm-up is practically unavoidable for importance-weighted KL projections—including on-policy distillation and self-training pipelines—through a support-overlap and importance-weighting view, and interpret DeepSeek-R1 and o1-style reasoning models as combining test-time Bayesian search with training-time KL amortization. Together, these pieces suggest a coherent Bayesian/KLprojection backbone behind many current post-training strategies for reasoning LMs.
Keywords: in-context learning; supervised fine-tuning; KL-regularized reinforcement learning; Gibbs posterior; Bayesian meta-learning
1
Part I: Few-shot ICL ≈ SFT (Bayesian View) 1.1 Supervised Fine-Tuning (SFT) as Maximum Likelihood Given a dataset D = {(xi , yi )}N i=1 , standard supervised fine-tuning maximizes the conditional loglikelihood: N X θSFT = arg max log pθ (yi | xi ). (1) θ
i=1
Equivalently, define the empirical data distribution N
p̂data (x, y) =
1 X δ(xi ,yi ) (x, y), N
(2)
i=1
so that the SFT objective can be written as θSFT = arg min E(x,y)∼p̂data − log pθ (y | x) . θ
(3)
Bayesian / KL interpretation. Assume the data are generated from an unknown teacher conditional distribution qtrue (y | x), and that p̂data is an empirical approximation to the joint qtrue (x, y) = qtrue (x) qtrue (y | x). Then E(x,y)∼qtrue − log pθ (y | x) = Ex∼qtrue (x) Ey∼qtrue (·|x) − log pθ (y | x) (4) h i (5) = Ex∼qtrue (x) H qtrue (· | x) + KL qtrue (· | x) ∥ pθ (· | x) , Here we expand the inner expectation using the standard cross-entropy decomposition H(q, p) = H(q) + KL(q∥p) applied to qtrue (· | x) and pθ (· | x), where H(·) is conditional entropy. The first term is independent of θ, so minimizing expected negative log-likelihood is equivalent to θ∗ = arg min Ex∼qtrue (x) KL qtrue (· | x) ∥ pθ (· | x) . (6) θ
In words, SFT learns pθ as the forward KL (maximum-likelihood) projection of the teacher qtrue onto the model family {pθ }.1
1.2 Hierarchical Bayesian model for tasks and contexts We now introduce a hierarchical Bayesian model for tasks, following the “ICL is Bayes” framework of Wakayama and Suzuki [24].
Task-level generative model. sampled as
Let F be a space of functions f : X → ∆(Y). A random task is f ∼ P (f ),
1
(7)
Here “projection” refers to minimizing the forward KL KL(qtrue ∥pθ ), also known as MLE or the m-projection in information geometry.
2
and conditioned on f , data pairs (x, y) are generated i.i.d. by x ∼ PX ,
y ∼ f (· | x).
(8)
For a given task f , a few-shot context (prompt) and a query point are generated as Dk = {(xi , yi )}ki=1 ,
(xk+1 , yk+1 ) ∼ PX × f (· | xk+1 ).
(9)
Bayes posterior over tasks and posterior predictive. Given context Dk , the posterior over tasks is k Y P (f | Dk ) ∝ P (f ) f (yi | xi ), (10) i=1
and the corresponding Bayes posterior predictive at a new query xk+1 is PBayes (y | xk+1 , Dk ) := Ef ∼P (f |Dk ) f (y | xk+1 ) .
(11)
Bayes-optimal in-context predictor. Consider any in-context prediction rule M that maps (Dk , xk+1 ) to a predictive distribution M (Dk , xk+1 ) ∈ ∆(Y). Under log-loss, define the ICL risk R(M ) := E − log M (Dk , xk+1 )(yk+1 ) , (12) where the expectation is over the hierarchical model above (sampling f , then Dk and (xk+1 , yk+1 )). Theorem 1 (Bayes-optimal in-context predictor; after Wakayama and Suzuki [24]). Define the Bayes posterior predictive MBayes (Dk , xk+1 ) := PBayes (· | xk+1 , Dk ).
(13)
Then: (i) MBayes is the unique minimizer of R(M ) over all measurable predictors M . (ii) For any M , we define its Bayes Gap as BayesGap(M ) := R(M ) − R(MBayes ).
(14)
In particular, the excess risk satisfies R(M ) − R(MBayes ) = BayesGap(M ).
(15)
In the squared-loss setting of Wakayama and Suzuki [24], this quantity coincides with their Bayes Gap term RBG (M ) in the decomposition R(M ) = RBG (M ) + RPV , and R(MBayes ) corresponds to their Posterior Variance component RPV . The full non-asymptotic bounds for the Bayes Gap and Posterior Variance are given in Wakayama and Suzuki [24]; here we only use that MBayes is the normative target of in-context prediction. This hierarchical model can be justified from a within-task exchangeability assumption via a de Finetti–type representation; see Appendix A.4. 3
Remark (Order of demonstrations). Real transformers process ordered sequences with positional embeddings, so the order of few-shot demonstrations in the prompt can affect predictions. Our exchangeability assumption applies at the level of the data-generating process: within a task, the examples (xi , yi ) are assumed to form an exchangeable sample from an underlying task distribution, so the Bayes posterior P (f | Dk ) and posterior predictive PBayes (y | xk+1 , Dk ) depend only on the multiset {(xi , yi )}, not on their order. Part I analyzes this ideal, order-symmetric Bayes predictor; practical in-context learners may introduce position-dependent biases, which we view as algorithmic deviations from this ideal rather than violations of the assumption.
1.3 In-context learning as amortized Bayesian inference Let Mθ (y | x, Dk ) denote the conditional distribution implemented by a pretrained Transformer with parameters θ. Pretraining across many tasks and contexts can be abstracted as minimizing the expected log-loss (16) R(θ) := E − log Mθ (Dk , xk+1 )(yk+1 ) , where the expectation is over the same hierarchical process as in Section 1.2. By Theorem 1, the risk-minimizing population predictor is MBayes . Thus, in the limit of infinite model capacity and pretraining data, any minimizer Mθ∗ of R(θ) must satisfy Mθ∗ (Dk , xk+1 ) ≈ MBayes (Dk , xk+1 ) = PBayes (· | xk+1 , Dk ).
(17)
Formally, under the assumption in Sec. 1.5 that the model class {Mθ } is rich enough to approximate MBayes , any global minimizer (or sequence of asymptotic minimizers) of R(θ) can be chosen to be arbitrarily close to MBayes in expected log-loss. This viewpoint matches recent constructions that explicitly train Transformers as amortized Bayesian inference procedures over latent variables and datasets (e.g., [22]), where a single forward pass maps each dataset D to an approximate posterior.
Interpretation.
In this Bayesian picture, few-shot ICL is best viewed as learning the mapping (Dk , xk+1 ) 7−→ PBayes (· | xk+1 , Dk )
(18)
in an amortized way. The Transformer does not explicitly perform gradient descent in parameter space at test time; instead, it approximates the Bayes posterior predictive directly in its forward pass. While our exposition centers on the Bayesian view, a complementary first-order / metalearning interpretation (ICL ≈ one GD step) has been studied extensively in prior work [5, 3, 13], which shows that Transformers can implement gradient-based updates in their forward passes. We do not re-derive those results here, as our goal is to foreground the Bayesian formulation.
1.4 From Bayes posterior to SFT: forward-KL projection over outputs We now connect the Bayes posterior predictive to supervised fine-tuning. Fix a context Dk and query xk+1 , and write qBayes (y | xk+1 , Dk ) := PBayes (y | xk+1 , Dk ) 4
(19)
for the ideal target distribution over outputs. Consider any parametric conditional model pθ (y | x) (suppressing Dk in notation). For this fixed pair (xk+1 , Dk ), the best approximation to qBayes within {pθ } under log-loss is θ∗ (xk+1 , Dk ) = arg min KL qBayes (· | xk+1 , Dk ) ∥ pθ (· | xk+1 ) , (20) θ
or equivalently θ∗ (xk+1 , Dk ) = arg max θ
X
qBayes (y | xk+1 , Dk ) log pθ (y | xk+1 ).
(21)
y
If we had direct samples y ∼ qBayes (· | xk+1 , Dk ) for many (xk+1 , Dk ), then performing SFT on those samples would be exactly the forward KL projection of the Bayes posterior predictive onto the family {pθ }. Aggregating over contexts and queries yields the population objective min E(Dk ,xk+1 ) KL qBayes (· | xk+1 , Dk ) ∥ pθ (· | xk+1 ) , (22) θ
and the standard SFT objective is its empirical approximation. Bayesian reading of “ICL ≈ SFT”.
Putting Sections 1.2-1.4 together:
• The ideal in-context predictor is the Bayes posterior predictive qBayes (· | xk+1 , Dk ). • A sufficiently expressive Transformer minimizing ICL risk implements an amortized approximation Mθ (· | xk+1 , Dk ) to qBayes . • Supervised fine-tuning (SFT) learns pθ as the forward KL projection of qBayes (or of an empirical approximation) onto the parametric family. Thus, in the Bayesian view, few-shot ICL ≈ Bayes posterior predictive ≈ KL-projected SFT solution,
(23)
and ”ICL ≈ SFT” is understood as both being different approximations to the same underlying Bayesian predictor. This distribution-level identification is the backbone for Part II, where KL-regularized RL (RLHF/RLVR) is shown to define a generalized Gibbs posterior over outputs, and reward-weighted SFT is exactly its KL projection onto {pθ }.
1.5 Assumptions and scope Regularity for Part I (Bayesian ICL and SFT). We assume: (i) a well-defined hierarchical prior P (f ) and data-generating mechanism as in Section 1.2 (this representation follows from a within-task exchangeability assumption; see Appendix A.4); (ii) log-loss and finite Y (or densities with respect to a common base measure); (iii) integrability conditions ensuring Fubini/Tonelli can be applied when exchanging expectations and logs; (iv) a model class {Mθ } rich enough that some Mθ∗ can approximate MBayes . The non-asymptotic generalization guarantees (Bayes Gap bounds, posterior variance behavior under task mixtures, and out-of-distribution analyses) are from Wakayama and Suzuki [24]. 5
Regularity for Part II (KL-RL). We assume β > 0; rewards r are bounded above so that exp(r/β) is integrable; and shared support: πref (y | x) > 0 whenever π(y | x) > 0. Interchange of ∇ and expectation is justified by dominated convergence. Regularity for Part IV (process-level). We assume a stationary MDP; well-defined Q/A functions; and advantage estimates (e.g., GAE) with bounded bias/variance. These are standard in KL-regularized policy optimization and RL-as-inference formulations. Remark (Meta-learning view as complementary, not primary). The original meta-learning perspective that ICL behaves like a single small gradient step on a contextual loss can be rigorously justified under additional smoothness and linearization assumptions, and connects attention to accumulated gradient updates. We treat that interpretation as complementary (see the discussion and citations summarized in Appendix A.3), while the main text takes the Bayesian formulation as the primary conceptual and mathematical foundation.
Part II: KL-Regularized RL (RLHF/RLVR) ↔ Reward-weighted SFT 2.1 KL-regularized RL objective and closed-form optimum Fix an input x (e.g., a prompt), a reference policy πref (· | x), a reward function r(· | x), and a temperature β > 0. The standard KL-regularized objective is max Ey∼π(·|x) [r(y | x)] − β KL π(· | x) ∥ πref (· | x) . (24) π
This form appears both in RLHF (human preference reward) and in RLVR (verifiable/correctness reward); the only difference is the source of r(y | x) [18, 21, 15, 25]. In practice, many RLHF-style algorithms (e.g., PPO-style KL-penalized RL, GRPO’s effective loss, and RLVR-style updates) optimize surrogate objectives that approximate (24) rather than this exact form. In what follows we work at the level of the underlying KL-regularized objective (24) and abstract away such implementation details. For discrete Y, we can write the objective as X X π(y | x) J(π) = π(y | x) r(y | x) − β π(y | x) log , (25) πref (y | x) y y P with the constraint y π(y | x) = 1. Introducing a Lagrange multiplier λ for this constraint, the Lagrangian is ! X X X π(y | x) L(π, λ) = π(y | x)r(y | x) − β π(y | x) log +λ 1− π(y | x) . (26) π (y | x) ref y y y Proposition 2 (KL-regularized RL ⇒ Gibbs posterior). The unique maximizer of (24) is X πref (y | x) exp r(y | x)/β ∗ π (y | x) = , Z(x) = πref (y | x) exp r(y | x)/β . (27) Z(x) y 6
Proof.
Take derivative of L w.r.t. π(y | x): ∂L π(y | x) = r(y | x) − β log + 1 − λ. ∂π(y | x) πref (y | x)
(28)
Setting ∂L/∂π(y | x) = 0 gives log
π(y | x) 1 = r(y | x) − λ − β . πref (y | x) β
(29)
Exponentiating both sides, λ+β r(y | x) exp − . (30) π(y | x) = πref (y | x) exp β β P The last factor does not P depend on y and is determined by normalization y π(y | x) = 1, which yields (27) with Z(x) = y πref (y | x) exp(r(y | x)/β). □ P Because the negative entropy term − y π(y | x) log π(y | x) is strictly concave in π(· | x), the objective in (24) is strictly concave with respect to π, guaranteeing that the stationary point above is the unique maximizer. Thus, KL-regularized RLHF/RLVR produces a Gibbs posterior over outputs: a generalized Bayes update with prior πref (· | x) and pseudo-likelihood proportional to exp(r(y | x)/β). This structure is classical in relative-entropy policy search and RL-as-inference formulations [20, 1, 12].
2.2 Reward-weighted SFT as KL projection of the Gibbs posterior We now show that reward-weighted SFT is exactly the forward-KL projection of π ∗ onto a parametric family {πθ }. Fix x and write qGibbs (y | x) := π ∗ (y | x) =
πref (y | x) exp(r(y | x)/β) . Z(x)
(31)
(This qGibbs is distinct from the teacher distribution qtrue used in Section 1.1.) Consider a parametric policy πθ (y | x) and the KL projection problem θ∗ = arg min KL qGibbs (· | x) ∥ πθ (· | x) . (32) θ
Throughout, when we write f (θ) ∝ g(θ) we mean that f (θ) = c + g(θ) for a constant c independent of θ; i.e., we drop additive terms that do not affect the optimizer. Expanding the KL, X qGibbs (y | x) KL qGibbs ∥πθ = qGibbs (y | x) log πθ (y | x) y X X = qGibbs (y | x) log qGibbs (y | x) − qGibbs (y | x) log πθ (y | x). y
|
y
{z
}
indep. of θ
7
(33) (34)
Therefore θ∗ = arg max θ
X
qGibbs (y | x) log πθ (y | x).
(35)
y
Substituting the Gibbs form of qGibbs , X y
1 X r(y | x) qGibbs (y | x) log πθ (y | x) = log πθ (y | x). πref (y | x) exp Z(x) y β
(36)
The normalizer Z(x) does not depend on θ, so maximizing the above is equivalent to maximizing X r(y | x) JRWSFT (θ; x) := πref (y | x) exp log πθ (y | x). (37) β y Proposition 3 (Reward-weighted SFT = forward-KL projection). For fixed x, any minimizer of (32) can be obtained by maximizing the reward-weighted SFT objective (37), and any maximizer of (37) minimizes (32). In particular, reward-weighted SFT learns πθ as the forward KL projection of the KL-regularized RL optimum π ∗ (i.e., qGibbs ) onto the parametric family {πθ }. In practice, one does not enumerate all y, but instead samples (x, y) from a data distribution whose y-marginal matches πref (· | x) (e.g., rollouts from the reference model or logged data), and forms the empirical objective X r(y | x) log πθ (y | x), (38) max exp θ β (x,y)
which is the familiar reward-weighted SFT loss used in many recent RLHF / RLVR formulations. Summary for Part II. KL-regularized RLHF/RLVR has a closed-form optimum π ∗ that is a Gibbs posterior (27). Reward-weighted SFT is exactly the KL projection of this posterior onto {πθ }. Thus, at the distribution level, KL-regularized RL (RLHF/RLVR) ≃ reward-weighted SFT.
(39)
This equivalence is meant at the level of the target Gibbs posterior qGibbs and its forward-KL projection; it does not claim that all practical RLHF algorithms exactly attain the closed-form optimum π ∗ in (27). This matches recent variational interpretations of RLHF that treat it as reward-weighted regression.
Part III: Reward-weighted Few-shot ICL (RW-ICL) We now lift the discussion from single (x, y) decisions to few-shot in-context prediction with rewards.
3.1 Generalized Bayesian posterior over outputs given context Let Dk = {(xi , yi )}ki=1 be a context, and xk+1 a query. As in Part I, the Bayes posterior predictive under a hierarchical task model is PBayes (y | xk+1 , Dk ) = Ef ∼P (f |Dk ) [f (y | xk+1 )]. 8
(40)
For our RLHF/RLVR setting, we augment this with a reward signal. Assume that for each potential output y at (xk+1 , Dk ) we have a scalar score R(y; Dk , xk+1 ),
(41)
which aggregates the relevant reward information (from human preferences, a verifier, or a reward model) given the entire context Dk and query xk+1 . We define a generalized Bayes posterior over outputs: R(y; Dk , xk+1 ) , qGibbs (y | xk+1 , Dk ) ∝ πref (y | xk+1 ) exp β
(42)
i.e., a Gibbs posterior with prior πref (· | xk+1 ) and pseudo-likelihood proportional to exp(R/β). This is the natural extension of Proposition 2 to the context-dependent case.
3.2 RW-ICL objective as KL projection of the generalized posterior Let Mθ (y | xk+1 , Dk ) denote the in-context predictor implemented by a meta-trained Transformer. A natural population objective for reward-weighted ICL is min E(Dk ,xk+1 ) KL qGibbs (· | xk+1 , Dk ) ∥ Mθ (· | xk+1 , Dk ) , (43) θ
where the expectation is taken over the same hierarchical meta-distribution on tasks and contexts as in Part I, together with whatever process generates R(y; Dk , xk+1 ). Expanding the KL and discarding terms independent of θ, (43) is equivalent to maximizing E(Dk ,xk+1 ) Ey∼qGibbs (·|xk+1 ,Dk ) log Mθ (y | xk+1 , Dk ) .
(44)
Substituting the Gibbs form of qGibbs from (42), E(Dk ,xk+1 ) Ey∼qGibbs [log Mθ (y | xk+1 , Dk )] X R(y; Dk , xk+1 ) ∝ E(Dk ,xk+1 ) πref (y | xk+1 ) exp log Mθ (y | xk+1 , Dk ). β y
(45) (46)
In practice, one approximates this expectation empirically by sampling contexts Dk , queries xk+1 , and candidate outputs y from the reference model (or from a replay buffer), and forming the empirical objective X R(y; Dk , xk+1 ) max exp log Mθ (y | xk+1 , Dk ), (47) θ β (Dk ,xk+1 ,y)
i.e., a reward-weighted in-context SFT objective. Proposition 4 (RW-ICL = KL projection of generalized Gibbs posterior). Under mild regularity conditions, any minimizer Mθ∗ of (43) is the forward KL projection of the generalized Gibbs posterior qGibbs (· | xk+1 , Dk ) onto the model family {Mθ (· | xk+1 , Dk )}. Stochastic gradient ascent on (47) implements this projection using samples from πref reweighted by exp(R/β). 9
Gradient-level correspondence. Although our exposition so far has been distributional (KL projections of Gibbs posteriors), the resulting stochastic-gradient updates have exactly the rewardweighted form emphasized in meta-learning views of ICL. For example, the empirical RW-ICL objective (47) has gradient X R(y;Dk ,xk+1 ) exp ∇θ LRW-ICL (θ) = − ∇θ log Mθ (y | xk+1 , Dk ), (48) β (Dk ,xk+1 ,y)
so a single SGD step is a reward-weighted in-context update. The reward-weighted SFT and AWSFT losses (37) and (59) yield the same update structure with Mθ replaced by πθ and R replaced by r or A. Thus, at the level of first-order optimization dynamics, RW-ICL, reward-weighted SFT, and KL-regularized RL share the same summed-score ∇θ log(·) weighted by exp(reward/β); the meta-learning results of Dai et al. [5], Akyürek et al. [3], Li et al. [13] can be viewed as constructive realizations of such updates in a single forward pass.
Relation to Part II. Comparing (42) to (27), we see that KL-regularized RLHF/RLVR defines the same Gibbs-family posterior over outputs, with R(y; Dk , xk+1 ) playing the role of r(y | x). Reward-weighted ICL meta-trains a Transformer to amortize the mapping (Dk , xk+1 ) 7−→ qGibbs (· | xk+1 , Dk ),
(49)
just as ordinary ICL in Part I amortizes the Bayes posterior predictive. Hence, at the distribution level, reward-weighted ICL (RW-ICL) ≈ KL-regularized RLHF/RLVR optimum, (50) up to approximation error in the KL projection.
Part IV: Process-level KL-RLVR and Advantage-weighted SFT We finally consider the full RL setting with trajectories and state-dependent policies, which is the regime of KL-RLVR and related algorithms.
4.1 Statewise KL-regularized objective and Gibbs policy Let s be a state and π(· | s) a stochastic policy over actions a. A typical stepwise KL-regularized objective is X max π(a | s)Q(s, a) − β KL π(· | s) ∥ πref (· | s) , (51) π
a
where Q(s, a) is an action-value function (for example, a critic or Monte Carlo return), and πref is a reference policy. This is the policy-improvement subproblem solved in many KL-regularized policy search methods. By the same Lagrangian argument as in Proposition 2, the unique maximizer is π ∗ (a | s) =
πref (a | s) exp(Q(s, a)/β) , Z(s)
Z(s) =
X a
10
πref (a | s) exp(Q(s, a)/β).
(52)
This is again a Gibbs posterior, now over actions conditional on state s, with log-likelihood Q(s, a)/β. It is often convenient to reparameterize in terms of the advantage X πref (a | s)Q(s, a), A(s, a) := Q(s, a) − V (s), V (s) :=
(53)
a
which only shifts Q by a baseline independent of a. Any such state-dependent baseline contributes a multiplicative factor exp(−V (s)/β) that is absorbed into the partition function Z(s), leaving the normalized policy unchanged. Hence, π ∗ (a | s) ∝ πref (a | s) exp(A(s, a)/β).
(54)
4.2 Advantage-weighted SFT (AWSFT) as KL projection Fix a state s and define qGibbs (a | s) := π ∗ (a | s) given by (52) or (54). For a parametric policy πθ (a | s), the KL projection problem is θ∗ (s) = arg min KL qGibbs (· | s) ∥ πθ (· | s) . θ
(55)
(56)
Expanding as before, this is equivalent to θ∗ (s) = arg max θ
X
qGibbs (a | s) log πθ (a | s).
(57)
a
Substituting the Gibbs form with advantages, X X A(s, a) qGibbs (a | s) log πθ (a | s) ∝ πref (a | s) exp log πθ (a | s). β a a Thus, across states s, the population objective is equivalent to maximizing X A(s, a) JAWSFT (θ) := Es∼d(s) πref (a | s) exp log πθ (a | s), β a
(58)
(59)
where d(s) is some state visitation distribution (e.g., from the reference or a replay buffer). In practice, we approximate (59) by empirical sums over state–action pairs (st , at ) with estimated advantages Â(st , at ): ! X Â(st , at ) log πθ (at | st ), (60) max exp θ β t which is exactly an advantage-weighted SFT (AWSFT) loss. This matches the advantage-weighted regression family of algorithms [19, 16] and entropy-regularized methods such as SAC [8], which all rely on exponentiated Q/A weighting. Proposition 5 (AWSFT = KL projection of statewise KL-RLVR optimum). For each state s, any minimizer of KL(qGibbs (· | s) ∥ πθ (· | s)) is a maximizer of the local AWSFT objective, and any maximizer of that objective minimizes the KL. Consequently, the global AWSFT objective (59) implements the forward KL projection of the stepwise KL-regularized RL optimum (52) onto the parametric policy family {πθ }. 11
4.3 Process-level RW-ICL alignment From the in-context perspective, a trajectory (or multi-step interaction) provides a context D = {(st , at , rt )}Tt=1 ,
(61)
and the goal is to predict or select future actions in-context. A reward-weighted in-context learner can be meta-trained to approximate, at each state s, the Gibbs posterior qGibbs (· | s) defined by the (advantage-based) KL-regularized RL objective: A(s, a) qGibbs (a | s) ∝ πref (a | s) exp . (62) β The corresponding RW-ICL meta-objective is the trajectory-level analogue of (43), with per-step pseudo-likelihoods proportional to exp(Â(st , at )/β) and model predictions Mθ (at | st , D). In particular, performing stochastic gradient ascent on this objective yields updates of exactly the same form as those obtained from the advantage-weighted SFT objective ! X Â(st , at ) JAWSFT (θ) := exp log Mθ (at | st , D), (63) β t up to choices of state distribution and context window. Thus, process-level RW-ICL and stepwise KL-RLVR share the same Gibbs-posterior backbone and are both realized as advantage-weighted SFT in parameter space.
4.4 Learning signals and credit assignment in SFT, RL, and ICL The equivalence results in Parts II and III operate at the level of objectives and first-order updates: under a suitable Gibbs posterior, KL-regularized RL, reward-weighted SFT, and reward-weighted ICL all reduce to minimizing a forward KL KL(q ∗ ∥ pθ ), and their stochastic gradients share the familiar “weighted score” form. Writing a generic forward-KL loss as X L(θ) = wt − log pθ (zt ) , (64) t
its gradient takes the form b θ L(θ) = − ∇
X
wt ∇θ log pθ (zt ).
(65)
t
for appropriate choices of samples zt and weights wt . However, as pointed out in the feedback, the source and granularity of the learning signal differ substantially across paradigms. In this section we disentangle these differences and show how they fit into our unified view.
Dense token-level supervision in SFT. Under standard supervised fine-tuning, the model is trained with a token-level cross-entropy loss. For an autoregressive language model with inputs x and target sequence y = (t1 , . . . , tn ), the SFT objective can be written as " n−1 # X LSFT (θ) = −E(x,y) log pθ ti+1 x, t≤i . (66) i=0
12
Equivalently, SFT learns from every mapping (x, y≤i ) 7→ yi+1 : each prefix within a sequence contributes a term to the loss and a gradient to the update. The learning signal is thus dense at the token level, and credit assignment is implicit in the data: human- or model-generated targets (x, y) already specify which next-token prediction should be encouraged at each position.
KL-regularized RL: sparse rewards + explicit credit assignment. In contrast, KL-regularized RL (including RLHF/RLVR-style objectives) typically starts from a sequence- or trajectory-level reward. In the sequence-level setting of Part II, the objective has the form h i JRL (π) = Ex∼D, y∼π(·|x) R(y | x) − β KL π(· | x) ∥ πref (· | x) , (67) and the naive policy gradient estimator looks like h i b θ JRL ≈ Ex,y∼π R(y | x) ∇θ log πθ (y | x) . ∇ θ
(68)
At first sight, this appears to support the claim that “RL can only learn once from the mapping x 7→ R(y)”: a single scalar reward per sampled sequence. In practice, however, modern RL algorithms do not update the policy directly from (68). Instead, they introduce a credit assignment step, in which the scalar reward is decomposed into per-step signals. In the step-wise KL-regularized MDP setting of Part IV, this yields a Gibbs-optimal policy π(a | s) ∝ πref (a | s) exp A(s, a)/β , and the corresponding forward-KL projection leads to the advantage-weighted SFT objective (cf. Eq. (59)–(60)): h i LAWSFT (θ) = −E(s,a) exp A(s, a)/β log πθ (a | s) . (69) The resulting policy gradient has exactly the weighted-score form (65): b θ LAWSFT = − ∇
X
exp
A(s b , a ) t
β
t
t
∇θ log πθ (at | st ).
(70)
Thus, while the raw RL signal is a single scalar reward per trajectory, RL algorithms recover the ability to “learn from each (st , at )” by explicitly estimating value functions V (s), action-values Q(s, a), or advantages A(s, a) and using them as per-step weights. The final policy update becomes a form of advantage- or reward-weighted SFT. In our framework, this is precisely the AWSFT/RWSFT connection established in Parts II and IV: at the level of updates, KL-regularized RL is equivalent to a particular choice of weights wt in (65). Few-shot ICL: dense supervision in pretraining, fast learning in activations. Few-shot in-context learning sits somewhat orthogonally to this discussion. As formalized in Part I, we view a parametric ICL predictor Mθ (y | x, Dk ) as an amortized Bayes approximation to the posterior predictive under a task prior P (f ): Z PBayes (y | x, Dk ) = pf (y | x) p(f | Dk ) df.
13
The parameters θ are learned using the same dense token-level supervision as in SFT: for example, the outer objective h i LICL (θ) = E(f,Dk ,xk+1 ,yk+1 ) − log Mθ yk+1 | xk+1 , Dk (71) is a standard log-loss aggregated over tasks and positions. In this sense, few-shot ICL does not introduce a new type of training signal: it consumes the same token-level (possibly reward-weighted) supervision as SFT/RW-ICL in the outer loop. The key difference is where and when adaptation happens. At test time, few-shot ICL does not take further gradients on θ. Instead, given a context Dk , the transformer uses attention and MLP activations to construct a task-specific internal state ϕθ (Dk ) in its forward pass, and the effective predictor for the query can be written schematically as Mθ (y | xk+1 , Dk ) ≈ fϕθ (Dk ) (y | xk+1 ) for some implicit inner model class {fϕ }. Recent theoretical and mechanistic studies have shown that, in controlled settings, the fast state ϕθ (Dk ) closely matches the output of explicit learning algorithms such as least squares or gradient descent on a task-specific inner objective; in this sense, ICL learning happens in activations rather than through online updates of θ.
Summary. Putting these pieces together, our framework unifies SFT, KL-regularized RL, and few-shot ICL at the level of target posteriors and first-order updates: all three aim to approximate a Bayes or Gibbs posterior, and their gradients can be written as weighted score sums of the form (65). At the same time, the origin and granularity of the learning signal differ: • SFT observes dense token-level supervision: every prefix (x, y≤i ) has an explicit label yi+1 , and credit assignment is implicit in the data. • KL-regularized RL starts from sparse sequence- or trajectory-level rewards R(y); it must perform explicit credit assignment via Q/A estimates, but the resulting policy update is an advantage- or reward-weighted SFT step. • Few-shot ICL relies on dense supervision during pretraining to learn θ as a meta-learner; at test time, it performs task-specific adaptation in the forward activations ϕθ (Dk ) without additional gradients. In this sense, our claimed equivalences are intentionally restricted to the level of objectives (posterior distributions) and update forms. They do not erase the very real differences in how SFT, RL, and ICL acquire and use their learning signals—differences which have important implications for sample efficiency, stability, and the kinds of structure (e.g., credit assignment and reasoning) each paradigm can readily capture.
14
Part V: Implications for Reasoning Models and Training Recipes 5.1 Posterior design and practical RLHF/RLVR recipes Parts II and IV showed that a broad class of KL-regularized RLHF/RLVR objectives admits a Gibbs-posterior optimum qGibbs (y | x) ∝ πref (y | x) exp(R(y; D, x)/β) and that reward- or advantage-weighted SFT is exactly the forward-KL projection of this posterior onto a parametric policy family. This matches a growing line of recent work that re-formulates RLHF as rewardweighted regression or weighted SFT rather than generic policy gradient [e.g. 11, 9, 6].2 From our perspective, practical RLHF/RLVR pipelines can be viewed as two-step procedures: 1. Posterior design. Choose a reference model πref and a reward or advantage signal R, which together define a target Gibbs posterior qGibbs (y | x) ∝ πref (y | x) exp(R(y; D, x)/β). 2. Projection. Fit a deployable policy πθ by minimizing the forward KL KL(qGibbs ∥πθ ), either directly via reward-/advantage-weighted SFT, or indirectly via a PPO-style optimizer that approximates the same projection. This view highlights three concrete design knobs: • Reward shaping as likelihood design. The reward R enters only through exp(R/β), so shifting and scaling R correspond to changing the effective likelihood ratio between highand low-quality outputs. Overly sharp or poorly calibrated rewards lead to posteriors with very low support overlap with πref , which our framework predicts to cause instability and over-optimization. • Temperature β as posterior concentration. Small β produces a highly concentrated qGibbs that strongly prefers top-reward outputs; large β keeps qGibbs close to πref . In our KL-projection picture, β controls how much the optimizer is allowed to move the policy away from the reference model before running into support-mismatch issues. • Offline vs on-policy projection. When rollouts come from πref or an early-stage model, the reward-weighted SFT losses in Parts II and IV implement an off-policy projection to the Gibbs posterior. PPO-style RLHF instead approximates the same target distribution with on-policy samples. Our analysis suggests that both are instances of the same forward-KL projection, but they trade off exploration (on-policy) against stability and reuse of pre-generated data (offline or replay-based).
5.2 Cold start and on-policy distillation Our KL-projection view also clarifies why recent pipelines based on on-policy distillation and iterative self-training almost always include a small supervised “cold start” phase before running high-variance on-policy updates. 2
Here we use shorthand placeholders for recent RLHF analyses that interpret PPO-style RLHF as approximately minimizing a forward-KL to a reward-weighted reference distribution, in line with our Propositions 2–5.
15
Empirically, several lines of work report that purely on-policy optimization from a random or weakly aligned prior tends to fail: DeepSeek-R1 distinguishes a pure-RL variant (R1-Zero) that exhibits severe style and safety issues compared to the full R1 pipeline with a small SFT warm-up; on-policy distillation methods for language and vision-language models emphasize that a “coldstart alignment” of student and teacher distributions is essential for successful online training; and recent diffusion-model distillation methods similarly cast reward-guided fine-tuning as iteratively projecting onto soft-optimal policies rather than running unconstrained RL [e.g. 7, 2, 4, 26, 23].3 In our framework, these observations are natural consequences of importance-weighted KL projection. Consider a target Gibbs posterior qGibbs (y | x) ∝ πref (y | x) exp(R(y; D, x)/β) and a sampling distribution µ(y | x) used to generate candidate outputs. The effective sample size of (y|x) the importance weights w(y) ∝ qGibbs collapses when the support of µ has little overlap with µ(y|x) the high-reward region under qGibbs . In that regime, both off-policy reward-weighted SFT and on-policy distillation degenerate: most samples receive negligible weight, and the few high-weight samples lead to unstable gradients. A small SFT or behavior-cloning warm-up can be understood as moving πref (or the student policy) into the high-reward region so that µ and qGibbs have non-trivial overlap. On-policy distillation steps then refine the policy by repeatedly projecting the current Gibbs-approximate teacher onto the student family via forward-KL, as in Parts II and III, but now in a regime where the importance weights have reasonable variance. Our framework thus predicts a theoretical necessity of cold start for OPD-style pipelines: without sufficient support overlap, Gibbs-posterior projection is not statistically feasible, regardless of the specific optimizer used. Interestingly, recent work on reward-guided fine-tuning of diffusion models adopts an almost identical recipe: simulate soft-optimal policies under a reward, then distill them via KL minimization into the base diffusion model [23]. From our perspective, these methods are simply applying the same Gibbs-posterior projection principle to non-autoregressive generative models.
5.3 Compute allocation and reasoning models Our analysis so far has been distributional and first-order: SFT, KL-regularized RL, and rewardweighted ICL all aim to approximate a (common) Bayes or Gibbs posterior, and their updates share the same weighted-score form. An orthogonal, but increasingly important, dimension in contemporary systems is how compute is allocated between training time and inference time. Recent “reasoning models” such as DeepSeek-R1 and OpenAI’s o1 family explicitly scale inferencetime compute via long chain-of-thought traces, multi-sample search, or tree-style exploration [7, 17, 10]. At a high level, these models combine two regimes: • Test-time inference (ICL / search). Given a fixed base policy, the model spends substantial compute per query to explore alternative reasoning paths and approximate a task-specific posterior over solutions. • Training-time amortization (RL / OPD / SFT). RLHF, RLVR, and on-policy distillation stages then amortize this search by projecting the high-reward posterior (often defined 3
For an accessible engineering perspective on on-policy distillation for post-training, see the blog post by Lu and Lab [14].
16
over multi-step chain-of-thoughts) back into the base model’s weights via KL-regularized updates. DeepSeek-R1 and related work report a somewhat surprising empirical finding: few-shot prompting often does not improve, and can even degrade, the performance of heavily RL-tuned reasoning models [7]. In our framework, this can be understood as a mismatch between the posterior that the model has been trained to amortize and the posterior implicitly defined by few-shot prompts. During RLHF/RLVR-style training, the model is typically optimized to approximate a Gibbs posterior qGibbs (y | x) under a distribution of zero-shot or lightly-structured prompts x. In contrast, few-shot prompting at inference time replaces x by a richer context (Dk , x) and implicitly targets a different posterior PBayes (y | x, Dk ), as in Part I. Unless the model has been meta-trained to perform reward-weighted ICL (RW-ICL) over such contexts, the learned mapping (Dk , x) 7→ Mθ (· | x, Dk ) need not approximate the desired posterior, and the additional context can act as distribution shift rather than useful information. From this perspective, few-shot prompting “fails” on some reasoning models not because ICL is inherently weaker than RL, but because the posterior level has changed (from token-level qGibbs (y | x) to in-context qGibbs (y | x, Dk )) without a corresponding change in the training objective. Our RW-ICL formulation suggests a natural remedy: if one wants RL-tuned models to benefit from few-shot prompts, the reward signal should be defined at the contextual level R(y; Dk , x) and the meta-objective should directly train Mθ to approximate the corresponding Gibbs posterior qGibbs (· | x, Dk ). Overall, the Bayesian/KL framework suggests that “System 2” reasoning via test-time search and ICL, and “System 1” behavior via amortized RL/SFT, are two ways of paying for the same posterior: either spend more compute per query to approximate it explicitly, or invest more training compute to bake it into the weights. Reasoning models such as DeepSeek-R1 and o1 can be viewed as deliberately combining both, using inference-time compute to explore high-reward regions of the posterior and training-time KL projection (via RLHF/RLVR/OPD) to amortize the resulting distributions into a fast at-test-time policy.
Final Conclusions Combining Parts I-IV, we obtain the following Bayesian chain: • Few-shot ICL (Bayesian view). Under a hierarchical task model, the Bayes-optimal incontext predictor is the posterior predictive PBayes (y | xk+1 , Dk ). A sufficiently expressive Transformer trained on many tasks approximates this predictor in an amortized way. • Generalized Bayes and KL-regularized RL. KL-regularized RLHF/RLVR defines a Gibbs posterior over outputs qGibbs (y | x, D) ∝ πref (y | x) exp(R(y; D, x)/β), where R is a reward or advantage signal. • Reward-weighted SFT / AWSFT. Reward-weighted SFT (for single-step problems) and advantage-weighted SFT (for RL) are exactly the forward KL projections of these Gibbs posteriors onto the parametric model family {pθ } or {πθ }. 17
• Reward-weighted ICL (RW-ICL). Meta-training a Transformer with reward-weighted incontext objectives makes its predictions Mθ (· | x, D) approximate the same Gibbs posteriors qGibbs (· | x, D) in an amortized fashion. In summary,
(i) Few-shot ICL ≈ (unweighted) in-context SFT as a forward-KL/MLE projection onto PBayes , (72) (ii) Reward-weighted ICL (RW-ICL) ≈ reward/advantage-weighted in-context SFT,
(73)
(iii) KL-regularized RLHF/RLVR ≃ reward-weighted SFT on the corresponding Gibbs posterior. (74) Here, “≈” is intended in an idealized sense at two levels: at the distribution level, as forward-KL projections onto the same posterior (Bayes or Gibbs) under our modeling assumptions; and at the first-order optimization level, as stochastic-gradient updates of the same weighted-score form described in Section 3.2. It does not claim that the corresponding practical algorithms are identical beyond these aspects. It is important to stress that this equivalence is deliberately confined to objectives and update forms. The source and granularity of the learning signal remain genuinely different across paradigms. Supervised fine-tuning observes dense token-level supervision from mappings (x, y≤i ) 7→ yi+1 ; KLregularized RL begins with sparse sequence- or trajectory-level rewards and must perform explicit credit assignment via value/advantage estimation before reducing to an advantage-weighted SFT update; and few-shot ICL reuses dense supervision in pretraining to learn a meta-learner θ, but performs task-specific adaptation at test time in the forward activations rather than via additional gradients. Our framework reconciles these methods at the level of their posterior targets and first-order dynamics, while leaving room for these process-level differences to drive distinct sampleefficiency and stability trade-offs in practice. All of these procedures are different faces of the same underlying mechanism: a (possibly generalized) Bayesian update of a reference model that produces a target posterior (either the true Bayes posterior PBayes or a designed Gibbs posterior qGibbs ), followed by a forward-KL (MLE) projection onto the parameterized policy family. Meta-learning / first-order interpretations (ICL ≈ one-step SFT / AWSFT) can then be seen as algorithmic realizations of these projection steps and are best viewed as complementary to the Bayesian backbone developed here. From an information-design perspective, vanilla SFT passively assumes that the data source already supplies the optimal information structure (the true posterior qtrue ). RLHF/RLVR first designs a target posterior qGibbs via the choice of reward R and temperature β, and then performs the same forward-KL/MLE projection as in SFT. In this sense, RLHF is SFT on a designed information structure: the optimization remains maximum likelihood, but the posterior being fit has been engineered to encode the desired incentives.
18
References [1] Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Rémi Munos, et al. Maximum a posteriori policy optimisation. arXiv preprint arXiv:1806.06920, 2018. URL https://arxiv. org/abs/1806.06920. [2] Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos, Matthieu Geist, and Olivier Bachem. On-policy distillation of language models: Learning from selfgenerated mistakes. In International Conference on Learning Representations, 2024. [3] Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou. What learning algorithm is in-context learning? investigations with linear models. arXiv preprint arXiv:2211.15661, 2022. URL https://arxiv.org/abs/2211.15661. [4] Walid Bousselham, Hilde Kuehne, and Cordelia Schmid. Vold: Reasoning transfer from llms to vision-language models via on-policy distillation. arXiv preprint, 2025. [5] Damai Dai, Yutao Sun, Li Dong, Yaru Hao, Shuming Ma, Zhifang Sui, and Furu Wei. Why can GPT learn in-context? language models implicitly perform gradient descent as metaoptimizers. arXiv preprint arXiv:2212.10559, 2022. URL https://arxiv.org/abs/2212. 10559. [6] Yuhao Du, Zhuo Li, Pengyu Cheng, Zhihong Chen, Yuejiao Xie, Xiang Wan, and Anningzhe Gao. Simplify rlhf as reward-weighted sft: A variational method. arXiv preprint, 2025. [7] Daya Guo et al. Deepseek-r1: Incentivizing reasoning capability in llms. arXiv preprint, 2025. [8] Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In ICML (PMLR v80), 2018. URL https://proceedings.mlr.press/v80/haarnoja18b/haarnoja18b.pdf. [9] Timo Kaufmann, Paul Weng, Viktor Bengs, and Eyke Hüllermeier. A survey of reinforcement learning from human feedback. Transactions on Machine Learning Research, 2025. ISSN 2835-8856. [10] Hanyu Lai, Xiao Liu, Junjie Gao, Jiale Cheng, Zehan Qi, Yifan Xu, Shuntian Yao, Dan Zhang, Jinhua Du, Zhenyu Hou, Xin Lv, Minlie Huang, Yuxiao Dong, and Jie Tang. A survey of posttraining scaling in large language models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2771–2791, 2025. [11] Nathan Lambert. Reinforcement Learning from Human Feedback. rlhfbook.com, 2024. Online textbook. [12] Sergey Levine. Reinforcement learning and control as probabilistic inference: Tutorial and review. arXiv preprint arXiv:1805.00909, 2018. URL https://arxiv.org/abs/1805.00909. [13] Yingcong Li, M. Emrullah Ildiz, Dimitris Papailiopoulos, and Samet Oymak. Transformers as algorithms: Generalization and stability in in-context learning. In ICML, 2023. URL https://proceedings.mlr.press/v202/li23l/li23l.pdf. [14] Kevin Lu and Thinking Machines Lab. On-policy distillation. Thinking Machines Lab: Connectionism, 2025. doi: 10.64434/tml.20251026. https://thinkingmachines.ai/blog/ on-policy-distillation. 19
[15] Youssef Mroueh et al. GRPO’s effective loss, dynamics, and success amplification. arXiv preprint arXiv:2503.06639, 2025. URL https://arxiv.org/abs/2503.06639. [16] Ashvin Nair, Murtaza Dalal, Abhishek Gupta, and Sergey Levine. Accelerating online reinforcement learning with offline datasets. arXiv preprint arXiv:2006.09359, 2020. URL https://arxiv.org/abs/2006.09359. [17] OpenAI. Learning to reason with llms. https://openai.com/index/ learning-to-reason-with-llms/, 2024. Announces the OpenAI o1 “reasoning” model. [18] Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, et al. Training language models to follow instructions with human feedback. In NeurIPS, 2022. URL https://proceedings.neurips.cc/paper_files/paper/2022/file/ b1efde53be364a73914f58805a001731-Paper-Conference.pdf. [19] Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine. Advantage-weighted regression: Simple and scalable off-policy reinforcement learning. arXiv preprint arXiv:1910.00177, 2019. URL https://arxiv.org/abs/1910.00177. [20] Jan Peters, Katharina Mülling, and Yasemin Altun. Relative entropy policy search. In AAAI, 2010. URL https://ojs.aaai.org/index.php/AAAI/article/view/7727/7588. [21] Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Chelsea Finn, et al. Direct preference optimization: Your language model is secretly a reward model. NeurIPS, 2023. URL https://arxiv.org/abs/2305.18290. [22] Arik Reuter, Tim G. J. Rudner, Vincent Fortuin, and David Rügamer. Can transformers learn full Bayesian inference in context? arXiv preprint arXiv:2501.16825, 2025. URL https: //arxiv.org/abs/2501.16825. [23] Xingyu Su, Xiner Li, Masatoshi Uehara, Sunwoo Kim, Yulai Zhao, Gabriele Scalia, Ehsan Hajiramezanali, Tommaso Biancalani, Degui Zhi, and Shuiwang Ji. Iterative distillation for reward-guided fine-tuning of diffusion models in biomolecular design. arXiv preprint, 2025. [24] Tomoya Wakayama and Taiji Suzuki. In-context learning is provably Bayesian inference: A generalization theory for meta-learning. arXiv preprint arXiv:2510.10981, 2025. URL https: //arxiv.org/abs/2510.10981. [25] Xumeng Wen, Zihan Liu, Shun Zheng, Zhijian Xu, Shengyu Ye, Zhirong Wu, Xiao Liang, Yang Wang, Junjie Li, Ziming Miao, Jiang Bian, and Mao Yang. Reinforcement learning with verifiable rewards implicitly incentivizes correct reasoning in base LLMs. arXiv preprint arXiv:2506.14245, 2025. URL https://arxiv.org/abs/2506.14245. [26] Tianzhu Ye, Li Dong, Zewen Chi, Xun Wu, Shaohan Huang, and Furu Wei. Black-box onpolicy distillation of large language models. arXiv preprint, 2025.
Appendix A: Bayesian ICL Results Used in the Main Text This appendix summarizes the key theoretical results from recent work on Bayesian views of incontext learning that we rely on in Part I: Wakayama and Suzuki [24] and Reuter et al. [22]. We 20
only state what is needed for our arguments; full technical details and proofs are in the original papers.
A.1 Risk decomposition and non-asymptotic bounds (Wakayama & Suzuki) Wakayama and Suzuki [24] formalize a hierarchical prompt-generating process over tasks and show that the in-context learning (ICL) risk admits an exact orthogonal decomposition under squared loss. Let M be any in-context predictor (not necessarily a Transformer), and let R(M ) denote its expected ICL risk under their meta-distribution. Then R(M ) = RBG (M ) + RPV ,
(75)
where: • RBG (M ) is the Bayes Gap: an excess-risk term measuring how far M is from the Bayesoptimal in-context predictor MBayes ; • RPV is the Posterior Variance: an irreducible term depending only on the task distribution and context length, not on M . In particular, MBayes is the unique minimizer of R(M ), and RPV captures the intrinsic task uncertainty. This is exactly the content we use in Theorem 1 in Part I: the Bayes posterior predictive is the normative target for in-context prediction, and any in-context learning algorithm is judged by its Bayes Gap RBG (M ). For a class of uniform-attention Transformers with feature dimension m, context length p and number of pretraining prompts N , Wakayama and Suzuki [24] further provide a non-asymptotic upper bound of the form m 1 − d2α e E RBG (Mθ̂ ) ≲ m eff + O + , (76) pN N where α ∈ (0, 1] is a Hölder exponent and deff is an effective dimension of the data manifold. The first term is an approximation error controlled by m, and the second term is a pretraining generalization error depending jointly on p and N . They also analyze mixtures of task types and input-distribution shift, showing that: • In mixed-task settings, the posterior over the task index concentrates quickly as the context grows, so the Bayes-optimal meta-algorithm rapidly specializes to the true task family. • Under input-distribution shift between pretraining and inference, only the Bayes Gap increases, with an out-of-distribution penalty proportional to a Wasserstein distance between input distributions, while the posterior variance term is intrinsic to the target domain. We do not need these bounds explicitly, but they support our assumption in Sec. 1.3 that a sufficiently expressive Transformer, trained on many tasks, can approximate MBayes up to a small Bayes Gap in the regimes of interest. 21
A.2 Transformers as amortized Bayesian inference (Reuter et al.) Reuter et al. [22] study whether in-context Transformers can approximate full Bayesian posterior inference over latent variables. They construct several synthetic scenarios (generalized linear models, Gaussian mixture models, etc.) where the ground-truth posterior over latents is available via high-quality MCMC, and then: • Train an in-context learner Mθ that maps each dataset D to a predictive distribution (or posterior sampler) over latent variables and responses. • Compare the distribution of samples from Mθ with those from gold-standard Bayesian methods (e.g., HMC, SGLD, variational approximations) using metrics such as C2ST, MMD, Wasserstein distance, and RMSE. • Investigate robustness under distribution shift between training and test datasets, showing that the in-context learner often remains close to the Bayesian baseline in out-of-distribution regimes. Overall, they find that suitably trained Transformers can act as strong amortized Bayesian inference engines in context, often matching or approaching fully Bayesian baselines on both synthetic and real-world datasets, and sometimes outperforming classical approximate methods in certain distributional metrics. This provides empirical support for the assumption used in Sec. 1.3 that a sufficiently expressive Transformer, trained on many tasks, can approximate the Bayes posterior predictive PBayes (y | x, Dk ) via a single forward pass Mθ (y | x, Dk ).
A.3 Relation to the main text (and to meta-learning views) In the main body of this note: • Part I uses Wakayama and Suzuki [24] to justify that the Bayes posterior predictive is the canonical target for in-context prediction, and that ICL risk decomposes into a Bayes Gap plus an irreducible posterior-variance term. • The amortized-inference perspective in Sec. 1.3 and the statement that a single Transformer forward pass can approximate a dataset-level Bayesian posterior are motivated by Reuter et al. [22]. We deliberately do not reproduce the gradient-based/meta-learning derivations in this appendix. Instead, we refer to Akyürek et al. [3], Dai et al. [5], and Li et al. [13] for rigorous analyses showing that, in linear and linearized regimes, Transformers can implement one-step gradient updates in their forward passes and that attention has an algebraic dual representation as accumulated gradient corrections. In our view, those results are best interpreted as algorithmic realizations of the Bayesian predictor described in Parts I–III, rather than as the primary conceptual foundation.
22
A.4 Task–prior representation under within-task exchangeability In Part I (Section 1.2), we adopted the hierarchical task–prior model f ∼ P (f ),
(x, y) | f ∼ pf (x, y)
to formalize few-shot in-context learning as Bayes posterior prediction over tasks. In this appendix, we justify that such a hierarchical representation is a canonical consequence of a much weaker and more intuitive assumption on within-task data: exchangeability. Throughout, exchangeability is an assumption on the data-generating process within a task episode, not on any particular neural architecture used to implement in-context learning. We first recall the notion of exchangeability and then invoke a classical representation theorem of de Finetti–Hewitt–Savage to obtain the hierarchical model used in the main text. A.4.1 Within-task exchangeability. Let X be the input space and Y the output space (e.g., labels or next tokens). Define the sample space Z := X × Y. We consider sequences Zi = (Xi , Yi ) ∈ Z,
i = 1, 2, . . .
generated under a single “task episode”. Definition A.4.1 (Finite and infinite exchangeability). A finite sequence Z1 , . . . , Zn is exchangeable if for every permutation π of {1, . . . , n}, d
(Z1 , . . . , Zn ) = (Zπ(1) , . . . , Zπ(n) ), i.e., its joint distribution is invariant under any reordering of the indices. An infinite sequence (Zi )i≥1 is exchangeable if for every n ≥ 1, the finite prefix (Z1 , . . . , Zn ) is exchangeable. Intuitively, for an exchangeable sequence, only the multiset {Z1 , . . . , Zn } matters, not the order. This is a standard formalization of “order-invariant” data within a task. Note that i.i.d. sequences are a special case of exchangeable sequences, but exchangeability is strictly weaker. In our setting, we model each task episode as inducing an exchangeable sequence of sample pairs: Assumption A.4.2 (Within-task exchangeability). For each task episode, the (conceptual) infinite sequence of observations (Zi )i≥1 = (Xi , Yi )i≥1 taking values in Z = X × Y is exchangeable. In practice, we only ever observe a finite dataset Dk = {(xi , yi )}ki=1 , but we assume it arises as a finite prefix of such an exchangeable sequence. This is the usual idealization used in probability theory: we posit that, if we were to continue sampling from the same task, the joint law of the extended sequence would remain exchangeable. A.4.2 A de Finetti–Hewitt–Savage representation. The sample space Z = X × Y is a standard Borel space in typical supervised learning settings (e.g., X , Y finite or subsets of Rd ), so we can invoke the general form of de Finetti’s representation theorem as extended by Hewitt and Savage. 23
We state an informal version sufficient for our purposes. Theorem A.4.3 (de Finetti–Hewitt–Savage representation; informal). Let (Zi )i≥1 be an infinite exchangeable sequence of random variables taking values in a standard Borel space Z. Then there exists a random probability measure µ on Z (a random element of P(Z)) such that, conditional on µ, the sequence is i.i.d. with common distribution µ: (Zi )i≥1 | µ i.i.d. ∼ µ,
µ ∼ Π,
for some probability law Π on P(Z). Moreover, the mixing law Π is unique. Equivalently, the joint law of (Zi )i≥1 can be written as a mixture of i.i.d. sequences. The original versions of this theorem were proved by de Finetti for Bernoulli sequences and generalized by Hewitt, Savage, and many others to arbitrary standard Borel spaces; see, e.g., standard modern references in probability theory for precise statements and proofs. Under Assumption A.4.2, the sequence (Zi ) generated by a given task episode satisfies the hypotheses of Theorem A.4.3, and hence admits such a mixture-of-i.i.d. representation.
A.4.3 From a random measure to a task prior. We now reinterpret the random measure µ from Theorem A.4.3 as a task, and its mixing distribution Π as a task prior P (f ). Let F := P(Z) denote the space of all probability measures on Z, equipped with the σ-algebra generated by the usual weak topology. For each realization µ ∈ F , define the joint law on Z by pµ (z) := µ({z}),
z ∈ Z,
or more generally pµ (z) as the Radon–Nikodym density of µ with respect to a suitable base measure if Z is continuous. We introduce the following identification: • Tasks. Each realization of µ is interpreted as a task f ∈ F, with joint law pf (x, y) over X × Y; • Task prior. The mixing law Π over µ is renamed as a prior P (f ) over tasks. Under this identification, Theorem A.4.3 can be restated as: There exists a random task f ∼ P (f ) and a conditional distribution (Xi , Yi ) | f ∼ pf (x, y) such that, for each task episode, the joint law of (Zi )i≥1 = (Xi , Yi )i≥1 is the same as that of an i.i.d. sequence drawn from pf . Formally, we obtain the hierarchical representation: f ∼ P (f ),
(Xi , Yi )i≥1 | f i.i.d. ∼ pf (x, y).
This is exactly the task–prior generative model used in Section 1.2 and the main Bayesian ICL analysis. 24
For supervised tasks, it is often convenient to factor pf (x, y) as pf (x, y) = pf (x) pf (y | x), and to interpret f primarily via its conditional component pf (y | x). In that case, we may succinctly write f : X → ∆(Y), with P (f ) a prior over such conditional task functions, as in the main text. We summarize this as a proposition: Proposition A.4.4 (Task–prior representation under exchangeability). Suppose that, for each task episode, the within-task observation sequence (Xi , Yi )i≥1 taking values in X ×Y is infinite and exchangeable (Assumption A.4.2). Then there exists a task space F and a prior P (f ) over F such that the joint distribution over data induced by a task episode admits the hierarchical representation f ∼ P (f ), (Xi , Yi ) | f i.i.d. ∼ pf (x, y). In particular, the existence of a task prior P (f ) is not an additional modeling assumption beyond within-task exchangeability; it is a canonical hierarchical reparameterization of the same joint law guaranteed by the de Finetti–Hewitt–Savage theorem. Proof sketch. Apply Theorem A.4.3 to the exchangeable sequence (Zi )i≥1 = (Xi , Yi )i≥1 on Z = X × Y to obtain a random measure µ on Z and a mixing law Π such that (Zi ) | µ are i.i.d. with common distribution µ, and µ ∼ Π. Define F := P(Z), identify each µ ∈ F with a task f , and set P (f ) := Π. Let pf (x, y) denote the joint law induced by f on X × Y. Then f ∼ P (f ),
(Xi , Yi ) | f i.i.d. ∼ pf (x, y)
is exactly the hierarchical generative process whose marginal law over (Xi , Yi )i≥1 coincides with the original exchangeable sequence. □ A.4.4 Remarks and scope. (1) What is the real modeling assumption? Proposition A.4.4 shows that the only substantive structural assumption we make at the task level is within-task exchangeability (Assumption A.4.2). Given this, the existence of a hierarchical task prior P (f ) is a representation theorem, not an extra ad hoc assumption. (2) Finite sequences and approximate exchangeability. In practice, we only observe finite datasets Dk = {(xi , yi )}ki=1 . The infinite-sequence formulation assumes that these finite datasets can be extended to an infinite exchangeable sequence. There also exist “finite de Finetti” theorems which provide approximate mixture representations for finitely exchangeable sequences; we do not rely on their technical details here, but conceptually they support the view that approximately exchangeable finite datasets are well-approximated by such hierarchical models. (3) Connection to Wakayama & Suzuki’s hierarchical model. Wakayama and Suzuki [14] start directly from a hierarchical task prior f ∼ P (f ) and derive non-asymptotic generalization bounds and a Bayes-optimal ICL risk decomposition. Our use of de Finetti–type representations clarifies that this hierarchical structure is not arbitrary: it is the canonical form of the joint distribution under within-task exchangeability. This provides a principled foundation for adopting their model in Part I. (4) Empirical plausibility. Whether within-task exchangeability is a good approximation for real pretraining corpora is ultimately an empirical question. In the main text, we view the task–prior 25
model as an effective theory of the latent task structure in natural data and discuss experimental designs (e.g., shuffling document structure, mixing tasks within prompts) that can probe the degree to which such a model is appropriate. (5) Order effects in practical in-context learners. Assumption A.4.2 and Proposition A.4.4 concern the joint law of within-task data (Xi , Yi )i≥1 under an exchangeable data-generating process. Real transformer models process ordered prompts with positional embeddings, and the order of demonstrations can affect predictions (e.g., via position-dependent attention weights). As discussed in Section 1.2, we treat such order effects as implementation-level deviations from the exchangeable Bayes benchmark characterized here, rather than violations of the underlying exchangeability assumption.
26