May, 2026
Exponential families from a single KL identity Marc Dymetman* ∗ Scientific Consultant
arXiv:2604.28036v1 [cs.LG] 30 Apr 2026
Abstract Exponential families encompass the distributions central to modern machine learning — softmax, Gaussians, and Boltzmann distributions — and underlie the theory of variational inference, entropy-regularized reinforcement learning, and RLHF. We isolate a simple identity for exponential families that expresses the KL difference KL( 𝑞 ∥ 𝑝𝜆 2 ) − KL( 𝑞 ∥ 𝑝𝜆 1 ) in terms of the log-partition function 𝐴 ( 𝜆 ) and the moment 𝜇 𝑞 . Remarkably, this identity together with the single fact that KL ≥ 0 (with equality iff 𝑝 = 𝑞) suffices, by direct substitution and rearrangement, to derive a cluster of results that are classically obtained by separate, heavier arguments: a generalized three-point identity for arbitrary reference distributions, Pythagorean theorems for I-projections and reverse I-projections, convexity of the log-partition function, identification of its Legendre dual in KL terms, the Gibbs variational principle, and the explicit optimizer in KL-regularized reward maximization, including the exponential tilting formula underlying entropy-regularized control and RLHF. Beyond these purely algebraic consequences, standard analytic arguments recover the gradient formula for the log-partition function, the Bregman representation of within-family KL divergence, and the surjectivity of the moment map. The note is self-contained.
1. Setup Let 𝑌 be a finite or countable set, and let 𝑎 be a strictly positive probability distribution on 𝑌 (the base distribution). Let 𝜙 : 𝑌 → ℝ𝑑 be a function (the sufficient statistic). For 𝜆 ∈ ℝ𝑑 , define the partition function ∑︁ 𝑍𝜆 = 𝑎 ( 𝑦 ) 𝑒𝜆 ·𝜙 ( 𝑦 ) , 𝑦 ∈𝑌
the log-partition function 𝐴 ( 𝜆 ) = log 𝑍 𝜆 , and the natural parameter space Λ = { 𝜆 ∈ ℝ𝑑 : 𝑍 𝜆 < ∞}. When 𝑌 is finite, Λ = ℝ𝑑 . For 𝜆 ∈ Λ, the exponential family member with natural parameter 𝜆 is 𝑝𝜆 ( 𝑦 ) =
𝑎 ( 𝑦 ) 𝑒𝜆 ·𝜙 ( 𝑦 ) . 𝑍𝜆
Note that 𝑝0 = 𝑎 and 𝐴 (0) = 0. Since 𝑎 ( 𝑦 ) > 0 for all 𝑦 ∈ 𝑌 , we have 𝑝𝜆 ( 𝑦 ) > 0 for all 𝑦 and 𝜆 ∈ Λ. Í For any distribution 𝑞 on 𝑌 such that 𝑦 𝑞 ( 𝑦 )| 𝜙 𝑗 ( 𝑦 )| < ∞ for all 𝑗, we write ∑︁ 𝜇𝑞 = 𝑞 ( 𝑦 ) 𝜙 ( 𝑦 ) ∈ ℝ𝑑 . 𝑦 ∈𝑌
When 𝑞 = 𝑝𝜆 , we write 𝜇 𝜆 = 𝜇 𝑝𝜆 ; this is always finite when 𝑌 is finite. We refer to Barndorff-Nielsen (1978), Brown (1986), Wainwright and Jordan (2008) for classical treatments of exponential families.
2. The identity Proposition 1 (KL difference identity). Let 𝜆 1 , 𝜆 2 ∈ Λ and let 𝑞 be a distribution on 𝑌 such that 𝜇 𝑞 exists. If 𝑌 is finite, then KL 𝑞 ∥ 𝑝𝜆 2 − KL 𝑞 ∥ 𝑝𝜆 1 = 𝐴 ( 𝜆 2 ) − 𝐴 ( 𝜆 1 ) + 𝜇 𝑞 · ( 𝜆 1 − 𝜆 2 ) (1)
Exponential families from a single KL identity
holds unconditionally. If 𝑌 is countably infinite, it holds whenever at least one of KL 𝑞 ∥ 𝑝𝜆 1 , KL 𝑞 ∥ 𝑝𝜆 2 is finite; in that case both are finite. Proof. Since 𝑝𝜆 ( 𝑦 ) > 0 for all 𝑦 , the log-ratio of two exponential family members is well-defined and affine in 𝜙: 𝑝𝜆 ( 𝑦 ) = ( 𝜆1 − 𝜆2) · 𝜙( 𝑦) + 𝐴 ( 𝜆2) − 𝐴 ( 𝜆1) . log 1 𝑝𝜆 2 ( 𝑦 ) Taking the expectation under 𝑞 (finite since 𝜇 𝑞 exists): 𝑝𝜆 1 𝔼𝑞 log = ( 𝜆 1 − 𝜆 2) · 𝜇𝑞 + 𝐴 ( 𝜆 2) − 𝐴 ( 𝜆 1) .
(2)
𝑝𝜆 2
The identity then follows from KL 𝑞 ∥ 𝑝𝜆 2 − KL 𝑞 ∥ 𝑝𝜆 1
𝑞 𝑞 = 𝔼𝑞 log − log 𝑝𝜆 2 𝑝𝜆 1
= 𝔼𝑞 log
𝑝𝜆 1
,
𝑝𝜆 2
where the difference on the left is well-defined (and both KLs are finite) since their difference equals the finite quantity (2). □ Part of the interest of identity (1) is that it relates three different types of quantities in a single linear equation: KL divergences (the left side), log-partition values (the 𝐴 ( 𝜆 ) terms), and moments (the 𝜇 𝑞 term). Different specializations move between these worlds: eliminating the log-partition terms yields the three-point identity (Section 3), while eliminating the KL terms yields convexity of 𝐴 ( 𝜆 ) (Section 4). Example 2 (Categorical distributions and softmax). Let 𝑌 = {1, . . . , 𝑘 } with uniform base measure 𝑎 ( 𝑦 ) = 1/𝑘, and let 𝜙 ( 𝑦 ) = 𝑒 𝑦 ∈ ℝ𝑘 be the 𝑦 -th standard basis vector, so 𝜆 · 𝜙 ( 𝑦 ) = 𝜆 𝑦 . The log-partition function is Í 𝐴 ( 𝜆 ) = log 1𝑘 𝑘𝑦 =1 𝑒𝜆 𝑦 = LSE( 𝜆 ) − log 𝑘, Í where LSE( 𝜆 ) = log 𝑦 𝑒𝜆 𝑦 is the log-sum-exp function, and the exponential family member is the softmax distribution: 𝑒𝜆 𝑦 𝑝𝜆 ( 𝑦 ) = Í𝑘
𝑗=1 𝑒
𝜆𝑗
= softmax( 𝜆 ) 𝑦 .
Since 𝜙 ( 𝑦 ) = 𝑒 𝑦 , for any distribution 𝑞 on 𝑌 the moment is 𝜇 𝑞 = 𝔼𝑞 [ 𝑒𝑌 ] = 𝑞 (the probability vector itself), and in particular 𝜇 𝜆 = 𝑝𝜆 . The identity. Using 𝐴 ( 𝜆 2 ) − 𝐴 ( 𝜆 1 ) = LSE( 𝜆 2 ) − LSE( 𝜆 1 ), identity (1) becomes 𝑘 ∑︁ KL 𝑞 ∥ 𝑝𝜆 2 − KL 𝑞 ∥ 𝑝𝜆 1 = LSE( 𝜆 2 ) − LSE( 𝜆 1 ) + 𝑞 𝑦 ( 𝜆 1,𝑦 − 𝜆 2,𝑦 ) . 𝑦 =1
This can be verified directly: since log 𝑝𝜆 ( 𝑦 ) = 𝜆 𝑦 − LSE( 𝜆 ), KL( 𝑞 ∥ 𝑝𝜆 ) = LSE( 𝜆 ) − where 𝐻 ( 𝑞) = − side.
Í
Í
𝑦 𝑞 𝑦 𝜆 𝑦 − 𝐻 ( 𝑞) ,
𝑦 𝑞 𝑦 log 𝑞 𝑦 , and the left-hand difference reduces immediately to the right-hand
KL-regularized optimization. For a reward 𝑟 : 𝑌 → ℝ and temperature 𝛽 > 0, Corollary 11 gives max 𝔼𝑞 [ 𝑟 ] − 𝛽 KL( 𝑞 ∥ 𝑎) = 𝛽 LSE( 𝑟 / 𝛽 ) − 𝛽 log 𝑘, 𝑞
2
Exponential families from a single KL identity
with unique maximizer 𝑞∗ ( 𝑦 ) = softmax( 𝑟 / 𝛽 ) 𝑦 ∝ 𝑒𝑟 ( 𝑦 )/ 𝛽 . Since KL( 𝑞 ∥ 𝑎) = log 𝑘 − 𝐻 ( 𝑞) for uniform 𝑎, the objective is equivalently 𝔼𝑞 [ 𝑟 ] + 𝛽 𝐻 ( 𝑞), the standard maximum-entropy objective in reinforcement learning Haarnoja et al. (2018), Ziebart (2010). The optimal policy is the Boltzmann distribution at temperature 𝛽 , and the maximum of this entropy-regularized objective is 𝛽 LSE( 𝑟 / 𝛽 ), the soft-maximum of 𝑟 .
Extension to general measurable spaces For readers interested in the fully general setting, the identity extends to arbitrary measurable spaces with no change to the algebra. Let (𝑌 , F , 𝑎) be a probability space and 𝜙 : 𝑌 → ℝ𝑑 a measurable function. The partition function, natural parameter space, and exponential family are defined as before with sums replaced by integrals: ∫ 𝑍𝜆 =
𝑒𝜆 · 𝜙 ( 𝑦 ) 𝑑𝑎 ( 𝑦 ) ,
𝑌
𝑝𝜆 ( 𝑦 ) =
𝑒𝜆 ·𝜙 ( 𝑦 ) 𝑍𝜆
(density with respect to 𝑎), and 𝜇 𝑞 = 𝔼𝑞 [ 𝜙] when 𝔼𝑞 [| 𝜙 𝑗 |] < ∞ for all 𝑗. The identity (1) holds under the same conditions as Proposition 1, with the countably infinite case replaced by the requirement that at least one of KL 𝑞 ∥ 𝑝𝜆 1 , KL 𝑞 ∥ 𝑝𝜆 2 is finite. The key additional step in the proof is to establish absolute continuity: since all 𝑝𝜆 have strictly positive densities with respect to 𝑎, KL 𝑞 ∥ 𝑝𝜆 1 < ∞ implies 𝑞 ≪ 𝑝𝜆 1 ∼ 𝑎 ∼ 𝑝𝜆 2 , so log( 𝑞/ 𝑝𝜆 2 ) is 𝑞-integrable and KL 𝑞 ∥ 𝑝𝜆 2 < ∞ follows by linearity of expectation. The integrability of 𝜇 𝑞 under 𝑝𝜆 for 𝜆 ∈ int( Λ) is guaranteed by a standard exponential moment argument: since 𝑍 𝜆 ±𝑡𝑒 𝑗 < ∞ for small 𝑡 > 0, the inequality | 𝜙 𝑗 | 𝑒𝜆 ·𝜙 ≤ 𝑡 −1 𝑒 ( 𝜆 ±𝑡𝑒 𝑗 ) ·𝜙 gives integrability.
3. Algebraic consequences All results in this section follow from the identity (1) and the single fact that KL( 𝑝 ∥ 𝑞) ≥ 0 with equality if and only if 𝑝 = 𝑞. Remarkably, KL nonnegativity—together with the equality condition KL( 𝑝 ∥ 𝑞) = 0 ⇐⇒ 𝑝 = 𝑞—is the sole ingredient behind all minimization and projection results in this note. No appeal to analysis, differentiation, or convexity theory is needed for any of these results. Each result is obtained by choosing specific values of 𝑞, 𝜆 1 , 𝜆 2 and rearranging.
3.1. Generalized three-point identity Setting 𝑞 = 𝑝𝜆 1 in (1) (which requires 𝜇 𝜆 1 to exist, e.g., 𝜆 1 ∈ int( Λ)) gives a specialized form of the identity: KL 𝑝𝜆 1 ∥ 𝑝𝜆 2 = 𝐴 ( 𝜆 2 ) − 𝐴 ( 𝜆 1 ) + 𝜇 𝜆 1 · ( 𝜆 1 − 𝜆 2 ) . (3) Using this to eliminate the log-partition difference in the general identity (1) yields a three-point identity. Corollary 3 (Generalized three-point identity). Under the conditions of Proposition 1, and assuming 𝜇 𝜆 1 exists, KL 𝑞 ∥ 𝑝𝜆 2 = KL 𝑞 ∥ 𝑝𝜆 1 + KL 𝑝𝜆 1 ∥ 𝑝𝜆 2 + ( 𝜇 𝑞 − 𝜇 𝜆 1 ) · ( 𝜆 1 − 𝜆 2 ) . (4) Proof. Subtract (3) from (1) (with the same 𝜆 1 , 𝜆 2 ): the log-partition terms cancel, leaving (4). □ The last term is an inner product between the moment mismatch 𝜇 𝑞 − 𝜇 𝜆 1 and the parameter difference 𝜆 1 − 𝜆 2 . Remark 4. The case where 𝑞 is an arbitrary distribution not necessarily in the exponential family does not appear to be readily found in the literature. It is this generality that makes the subsequent projection results directly applicable to arbitrary distributions: the Pythagorean theorem
3
Exponential families from a single KL identity
(Corollary 5) characterizes the reverse I-projection of any 𝑞 onto the family, and the I-projection corollary (Corollary 7) characterizes the I-projection of 𝑎 onto any moment slice.
3.2. Pythagorean theorem and reverse I-projection When the inner product in (4) vanishes, one obtains a Pythagorean theorem. Corollary 5 (Pythagorean theorem). If ( 𝜇 𝑞 − 𝜇 𝜆 1 ) · ( 𝜆 1 − 𝜆 2 ) = 0, then KL 𝑞 ∥ 𝑝𝜆 2 = KL 𝑞 ∥ 𝑝𝜆 1 + KL 𝑝𝜆 1 ∥ 𝑝𝜆 2 . In particular, if 𝑝𝜆 1 is the moment-matching member of the family for 𝑞 (i.e., 𝜇 𝜆 1 = 𝜇 𝑞 ), then this holds for all 𝜆 2 ∈ Λ, and consequently 𝑝𝜆 1 is the unique closest member of the family to 𝑞 in KL divergence (the reverse I-projection, or reverse information projection, of 𝑞 onto the family, in the terminology of Csiszár and Matúš (2003)). Proof. When 𝜇 𝜆 1 = 𝜇 𝑞 , the inner product vanishes for every 𝜆 2 , and (4) reduces to the stated identity. Since KL 𝑝𝜆 1 ∥ 𝑝𝜆 2 ≥ 0, this gives KL 𝑞 ∥ 𝑝𝜆 2 ≥ KL 𝑞 ∥ 𝑝𝜆 1 for all 𝜆 2 , with equality iff 𝑝𝜆 2 = 𝑝𝜆 1 . □ Remark 6. The existence of a moment-matching 𝜆 1 (i.e., 𝜇 𝜆 1 = 𝜇 𝑞 ) is a separate question, addressed in Section 4.3. The above result says: if such a 𝜆 1 exists, then it yields the unique reverse I-projection.
3.3. I-projection onto moment slices The identity also characterizes 𝑝𝜆 as the I-projection of the base distribution 𝑎 onto the moment slice M 𝜇 . Equivalently, 𝑝𝜆 is the maximum-entropy distribution subject to the moment constraint, without Lagrange multipliers. Given 𝜇 ∈ ℝ𝑑 , define the moment slice M 𝜇 = {𝑞 : 𝔼𝑞 [ 𝜙] = 𝜇 }, i.e., the set of all distributions with mean parameter 𝜇 . Corollary 7 (I-projection onto M 𝜇 ). If 𝜆 ∈ Λ satisfies 𝜇 𝜆 = 𝜇 , then 𝑝𝜆 is the unique minimizer of KL( 𝑞 ∥ 𝑎) over 𝑞 ∈ M 𝜇 (equivalently, the unique maximizer of entropy relative to the distribution 𝑎 subject to the moment constraint). Proof. Setting 𝜆 2 = 0 and 𝜆 1 = 𝜆 in (4), and using 𝑝0 = 𝑎: KL( 𝑞 ∥ 𝑎) = KL( 𝑞 ∥ 𝑝𝜆 ) + KL( 𝑝𝜆 ∥ 𝑎) + ( 𝜇 𝑞 − 𝜇 𝜆 ) · ( 𝜆 − 0) . For 𝑞 ∈ M 𝜇 with 𝜇 𝜆 = 𝜇 , we have 𝜇 𝑞 = 𝜇 𝜆 , so the last term vanishes: KL( 𝑞 ∥ 𝑎) = KL( 𝑞 ∥ 𝑝𝜆 ) + KL( 𝑝𝜆 ∥ 𝑎) . Since KL( 𝑞 ∥ 𝑝𝜆 ) ≥ 0 with equality iff 𝑞 = 𝑝𝜆 , the minimum of KL( 𝑞 ∥ 𝑎) over M 𝜇 is uniquely attained at 𝑞 = 𝑝𝜆 . □ Remark 8. This says that 𝑝𝜆 is the I-projection of the base distribution 𝑎 onto the moment slice M 𝜇 . As in Corollary 5, the existence of a 𝜆 with 𝜇 𝜆 = 𝜇 is addressed separately in Section 4.3.
4
Exponential families from a single KL identity
E
𝑝0 = 𝑎
I-proj
rev. I-proj
M𝜇
𝑝𝜆
𝑞
Figure 1: The reverse I-projection of 𝑞 ∈ M 𝜇 onto the exponential family E and the I-projection of 𝑎 = 𝑝0 onto the moment slice M 𝜇 both yield 𝑝𝜆 . The right angle at 𝑝𝜆 reflects Corollary 5.
3.4. Gibbs variational principle (ELBO) Corollary 9. For all 𝜆 ∈ Λ and all distributions 𝑞 with 𝔼𝑞 [| 𝜙 𝑗 |] < ∞ for all 𝑗 and KL( 𝑞 ∥ 𝑎) < ∞, 𝐴 ( 𝜆 ) = 𝜆 · 𝜇 𝑞 − KL( 𝑞 ∥ 𝑎) +KL( 𝑞 ∥ 𝑝𝜆 ) .
|
{z
ELBO( 𝑞,𝜆 )
(5)
}
In particular, 𝐴 ( 𝜆 ) = sup 𝜆 · 𝜇 𝑞 − KL( 𝑞 ∥ 𝑎) ,
𝑞
and the supremum is attained uniquely at 𝑞 = 𝑝𝜆 . Proof. Set 𝜆 1 = 0 in (1), so 𝑝𝜆 1 = 𝑎 and 𝐴 (0) = 0. For 𝜆 2 = 𝜆 : KL( 𝑞 ∥ 𝑝𝜆 ) = KL( 𝑞 ∥ 𝑎) + 𝐴 ( 𝜆 ) − 𝜆 · 𝜇 𝑞 . Rearranging gives (5), and the variational characterization follows from KL( 𝑞 ∥ 𝑝𝜆 ) ≥ 0 with equality iff 𝑞 = 𝑝𝜆 . □ Remark 10. The decomposition (5) is widely known in machine learning as the ELBO (Evidence Lower Bound) decomposition Blei et al. (2017),Jordan et al. (1999): the log-evidence 𝐴 ( 𝜆 ) equals the ELBO plus the KL gap. In statistical physics, the same result appears as the variational free energy principle, with − 𝐴 ( 𝜆 ) as free energy, KL( 𝑞 ∥ 𝑎) as negative entropy, and 𝜆 · 𝜇 𝑞 as expected energy. In probability and large deviations, it is the Gibbs variational principle; see (Polyanskiy and Wu, 2025, Chapter 4) for a modern treatment and historical references. The variational characterization 𝐴 ( 𝜆 ) = sup𝑞 { 𝜆 · 𝜇 𝑞 − KL( 𝑞 ∥ 𝑎)} directly yields the solution to KL-regularized reward maximization: setting 𝑑 = 1, 𝜙 ( 𝑦 ) = 𝑟 ( 𝑦 ), 𝜆 = 1/ 𝛽 and multiplying through by 𝛽 gives sup𝑞 {𝔼𝑞 [ 𝑟 ] − 𝛽 KL( 𝑞 ∥ 𝑎)} = 𝛽 𝐴 (1/ 𝛽 ), which is the content of the next corollary. We present it separately because the explicit Boltzmann form of the optimizer 𝑞∗ ∝ 𝑎 𝑒𝑟/ 𝛽 , the role of the temperature parameter 𝛽 , and the connections to the reinforcement learning and RLHF literature warrant a dedicated statement.
5
Exponential families from a single KL identity
3.5. KL-regularized reward maximization Corollary 9 directly yields the solution to KL-regularized reward maximization, a key result in reinforcement learning. Setting 𝑑 = 1, 𝜙 ( 𝑦 ) = 𝑟 ( 𝑦 ), 𝜆 = 1/ 𝛽 in the variational characterization and multiplying by 𝛽 gives the following. Corollary 11. Let 𝑑 = 1, 𝜙 ( 𝑦 ) = 𝑟 ( 𝑦 ) a reward function, and 𝛽 > 0 a regularization parameter such that 1/ 𝛽 ∈ Λ (i.e., 𝔼𝑎 [ 𝑒𝑟/ 𝛽 ] < ∞; this is automatic when 𝑌 is finite, or more generally when 𝑟 is bounded above). Then max 𝔼𝑞 [ 𝑟 ] − 𝛽 KL( 𝑞 ∥ 𝑎) = 𝛽 𝐴 (1/ 𝛽 ) , 𝑞
where the maximum is over distributions 𝑞 with 𝔼𝑞 [| 𝑟 |] < ∞ and KL( 𝑞 ∥ 𝑎) < ∞, and the unique maximizer is the exponential family member 𝑞∗ = 𝑝1/ 𝛽 , i.e., 𝑞∗ ( 𝑦 ) ∝ 𝑎 ( 𝑦 ) 𝑒𝑟 ( 𝑦 )/ 𝛽 . Proof. Set 𝑑 = 1, 𝜙 = 𝑟 , 𝜆 = 1/ 𝛽 in Corollary 9 and multiply through by 𝛽 .
□
Remark 12. The exponential form of the optimal distribution in KL-regularized optimization has deep roots; the distribution 𝑝1/ 𝛽 is variously called the Boltzmann or Gibbs distribution in statistical physics, and the softmax or Boltzmann policy in reinforcement learning. In control theory, Todorov (2007) introduced linearly solvable MDPs with KL control costs, where the optimal policy takes this exponential form; related results were obtained independently by Kappen (2005) via path integral methods. In maximum entropy reinforcement learning, Ziebart (2010) derived the same form for the entropy-regularized setting (the special case 𝑎 = uniform); the soft actorcritic framework of Haarnoja et al. (2018) extends these ideas to the deep RL setting. The LLM fine-tuning literature largely rediscovered these results independently, with the exponential form appearing as both an optimal policy and a modeling target. Khalifa et al. (2021) give an early application that explicitly uses exponential families and moment constraints, framing fine-tuning as forward-KL minimization toward a target distribution; Ouyang et al. (2022) established the canonical RLHF setting; (Korbak et al., 2022a, Theorem 1) provide a clean statement of the equivalence between KL-regularized reward maximization and reverse-KL minimization toward 𝑝1/ 𝛽 , and Korbak et al. (2022b) offer a concurrent Bayesian interpretation of the same equivalence; Rafailov et al. (2023) exploit the exponential form to derive a closed-form policy optimization objective, connecting RLHF to direct preference optimization. Together, these threads illustrate how the same variational principle has been independently rediscovered across communities, each time yielding new algorithmic and conceptual insights.
3.6. Convexity of 𝐴 and the supporting hyperplane property Corollary 13. Assume that 𝜇 𝜆 exists for every 𝜆 ∈ Λ. Then, for all 𝜆 1 , 𝜆 2 ∈ Λ, 𝐴 ( 𝜆 2) ≥ 𝐴 ( 𝜆 1) + 𝜇 𝜆1 · ( 𝜆 2 − 𝜆 1) .
(6)
In other words, 𝐴 is convex on Λ, and the hyperplane through ( 𝜆 1 , 𝐴 ( 𝜆 1 )) with slope 𝜇 𝜆 1 is a global supporting hyperplane of 𝐴. Proof. The specialized identity (3) gives 𝐴 ( 𝜆 2 ) − 𝐴 ( 𝜆 1 ) − 𝜇 𝜆 1 · ( 𝜆 2 − 𝜆 1 ) = KL 𝑝𝜆 1 ∥ 𝑝𝜆 2 ≥ 0.
□
Remark 14. If the parametrization is injective ( 𝜆 1 ≠ 𝜆 2 implies 𝑝𝜆 1 ≠ 𝑝𝜆 2 ), then the inequality in (6) is strict for 𝜆 1 ≠ 𝜆 2 , since the gap equals KL 𝑝𝜆 1 ∥ 𝑝𝜆 2 , which vanishes only when 𝑝𝜆 1 = 𝑝𝜆 2 . Thus 𝐴 is strictly convex.
6
Exponential families from a single KL identity
3.7. The dual function 𝐴∗ and Legendre duality Define the Legendre dual of 𝐴 as 𝐴∗ ( 𝜇 ) = sup { 𝜆 · 𝜇 − 𝐴 ( 𝜆 )} . 𝜆 ∈Λ
By definition, 𝐴∗ is convex in 𝜇 (as a supremum of affine functions), and the Legendre–Fenchel inequality 𝐴 ( 𝜆 ) + 𝐴∗ ( 𝜇 ) ≥ 𝜆 · 𝜇 holds for all 𝜆, 𝜇 . Corollary 15 (Identification of 𝐴∗ ). If there exists 𝜆 ∈ Λ with 𝜇 𝜆 = 𝜇 , then the supremum in the definition of 𝐴∗ ( 𝜇 ) is attained at this 𝜆 , and 𝐴∗ ( 𝜇 ) = KL( 𝑝𝜆 ∥ 𝑎) .
(7)
Proof. From the specialized identity (3) with 𝜆 2 = 0 (so 𝑝𝜆 2 = 𝑎 and 𝐴 (0) = 0): KL( 𝑝𝜆 ∥ 𝑎) = 𝜆 · 𝜇 𝜆 − 𝐴 ( 𝜆 ) = 𝜆 · 𝜇 − 𝐴 ( 𝜆 ) . So 𝜆 · 𝜇 − 𝐴 ( 𝜆 ) = KL( 𝑝𝜆 ∥ 𝑎). It remains to show this is the supremum over all 𝜆 ′ ∈ Λ. For any 𝜆 ′ ∈ Λ, the supporting hyperplane property (6) with 𝜇 𝜆 1 = 𝜇 𝜆 = 𝜇 gives: 𝐴( 𝜆′) ≥ 𝐴( 𝜆) + 𝜇 · ( 𝜆′ − 𝜆),
which rearranges to 𝜆 ′ · 𝜇 − 𝐴 ( 𝜆 ′ ) ≤ 𝜆 · 𝜇 − 𝐴 ( 𝜆 ) = KL( 𝑝𝜆 ∥ 𝑎). Hence 𝐴∗ ( 𝜇 ) = KL( 𝑝𝜆 ∥ 𝑎) ≥ 0, with the nonnegativity following from KL ≥ 0. □ Remark 16. The identification 𝐴∗ ( 𝜇 ) = KL( 𝑝𝜆 ∥ 𝑎) gives a probabilistic interpretation of the Legendre dual: 𝐴∗ ( 𝜇 ) is the KL divergence from the moment-matching exponential family member 𝑝𝜆 to the base distribution 𝑎. Together with 𝐴 ( 𝜆 ) = 𝜆 · 𝜇 − 𝐴∗ ( 𝜇 ) (from the proof above), this recovers the Legendre–Fenchel relation 𝐴 ( 𝜆 ) + 𝐴∗ ( 𝜇 ) = 𝜆 · 𝜇 as an identity (not merely an inequality) at the matching pair ( 𝜆, 𝜇 ).
4. Analytic complements The results of Section 3 required nothing beyond KL(· ∥ ·) ≥ 0. We now collect the results that require one additional analytic input: the differentiability of 𝐴 on int( Λ), and a geometric argument for the surjectivity of the moment map.
4.1. Differentiability and the gradient of 𝐴 Remark 17. The log-partition function 𝐴 ( 𝜆 ) is in fact differentiable on int( Λ), with ∇ 𝐴 ( 𝜆 ) = 𝜇 𝜆 . Í When 𝑌 is finite, this is immediate since 𝐴 ( 𝜆 ) = log 𝑦 𝑎 ( 𝑦 ) 𝑒𝜆 ·𝜙 ( 𝑦 ) is a finite sum of smooth functions. In general, it follows from differentiating under the integral sign, justified by dominated convergence using the fact that 𝑍 𝜆 converges in a neighborhood of any 𝜆 ∈ int( Λ); see, e.g., (Brown, 1986, Chapter 2) or (Wainwright and Jordan, 2008, Proposition 3.1). This is the only analytic input in this note that does not follow from the identity.
4.2. Bregman divergence interpretation Given differentiability, the specialized identity (3) acquires a classical interpretation. Recall that for a differentiable convex function 𝐹 , the Bregman divergence is defined as 𝐵 𝐹 ( 𝑥, 𝑦 ) = 𝐹 ( 𝑥 ) − 𝐹 ( 𝑦 ) − ∇ 𝐹 ( 𝑦 ) · ( 𝑥 − 𝑦 ) Bregman (1967). Using ∇ 𝐴 ( 𝜆 1 ) = 𝜇 𝜆 1 (Remark 17), the specialized identity (3) becomes KL 𝑝𝜆 1 ∥ 𝑝𝜆 2 = 𝐴 ( 𝜆 2 ) − 𝐴 ( 𝜆 1 ) − ∇ 𝐴 ( 𝜆 1 ) · ( 𝜆 2 − 𝜆 1 ) , (8)
7
Exponential families from a single KL identity
which is precisely the Bregman divergence of the log-partition function evaluated at ( 𝜆 2 , 𝜆 1 ). This is the well-known representation of within-family KL divergence as a Bregman divergence Banerjee et al. (2005). Remark 18. With the Bregman interpretation in hand, the generalized three-point identity (Corollary 3) can be compared with the standard three-point property of Bregman divergences Banerjee et al. (2005); see also Nielsen and Nock (2010). When 𝑞 = 𝑝𝜆 0 for some 𝜆 0 ∈ Λ, identity (4) reduces to the standard Bregman three-point property, with each KL divergence replaced by the corresponding Bregman divergence of 𝐴 ( 𝜆 ) via (8): KL 𝑝𝜆 0 ∥ 𝑝𝜆 2 = KL 𝑝𝜆 0 ∥ 𝑝𝜆 1 + KL 𝑝𝜆 1 ∥ 𝑝𝜆 2 + ( 𝜇 𝜆 0 − 𝜇 𝜆 1 ) · ( 𝜆 1 − 𝜆 2 ) . However, the general form (4), where 𝑞 is an arbitrary distribution outside the exponential family, does not seem to be readily found in the literature. This generality is what makes the Pythagorean theorem (Corollary 5) and the maximum entropy characterization (Corollary 7) directly applicable to reverse I-projections of arbitrary distributions onto the exponential family. The Pythagorean theorem for exponential families is classically developed in the framework of information geometry Amari (1985),Amari and Nagaoka (2000), using the dually flat structure of the manifold. A related Pythagorean property for I-projections in a more general setting (without explicit exponential family structure) was established by Csiszár (1975).
4.3. Surjectivity of the moment map Several results in Section 3 are conditional on the existence of a parameter 𝜆 with 𝜇 𝜆 = 𝜇 . In the discrete setting this can be shown directly. Assume throughout this subsection that 𝑌 is finite or countably infinite, so that 𝑀 = conv( 𝜙 (𝑌 )) .
Proposition 19. Assume Λ = ℝ𝑑 (which holds automatically when 𝑌 is finite, and is an additional assumption when 𝑌 is countably infinite, e.g. when 𝜙 is bounded). For any 𝜇 ∈ int( 𝑀 ), there exists 𝜆 ∗ ∈ ℝ𝑑 such that 𝜇 𝜆 ∗ = 𝜇 . Proof. Set 𝑓 ( 𝜆 ) = 𝐴 ( 𝜆 ) − 𝜆 · 𝜇 . By Remark 17, ∇ 𝑓 ( 𝜆 ) = 𝜇 𝜆 − 𝜇 , so it suffices to show that 𝑓 attains its minimum. Growth at infinity. Since 𝜇 ∈ int( 𝑀 ) ⊂ conv( 𝜙 (𝑌 )), by Carathéodory’s theorem there exist finitely many points 𝑦1 , . . . , 𝑦𝑚 ∈ 𝑌 such that 𝜇 ∈ int(conv{𝜙 ( 𝑦1 ) , . . . , 𝜙 ( 𝑦𝑚 )}). Hence there exists 𝑟 > 0 with 𝐵 ( 𝜇, 𝑟 ) ⊂ conv{𝜙 ( 𝑦1 ) , . . . , 𝜙 ( 𝑦𝑚 )}. For any 𝜆 ≠ 0, let 𝑢 = 𝜆 /∥ 𝜆 ∥. Since 𝜇 + 𝑟𝑢 lies in conv{𝜙 ( 𝑦1 ) , . . . , 𝜙 ( 𝑦𝑚 )}, there exists some 𝑖 ∈ {1, . . . , 𝑚 } with 𝑢 · 𝜙 ( 𝑦𝑖 ) ≥ 𝑢 · 𝜇 + 𝑟. Using this 𝑦𝑖 and the definition of 𝐴, 𝐴 ( 𝜆 ) ≥ 𝜆 · 𝜙 ( 𝑦𝑖 ) + log 𝑎 ( 𝑦𝑖 ) ,
and therefore 𝑓 ( 𝜆 ) ≥ 𝑟 ∥ 𝜆 ∥ + log 𝑎 ( 𝑦𝑖 ) ≥ 𝑟 ∥ 𝜆 ∥ + min log 𝑎 ( 𝑦𝑖 ) . 1≤ 𝑖 ≤ 𝑚
Thus 𝑓 ( 𝜆 ) → +∞ as ∥ 𝜆 ∥ → ∞.
8
Exponential families from a single KL identity
Conclusion. Under the assumption Λ = ℝ𝑑 , the function 𝐴 is smooth (Remark 17), so 𝑓 is continuous. The linear lower bound shows that for large enough 𝑅, the minimum of 𝑓 over all of ℝ𝑑 is achieved inside the closed ball ∥ 𝜆 ∥ ≤ 𝑅; since 𝑓 is continuous on this compact set, it attains its minimum there at some 𝜆 ∗ . Then ∇ 𝑓 ( 𝜆 ∗ ) = 0 gives 𝜇 𝜆 ∗ = 𝜇 . □ The analogous result for exponential families on general measurable spaces is established in (Brown, 1986, Chapter 3).
5. Discussion This note is organized around a sharp separation between algebraic consequences of the KL difference identity (1) (Section 3), which require nothing beyond KL(· ∥ ·) ≥ 0, and analytic complements (Section 4), which additionally use the differentiability of 𝐴 on int( Λ). The identity itself is a one-line calculation: the log-ratio log( 𝑝𝜆 1 / 𝑝𝜆 2 ) is affine in 𝜙, so its expectation under any 𝑞 involves only 𝐴 ( 𝜆 1 ), 𝐴 ( 𝜆 2 ), and 𝜇 𝑞 . The purely algebraic results—the generalized three-point identity, the Pythagorean theorem, the I-projection and reverse I-projection characterizations, the convexity of 𝐴 and identification of 𝐴∗ , the ELBO decomposition, and KL-regularized reward maximization—are all obtained by the same recipe: choose 𝑞, 𝜆 1 , 𝜆 2 in the identity, and apply KL(· ∥ ·) ≥ 0 with its equality characterization. The analytic section adds one external input—differentiability of 𝐴—which enables the Bregman divergence interpretation, connects to the classical literature, and provides the gradient characterization ∇ 𝐴 ( 𝜆 ) = 𝜇 𝜆 needed for the surjectivity proof: • The convexity of 𝐴 and the supporting hyperplane property are algebraic; promoting the slope 𝜇 𝜆 to a gradient requires differentiability, but a mild argument suffices: term-by-term differentiation for finite 𝑌 , or dominated convergence in general Brown (1986),Wainwright and Jordan (2008). • The generalized three-point identity (4), proved algebraically, is recognized as a generalization of the Bregman three-point property of Banerjee et al. (2005) once the Bregman interpretation is available. • The Pythagorean theorem is classically developed using the dually flat structure of information geometry Amari (1985),Amari and Nagaoka (2000). Here it is a direct consequence of the three-point identity applied to arbitrary distributions 𝑞 outside the family, a generality not readily found in the standard Bregman divergence literature. A related Pythagorean property for I-projections was established by Csiszár (1975). • Surjectivity of the moment map requires differentiability of 𝐴 together with a standard geometric argument about the interior of the moment space M. We do not address the rich analytic structure of 𝐴 (higher-order cumulants, the Fisher information metric, analyticity on int( Λ)), for which we refer to Brown (1986),Wainwright and Jordan (2008). Our aim is limited to showing how far a single algebraic identity can reach. This note arose from the study of KL-regularized control problems, where the identity provides a natural framework for analyzing the geometry of optimal policies. This connection will be developed elsewhere.
9
Exponential families from a single KL identity
AI Disclosure The author used Claude (Anthropic, claude.ai) and ChatGPT (OpenAI) during the preparation of this manuscript for assistance with exposition, structuring arguments, and reviewing text and proof drafts. The author reviewed and edited all AI-assisted content and takes full responsibility for the content of this paper.
References S.-i. Amari. Differential-Geometrical Methods in Statistics, volume 28 of Lecture Notes in Statistics. Springer, 1985. 8, 9 S.-i. Amari and H. Nagaoka. Methods of Information Geometry. Translations of Mathematical Monographs. American Mathematical Society, 2000. 8, 9 A. Banerjee, S. Merugu, I. S. Dhillon, and J. Ghosh. Clustering with Bregman divergences. Journal of Machine Learning Research, 6:1705–1749, 2005. 8, 9 O. Barndorff-Nielsen. Information and Exponential Families in Statistical Theory. Wiley, 1978. 1 D. M. Blei, A. Kucukelbir, and J. D. McAuliffe. Variational inference: A review for statisticians. Journal of the American Statistical Association, 112(518):859–877, 2017. 5 L. M. Bregman. The relaxation method of finding the common point of convex sets and its application to the solution of problems in convex programming. USSR Computational Mathematics and Mathematical Physics, 7(3):200–217, 1967. 7 L. D. Brown. Fundamentals of Statistical Exponential Families with Applications in Statistical Decision Theory. Institute of Mathematical Statistics, Hayward, CA, 1986. 1, 7, 9 I. Csiszár. I-divergence geometry of probability distributions and minimization problems. The Annals of Probability, 3(1):146–158, 1975. 8, 9 I. Csiszár and F. Matúš. Information projections revisited. IEEE Transactions on Information Theory, 49(6):1474–1490, 2003. 4 T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In Proceedings of the 35th International Conference on Machine Learning (ICML), pages 1861–1870, 2018. 3, 6 M. I. Jordan, Z. Ghahramani, T. S. Jaakkola, and L. K. Saul. An introduction to variational methods for graphical models. Machine Learning, 37(2):183–233, 1999. 5 M. Khalifa, H. Elsahar, and M. Dymetman. A distributional approach to controlled text generation. In 9th International Conference on Learning Representations (ICLR), 2021. 6 H. J. Kappen. Path integrals and symmetry breaking for optimal control theory. Journal of Statistical Mechanics: Theory and Experiment, 2005(11):P11011, 2005. 6 T. Korbak, H. Elsahar, G. Kruszewski, and M. Dymetman. On reinforcement learning and distribution matching for fine-tuning language models with no catastrophic forgetting. In Advances in Neural Information Processing Systems (NeurIPS), 2022a. 6 T. Korbak, E. Perez, and C. L. Buckley. RL with KL penalties is better viewed as Bayesian inference. arXiv preprint arXiv:2205.11275, 2022b. 6 F. Nielsen and R. Nock. Entropies and cross-entropies of exponential families. In Proceedings of the IEEE International Conference on Image Processing (ICIP), pages 3621–3624, 2010. 8 L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems (NeurIPS), 2022. 6 Y. Polyanskiy and Y. Wu. Information Theory: From Coding to Learning. Cambridge University Press, 2025. 5
10
Exponential families from a single KL identity
R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn. Direct preference optimization: Your language model is secretly a reward model. In Advances in Neural Information Processing Systems (NeurIPS), 2023. 6 E. Todorov. Linearly-solvable Markov decision problems. In Advances in Neural Information Processing Systems (NIPS), pages 1369–1376, 2007. 6 M. J. Wainwright and M. I. Jordan. Graphical models, exponential families, and variational inference. Foundations and Trends in Machine Learning, 1(1–2):1–305, 2008. 1, 7, 9 B. D. Ziebart. Modeling Purposeful Adaptive Behavior with the Principle of Maximum Causal Entropy. PhD thesis, Carnegie Mellon University, 2010. 3, 6
11