Conceptio › Archive › arXiv CS
arXiv CSopen access

A Generalization of Amari's Bayesian Duality

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

A Generalization of Amari’s Bayesian Duality Mohammad Emtiyaz Khan1,2,3 and Thomas Möllenhoff1 1 RIKEN Center for Advanced Intelligence Project, 1-4-1 Nihonbashi,

Chuo-ku, Tokyo, 103-0027, Tokyo, Japan. 2 Department of Computer Science, Technische Universität Darmstadt,

arXiv:2609.09126v1 [cs.AI] 8 Sep 2026

Hochschulstraße 10, Darmstadt, 64289, Hessen, Germany. 3 The Hessian Center for Artificial Intelligence, Landwehrstraße 50a,

Darmstadt, 64293, Hessen, Germany.

Contributing authors: [email protected]; [email protected]; Abstract Amari’s contributions to information geometry and machine learning are well known. Here, we revisit Amari’s work on Bayesian duality which has not received as much attention. We connect Amari’s Bayesian duality to a convex duality of Bayes’ rule. Using this connection, we present a generalization of Amari’s Bayesian duality and discuss its relevance for modern artificial intelligence. Keywords: Bayes’ Rule, Information Geometry, Convex Duality

1 Introduction Amari has made many significant contributions to information geometry and machine learning. He was one of the first to use stochastic gradient descent to train neural networks (Amari, 1967). He proposed recurrent neural networks that were earlier versions of what is now called Hopfield networks (Amari, 1979). He also established a connection between em and EM algorithms (Amari, 1995a,b) and put forward a proposal to use natural-gradient methods to train neural networks (Amari, 1998). All these works (and many others) are now foundational concepts in machine learning and have led to new advances in artificial intelligence. Here, we revisit a relatively less known work by Amari (1996) on the development of a Bayesian duality theory. This work aimed to explain the ‘basic but primitive’

1

Posterior manifold

p(θ | y)

Likelihood manifold

Bijection

Posterior manifold

p(y | θ)

q(θ)

(a) Amari’s Bayesian duality

Likelihood manifold

Convex Duality

t(θ)

(b) Our generalization

Fig. 1: Amari (1996) proposed a dual structure connecting the manifolds of posterior and likelihood (Panel a). In his case, a likelihood function p(y|θ) over data vector y given parameter vector θ is connected to a unique posterior p(θ|y) and vice versa. His framework relied on a simple posterior form. We present a generalization by using convex duality which not only recovers Amari’s case but applies much more generally.

mechanisms of information processing in the brain and was motivated by the recent works of that time on models such as mixture-of-expert nets, the Helmholtz machine, and the Ying-Yang machine. The work did not receive much attention back then and the focus of the machine-learning community also partly drifted away from such ideas. Our goal here is to connect Amari’s Bayesian duality to other works in the machine-learning literature, so as to generalize its scope and discuss its relevance. We will start with a description of the basic setup of Amari (1996) where he introduced a new duality structure for the Bayesian framework from the point of view of information geometry. He focused on a specific case where the likelihood and posterior have the same exponential-family (EF) forms and showed that there is a bijection between the two quantities where the roles of their dual coordinates are interchanged (Fig. 1a). The theory however does not apply to the majority of Bayesian cases. For example, in conjugate Bayesian models, the form of the posterior matches the prior, not the likelihood. We will extend the scope of Amari’s theory by proposing a more general Bayesian duality through a different mathematical framework that relies on the convex duality of a variational formulation of Bayes’ rule (Fig. 1b). This will also enable us to connect to a much broader literature in Bayesian inference and convex optimization. We will conclude the paper by discussing the relevance of Bayesian duality for modern AI.

2 Amari’s Bayesian Duality Theory We start with a brief description of Amari’s Bayesian duality theory. With a focus on the mechanisms of information processing in the brain, Amari attempted to explain the dynamic interactions between lower and higher systems. We will first discuss his framework and return to the original motivation again later in the last section. Amari introduced a new duality structure for a simple case of Bayesian inference by using ideas from information geometry. In his example, we have an observation vector y consisting of scalar entries yi which are connected to the parameter vector θ

2

of scalar entries θi . The two vectors are assumed to be of the same size (although an extension to different lengths was also briefly mentioned). Amari described a duality theory associated with the likelihood function and posterior density, denoted as, p(y | θ)

and

p(θ | y),

(1)

respectively. The notation p is used to denote both densities over y and θ, which is an abuse of notation for simplicity but is standard practice in many Bayesian texts.

2.1 The Likelihood Function Amari considered a simple case where both the likelihood and posterior take a canonical exponential-family (EF) form. Let us define the likelihood to be

h i p(y | θ) = hlik (y) exp θ ⊤ y − Alik (θ) .

(2)

This is a familiar form: Alik (θ) is the log-partition function and hlik (y) is a function of y, sometimes referred to as the base measure (Wainwright and Jordan, 2008). The set of such densities is a manifold (denote it by Slik ) with two coordinate systems consisting of natural parameters (denoted by λlik ) and expectation parameters (denoted by µlik ), respectively. For the EF in Eq. 2, we have the following: λlik = θ,

µlik = Ep(y|θ) [y].

(3)

It is a well-known fact that these two quantities provide two different but equivalent ways to parameterize EFs (for example, see Barndorff-Nielsen (1978)). The coordinates are dual of each other due to their relationship through the Legendre transform. Specifically, we have µlik = ∇Alik (λlik ) and λlik = ∇A∗lik (µlik ) where A∗lik is the convex dual of Alik . In information geometry, this is referred to as the dually-flat structure.

2.2 The Posterior Distribution Now, let us introduce the Bayesian framework where a posterior density is obtained by using the above likelihood along with a prior denoted by p(θ). Using Bayes’ rule, we can write the posterior density as p(θ | y) =

p(y | θ)p(θ) , p(y)

(4)

R where the marginal density is denoted by p(y) = p(y | θ)p(θ)dθ. Amari considered a case where the form of the posterior distribution p(θ | y) with respect to θ is the same as the form of p(y | θ) with respect to y. This differs from the more popular case of conjugate-Bayes where this is not necessarily the case. There, the form of the posterior matches the form of the prior. The likelihood can be expressed in the same form with respect to θ, but it is rarely the case that is has the same form also with respect to y. Amari considered a restrictive case where this is possible by 3

absorbing the prior into the base measure of the likelihood. This is shown below where we first substitute the likelihood from Eq. 2, then rearrange the numerator to write it in a similar form as the likelihood:

h i hlik (y) exp θ ⊤ y − Alik (θ) p(θ) p(y) h i = p(θ) exp[−Alik (θ)] exp θ ⊤ y − (log p(y) − log hlik (y)) |{z} | | {z } {z } =λpost =hpost (θ) =Apost (y)   = hpost (θ) exp y⊤ θ − Apost (y) .

p(θ | y) =

(5)

In the last line, we redefined the base measure and log-partition function to write the posterior in the exact same form as the likelihood in Eq. 2. Note that this rearrangement is only possible if absorbing p(θ) still defines a valid EF distribution where y is the natural parameter (which is rarely the case as mentioned earlier). However, when it is indeed possible to do so, then, similarly to the likelihood, the posterior also forms a manifold (denote it by Spost ) with dual-coordinate systems consisting of the natural and expectation parameters defined below: λpost = y,

µpost = Ep(θ|y) [θ].

(6)

Together, the likelihood and posterior give rise to two distinct manifolds.

2.3 Bayesian Duality Amari (1996) showed that the two manifold structures associated with the likelihood and posterior, respectively, are connected to each other through a bijection (Fig. 1a) and the roles of the dual coordinates are also interchanged. This can be more clearly seen in the table below which summarizes the density, natural parameters, sufficient statistics, and expectation parameters, respectively. If all instances of y and θ are swapped in the top row, we get the bottom row. Name Likelihood Posterior

Density   p(y|θ) = hlik (y) exp θ ⊤ y − Alik (θ)   p(θ|y) = hpost (θ) exp y⊤ θ − Apost (y)

Nat. param.

Suff. Stat.

Exp. param.

θ

y

Ep(y|θ) [y]

y

θ

Ep(θ|y) [θ]

This connection between the likelihood and posterior is formalized by Amari via the new notion of Bayesian duality. Due to bijection, there is a unique mapping between the two manifolds. Therefore, for a likelihood p(y | θ) in Slik , there exists a unique posterior p(θ | y) in Spost , and vice versa. The roles of natural parameters and sufficient statistics are swapped. Amari referred to this as a new dual structure associated with Bayesian inference, giving rise to a new Bayesian duality theory. To further illustrate the point, we give a simple example of the dual structure for a case where both likelihood and posterior are isotropic Gaussians.

4

Ex 1. Consider a Gaussian likelihood function for y written in an EF form:

h i p(y|θ) = N (y|θ, I) = (2π)−D/2 exp(− 21 y⊤ y) exp y⊤ θ − 12 θ ⊤ θ , | {z } | {z } hlik (y)

(7)

Alik (θ)

For this likelihood, the natural parameter is λlik = θ and the expectation parameter is also the same because µlik = ∇Alik (θ) = θ. The dual-coordinate pair is simply (λlik , µlik ) = (θ, θ). Let us turn to the posterior density next. We want the posterior to have the same form as the likelihood. A Gaussian prior ensures that the posterior is Gaussian as well, but it does not ensure that the posterior covariance is also identity. However, this happens if we choose a uniform prior. In that case, the marginal density is also uniform, as shown below, Z Z p(y) = N (y | θ, I)dθ = N (θ | y, I)dθ = 1. (8) As a result, we can write the posterior as an isotropic Gaussian density: p(θ | y) =

N (y | θ, I)p(θ) = N (θ | y, I). p(y)

(9)

This also shows that the posterior mean is simply equal to y. By writing the density in the EF form as in Eq. 7 we can show that the natural parameter is also equal to y and that the posterior takes the same form as the likelihood:

h i p(θ | y) = (2π)−D/2 exp(− 12 θ ⊤ θ) exp θ ⊤ y − 12 y⊤ y , | {z } | {z } hpost (θ)

(10)

Apost (y)

Therefore, the dual-coordinate pair is (λpost , µpost ) = (y, y). As Amari pointed out, the roles of θ and y are interchanged for manifolds Slik and Spost and there is a trivial bijection between them.

2.4 How to Generalize Amari’s Framework? The main limitation of Amari’s Bayesian duality is the restriction that the likelihood and posterior need have the same form. Amari did show an example on the Boltzmann machine but applying it to even simple machine-learning models runs into difficulties. We will now discuss this difficulty for a simple ridge regression problem. We will show that a dual structure can still be found but handling the general cases requires a different mathematical framework. In the next section, we will present one such framework based on convex duality. Consider a one-dimensional ridge regression problem with a scalar parameter θ ∈ R to model N scalar outputs yi ∈ R given scalar inputs xi ∈ R. We denote the vectors of outputs and inputs by y and x respectively (both of length N ). We assume a Gaussian 5

likelihood with variance 1 and a standard Gaussian prior: p(y | θ) = N (y | xθ, I)

p(θ) = N (θ | 0, 1).

(11)

With this choice, the posterior is a Gaussian density as well: p(θ | y) = N (θ | m∗ , σ∗2 ) where m∗ = σ∗2 x⊤ y and σ∗2 = 1/(x⊤ x + 1).

(12)

Unlike the likelihood, the posterior has a variance that is not equal to 1. The likelihood and posterior therefore take different EF forms, and interchanging the roles of the dual coordinates does not make much sense. However, it turns out that we can still find a mapping between the two manifolds. We will first describe the interchanging process and give a formal mathematical framework in the next section. The key idea is to express the likelihood in terms of the posterior’s sufficient statistics. For example, for the posterior in Eq. 12, the sufficient statistic is a 2D vector T(θ) = (θ, θ2 )⊤ ,

(13)

and it is therefore possible to write the posterior in an EF form as p(θ | y) ∝ exp(⟨λpost , T(θ)⟩),

(14)

where ⟨·, ·⟩ is an inner product and λpost is the natural parameter of the posterior; see Khan and Rue (2023, Sec. 5). Our idea is to first write the likelihood in the same e lik such that form, for example, we can find a vector λ

e lik , T(θ)⟩), p(y | θ) ∝ t(θ) = exp(⟨λ

(15)

and then write a dual problem as an inverse mapping to find t(θ), giving rise to a dual structure. This is the basic idea of our setup. We will now show how to write the likelihood in this form and that doing so yields an interchanging of the coordinates, which is similar in spirit to Amari’s case. Ex 2. To show this, we first write the likelihood in the canonical EF form as in the previous example and then redefine the sufficient statistics and natural parameters to write them as an inner product,

  N p(y | θ) = (2π)− 2 exp(− 12 y⊤ y) exp y ⊤ (xθ) − 21 x⊤ xθ2 |{z} |{z} | {z } | {z } Tlik (y) λlik hlik (y) Alik (λlik ) + # "*    xθ y   ,  = hlik (y) exp − 21 x⊤ xθ2 1 |{z} | {z } b lik (y) T

b lik λ

6

(16)

In the second line, we have defined an inner product between (N + 1)-dimensional b lik vectors where we denote the new sufficient statistics and natural parameter by T b and λlik , respectively. Next, we write the term inside the exponential as an inner product between 2D vectors which yields a form in terms of T(θ) on the right,

+ *   + *   y xθ x⊤ y θ   ,  =  ,   . 1 ⊤ 1 ⊤ 2 1 − 2 x xθ −2x x θ2 |{z} | | {z } | {z } {z } b lik (y) T

b lik λ

e lik λ

(17)

T(θ)

The whole rearrangement can be more compactly written as

b lik ⟩ = ⟨λ e lik , T(θ)⟩, b lik (y), λ ⟨T

(18)

which shows an interchanging of the sufficient statistic and natural parameter. The example suggests that, do find a dual mapping, we rewrite Bayes’ rule in a form where we replace the likelihood by its mapping t(θ) in the space of T(θ). p(θ | y) ∝ t(θ)p(θ)

(19)

This way the forward map is simply Bayes’ rule, while the reverse map is to find t(θ) given the posterior p(θ | y). We will show next that these operations are written as duals of each other through convex duality. We will use it to generalize Amari’s Bayesian duality to general Bayesian cases.

3 Bayesian Duality via Convex Duality We will now derive a dual structure of Bayes’ rule by using convex duality. This framework removes the restrictions imposed in Amari’s work. We do not claim that this solution is better than using information geometry. Rather, we think that this viewpoint is helpful to connect Amari’s work to other works that also use convex duality for Bayesian inference. The key idea is to write Bayes’ rule as an optimization problem and then use its dual to find an inverse mapping to search for the likelihood e T(θ)⟩) in a dual space. mapping t(θ) = exp(⟨λ,

3.1 Variational Formulation of Bayes’ Rule We start by writing a variational form of Bayes’ rule, also known as Gibbs’ variational principle in information theory. It is well known that Bayes’ rule can be written as an optimization problem in the space of all densities (denoted by P) over θ. Consider the following optimization problem over candidate densities q ∈ P: q∗ (θ) = arg sup Eq(θ) [log p(y|θ)] − DKL (q(θ) ∥ p(θ)), q∈P

7

(20)

where the second term on the right is the Kullback-Leibler divergence (KLD) between the candidate q(θ) and the prior p(θ). It is well-known that the optimal q∗ is the posterior density, that is, q∗ (θ) = p(θ | y). One can also show that the optimal value is simply the log-partition function. More precisely, we have Z log p(y|θ)p(θ)dθ = sup Eq(θ) [log p(y|θ)] − DKL (q(θ) ∥ p(θ)). (21) q∈P

The left hand side is the log-partition function log p(y). The result is easy to verify by simply setting q = q∗ on the right-hand side:

 Eq∗ (θ) [log p(y|θ)] − DKL (q∗ (θ) ∥ p(θ)) = Eq∗ (θ)

 p(y|θ)p(θ) log = log p(y). q∗ (θ)

(22)

This result is attributed to Gibbs (1902) and is sometimes referred to as the Gibbs variational principle. The connection to Bayes’ rule was formalized by Jaynes, starting from his work on the maximum-entropy principle (Jaynes, 1957). A general result that goes beyond likelihood functions and applies to generic losses is given in a series of papers by Donsker and Varadhan (1975a,b, 1976, 1983).

3.2 A Dual of the Gibbs Variational Formulation Amari’s bijection between posterior and likelihood can be derived as a special case of a dual version of the Gibbs variational principle. The goal of the dual problem is to find a likelihood function given a posterior distribution q(θ) ∈ P. For this purpose, let us define another space P ∗ consisting of valid likelihood functions. For instance, we can choose potential likelihoods t(θ) from a function space F for which the log-partition is finite. Formally, we define P∗ =



Z t ∈ F : log

 t(θ)p(θ)dθ < ∞ .

Given such a space, we can write the dual problem as the following: Z DKL (q(θ) ∥ p(θ)) = sup Eq(θ) [log t(θ)] − log t(θ)p(θ)dθ.

(23)

(24)

t∈P ∗

This dual problem recovers the KLD between the posterior q(θ) and the prior p(θ) by solving an optimization problem over a function space P ∗ . We propose it as a generalization of Amari’s Bayesian duality (Fig. 1b). We refer to Eq. 24 as the dual of the Gibbs’ variational principle, even though different names are used in the literature. For instance, in information theory, this is sometimes referred to as the DonskerVaradhan formula (Polyanskiy and Wu, 2025, Thm. 4.6). For a given y, if we set q(θ) = q∗ (θ) (the posterior p(θ | y) from Eq. 20), then the assertion is that an optimal solution t∗ (θ) recovers the likelihood function p(y | θ) up to a constant factor that is absorbed in the normalization. More precisely, when

8

t∗ (θ) is plugged in Bayes’ rule, it will yield the posterior p(θ | y). The equality can be verified by substituting the optimal t(θ) = p(y | θ) and q = q∗ ,



Z Eq∗ (θ) [log p(y|θ)] − log

p(y|θ)p(θ)dθ = Eq∗ (θ)

   p(y|θ) q∗ (θ) log = Eq∗ (θ) log . p(y) p(θ)

3.3 Amari’s Bayesian Duality via Convex Duality Before giving more details, we connect the above duality to Amari’s bijection by applying it to the Gaussian case discussed in Ex 1. Ex 3. We start by writing the variational form of Eq. 20. Because we know that the posterior is an isotropic Gaussian, we can simplify the optimization by restricting it to that space, that is, we set q(θ) = N (θ | m, I). This enables us to write the optimization problem over the mean m. This is shown below: m∗ = arg sup EN (θ | m,I) [log N (y | θ, I)] − DKL (N (θ | m, I) ∥ p(θ)) m

= arg sup − 12 (y − m)⊤ (y − m).

(25)

m

We see that Bayes’ rule in this case reduces to optimization of a quadratic function with maximum m∗ = y. A similar derivation can be used to write the dual form as a quadratic problem. Mimicking the form of the likelihood, we define log t(θ) to be a quadratic,

  ⊤ e θ − 1 θ⊤ θ , t(θ) = exp λ 2

(26)

e as the free parameter. Given a posterior q∗ = N (θ | m∗ , I), our goal then with λ e ∗ . The dual problem in Eq. 24 can then be written as is to find the optimal λ Z e ∗ = arg sup EN (θ|m ,I) [log t(θ)] − log t(θ)p(θ)dθ λ ∗ e λ

Z i  ⊤  h ⊤ e θ − 1 θ ⊤ θ p(θ)dθ e θ − 1 θ ⊤ θ − log exp λ = arg sup EN (θ|m∗ ,I) λ 2 2 e λ

e ⊤ m∗ − 1 m⊤ m∗ − D log(2π) − 1 λ e⊤λ e = arg sup λ ∗ 2 2 2 e λ e ⊤ m∗ − 1 λ e ⊤ λ. e = arg sup λ 2 e λ

e ∗ = m∗ = y. Therefore, we have This is a quadratic function whose solution is λ t∗ (θ) ∝ N (y | θ, I). We note that the function t(θ) need not be a probability density, but its role is identical. If we replace p(y | θ) in Eq. 20 by t∗ (θ), then we recover the posterior q∗ (θ) which is also equal to the posterior p(θ | y).

9

This example illustrates that the primal-dual formulation of the Gibbs’ variational principle can realize the bijection proposed by Amari when we use the specific form of the likelihoods and posterior. This dual version however applies much more generally. In fact, it applies even when likelihoods are replaced by generic loss functions. We will now show an application to the ridge regression case.

3.4 Bayesian Duality of Ridge Regression We return to the example of ridge regression shown in Eq. 11. Amari’s framework is difficult to apply to this case and we will show that by using the dual problem over t(θ), we can solve this issue. We showed in Eq. 17 of Ex 2 that the optimal likelihood e lik = (x⊤ y, − 1 x⊤ x)⊤ . We will now show that this solution mapping is a 2D vector: λ 2 can be recovered by solving the dual problem. Ex 4. We restrict log t(θ) to be a quadratic function with two scalar parameters e1 and λ e2 , as shown below: λ

  e1 θ − 1 λ e 2 . t(θ) = exp λ 2 2θ

(27)

e1,∗ and λ e2,∗ . Given a posterior q∗ = N (θ | m∗ , σ∗2 ), our goal is to find the optimal λ The dual problem in Eq. 24 can be written as Z e λ∗ = arg sup EN (θ|m∗ ,σ∗2 ) [log t(θ)] − log t(θ)p(θ)dθ e λ 2 2 e e1 m∗ − 1 λ = arg sup λ 2 2 (m∗ + σ∗ ) − log

Z

  e1 θ − 1 λ e2 θ2 p(θ)dθ. exp λ 2

(28)

e λ

The integral is available in closed form:

Z log

  e1 θ − 1 λ e2 θ2 p(θ)dθ = − 1 log(λ e2 + 1) + 1 exp λ 2 2 2

e2 λ 1 e2 + 1 λ

.

(29)

Substituting this in the optimization problem and taking derivatives, we get

e1,∗ = m∗ /σ 2 = x⊤ y, λ ∗

e2,∗ = 1/σ 2 − 1 = x⊤ x, λ ∗

(30)

which is the desired solution. A detailed derivation is in Sec. A.1. The dual problem yields the optimal mapping for the likelihood in the T(θ) space. Using the interchanging operation shown in Eq. 18 it may be possible to then write an appropriate likelihood. Unlike Amari’s example, this operation will not be a bijection in general because there could be several ways to aggregate the information. It is also possible to write the dual problem directly in the data space. For instance, we can rewrite the dual problem in the space of the data vector y which is of length

10

N . This is the space of linear predictions denoted by f = xθ.

(31)

This formulation can recover the likelihood exactly, as shown next. Ex 5. The dual problem can be written in the f -space by using a pushforward of the distribution. Specifically, we can write the prior and posterior in the f -space using a change of variable. p(f ) = N (f | 0, xx⊤ ),

q∗ (f ) = N (f | xm∗ , σ∗2 xx⊤ ).

(32)

The prior is degenerate but that does not pose a problem as long as the integrals with respect to it are finite. The likelihood candidate can also be mapped in that space by redefining the free parameters to be of appropriate sizes:

  t(f ) = exp a⊤ f − 12 f ⊤ Bf ,

(33)

where a is an N -length real vector while B is an N × N real matrix. With these, the dual problem can be written in the f -space, as an optimization over a and B, Z (a∗ ,B∗ ) = arg sup Eq∗ (f ) [log t(f )] − log t(f )p(f )df (34) a,B Z   = arg sup a⊤ xm∗ − 12 x⊤ Bx(m2∗ + σ∗2 ) − log exp a⊤ f − 12 f ⊤ Bf p(f )df . a,B

The integral in the last term can be obtained by changing back the variable to θ e1 = a⊤ x and λ e2 = x⊤ Bx. Using and then substituting in Eq. 29 the following: λ this, we can write the optimization problem as follows: a⊤ xm∗ − 12 x⊤ Bx(m2∗ + σ∗2 ) + 12 log(x⊤ Bx + 1) − 21

(a⊤ x)2 . x⊤ Bx + 1

(35)

Taking the derivatives, we get the optimality condition that is identical to Eq. 30, x⊤ a∗ = m∗ /σ∗2 = x⊤ y,

x⊤ B∗ x = 1/σ∗2 − 1 = x⊤ x.

(36)

These are satisfied by a∗ = y and B∗ = I, so a solution recovers the likelihood,

 t∗ (f ) = exp y⊤ xθ − 21 x⊤ xθ2 ∝ N (y | xθ, I) = p(y | θ).

(37)

This dual formulation yields the likelihood in p(y | θ) space, but since the dual solution is not unique, there is no bijection between posterior and likelihood. The formulation automatically performs the interchanging operation shown in Eq. 18 similarly to Amari’s case.

11

Train

Feedforward Concept Machine

θ

y

Lower System

LLM

θ

y

Feedback

Trace

(a)

(b)

Data

Fig. 2: Amari’s motivation was to explain the dynamic interaction between a higherorder system (concept machine) and a lower order system, through an iterative feedforward and feedback mechanism. In modern times, Bayesian duality could be useful to study the relationship between the model and data. For example, the dual problem can be used to trace the knowledge of an LLM.

4 Relevance for AI and Connections to Other Fields So far, we have shown that Amari’s Bayesian duality can be generalized through convex duality where the dual problem can be seen as an inverse (and sometimes bijective) mapping from the posterior manifold to the likelihood manifold. But, why is this relevant for machine learning and AI? Amari’s original proposal was focused on information processing in the brain. At the time, he aimed to develop a new theory of “dynamic interaction between a lower and a higher neural systems” (Fig. 2a). The lower order system, modeled by y, is connected to a sensory input while the higher order system is modeled by θ which is connected to a ‘concept’ machine. He imagined Bayesian duality as a stochastic version of the autoassociative-memory model where the concepts are encoded as binary vectors in the higher system. When a new sensory input is shown, the concept machine uses a ‘feedforward’ process to estimate a concept that best explains it. This is the inference part to get p(θ|y). Then, the higher neural system stimulates the lower one by proposing a new y through the feedback process. This is realized by p(y|θ). This dynamic interaction was the motivation for Amari to develop Bayesian duality, which he proposed to implement through an e-m procedure. For us too, the interaction between y and θ is the most attractive part of Bayesian duality. In our research, we are interested in Bayesian duality as a tool to study the relationship between the model and data. For example, given a Large Language Model (LLM) trained on a large data, we would like to understand the model’s knowledge. This question can be framed as the inverse mapping: we want to find a compact set of data examples (or prompts) that can summarize the main crux of the LLM’s knowledge. The dual problem can be used to address such problems: given p(θ|y), the mapping t(θ) in the likelihood space can be used to ‘trace’ a compact summary of the model’s knowledge (Fig. 2b). Amari himself was interested in understanding the underlying concepts encoded in θ, which he modeled through a subspace in the manifold spanned by a smaller set of coordinates. Our work here connects Amari’s ideas to other works that rely on convex duality to find compact representations. One of the earliest works on this is the representer theorem proposed by Kimeldorf and Wahba (1970) in a Bayesian context. Later, this

12

led to the representer theorem used in kernel methods (Schölkopf et al., 2001) and the support vector machine (SVM) where compact representations in the data space are found by solving an efficient convex-optimization problem. A representer theorem for Gaussian process models was proposed in Csató and Opper (2002, Lemma 1). General versions of such dualities have also been studied extensively, for example, by Altun and Smola (2006) for function spaces and by Dudı́k et al. (2007) to generalize the maximum entropy principle. An SVM-style generalization of Bayesian inference is proposed by Zhu et al. (2014). We also note that an informal mention of Bayesian duality is made in Robert (2007) for Bayes’ rule as an inversion but Amari’s usage is more precise. More recently, such dualities are used to obtain primal-dual adversarial interpretations of Bayes (Husain and Knoblauch, 2022; Möllenhoff and Khan, 2023). In fact, Bayesian duality can recover some non-Bayesian applications of convex duality in machine learning. For instance, the dual problem for ridge regression of Eq. 11 is a classical reformulation of minimization over θ as a maximization over an N -length real-valued dual vector α (Rockafellar, 1967). This is shown below, min − log p(y|θ) − log p(θ) θ

⇐⇒

max y⊤ α − 21 α⊤ (I + xx⊤ )α. α

(38)

We can show that, for this case, the optimal dual vector α∗ is related to t∗ (f ∗ ) in Eq. 33 as follows, α∗ = y − f ∗ = ∇ log t∗ (f ∗ ), (39) where f ∗ = xθ∗ is the optimal prediction. A proof is given in Sec. A.2. In general, such equivalences between the non-Bayesian and Bayesian dualities are to be expected due to their common origins in convex duality. We also expect Bayesian duality with more expressive posterior forms to contain as special cases other dualities based on less flexible posteriors and also those that arise in non-Bayesian scenarios. Recent works have tried to build such versions of Bayesian dualities by using variational approximations. Such methods use dual structures associated with simple posterior approximations (such as Gaussian approximations to the posterior). This makes it possible to derive practical algorithms that can exploit the benefits of Bayesian duality. One such early work is by Khan et al. (2013) which proposed a dual formulation for Gaussian variational inference. More recently, Adam et al. (2021) proposed the use of such dual variables for approximate inference in Gaussian processes. Khan (2025) uses dual representations to unify knowledge-adaptation methods, while Möllenhoff et al. (2026) used a dual structure to generalize ADMM for federated learning. The work started by Amari on Bayesian duality is now finally having an impact on modern AI and we hope that it will continue to benefit the community. Acknowledgments. This work is supported by JST CREST Grant Number JPMJCR2112. We thank many current and past members of the ABI team and the Bayes-duality project for discussions over the years which played an important role in shaping the ideas presented in this paper.

13

A Appendix A.1 Derivation of the Dual Problem for Ridge Regression To the integral in Eq. 29, we use a change of variables and the fact that √ R evaluate exp(−z 2 /2)dz = 2π. First, we bring the integral into a form amenable to the change of variables. Z Z     1 1e 2 e e1 θ − 1 (λ e2 + 1)θ2 dθ exp λ1 θ − 2 λ2 θ p(θ)dθ = √ exp λ (40) 2 2π   !Z !2 e1 e2 1 λ λ 1 1 e2 + 1) θ −  dθ. = √ exp exp − 2 (λ (41) e2 + 1) e2 + 1 2π 2(λ λ Then, setting z =

q   e2 + 1 θ − λe1 , dθ = √ dz λ e +1 λ

gives,

e2 +1 λ

2

Z

!2   2 Z z e2 + 1) θ −  dθ = q 1 exp − 12 (λ dz . exp − e2 + 1 2 λ e λ2 + 1 | {z } e1 λ

(42)

√ = 2π

Substituting that back into Eq. 41 and taking the log gives, log q

1

exp

e2 + 1 λ

e2 λ 1

! e2 + 1) + 1 = − 21 log(λ 2

e2 + 1) 2(λ

e2 λ 1 e2 + 1 λ

.

(43)

This is the result shown in the main paper. We now differentiate the full objective in e1 and λ e2 and set the resulting derivatives to zero. Eq. 28 with respect to λ m∗ −

e1 λ e2 + 1 λ

= 0,

−m2∗ − σ∗2 +

1

e2 + 1 λ

+

e2 λ 1 = 0. e 2 )2 (1 + λ

(44)

e1 = m∗ (λ e2 + 1). Inserting that into the equation on Solving the first equation yields λ the right, we arrive at σ∗2 =

1 e λ2 + 1

⇔

e2 = 1 − 1. λ σ∗2

(45)

e1 , we get λ e1 = m∗ /σ 2 . Since m∗ = σ 2 x⊤ y Plugging that back into the expression for λ ∗ ∗ ⊤ ⊤ 2 e1 = x y. Since σ = 1/(x x+1), we get λ e2 = x⊤ x as claimed. as per Eq. 12, we have λ ∗

14

A.2 Connection to Convex Dual of Ridge Regression The primal solution is just the ridge solution which has the following form: θ∗ = (x⊤ x + 1)−1 x⊤ y

(46)

The dual solution is obtained from the stationarity condition of the dual problem: y − (I + xx⊤ )α∗ = 0

=⇒

α∗ = (xx⊤ + I)−1 y.

(47)

From here, we can represent θ∗ in terms of α∗ by using the matrix inversion lemma, θ∗ = (x⊤ x + 1)−1 x⊤ y = x⊤ (xx⊤ + I)−1 y = x⊤ α∗

(48)

Using this, we can prove Eq. 39 from the stationarity condition of the dual problem: y − (I + xx⊤ )α∗ = 0

=⇒

y − α∗ − xθ∗ = 0

=⇒

α∗ = y − f ∗ .

(49)

This is equivalent to 1 ⊤ ∇ log t∗ (f ∗ ) = ∇(a⊤ ∗ f ∗ − 2 f ∗ B∗ f ∗ ) = a∗ − B∗ f ∗ = y − f ∗ .

(50)

Conflict of Interest: One of the authors has co-authored several papers with Dr. Frank Nielsen. There are no other obvious conflicts of interest to report at this moment.

References Amari, S.: A theory of adaptive pattern classifiers. IEEE Transactions on Electronic Computers (3), 299–307 (1967) Amari, S.: Theory of self-organizing nerve nets with special reference to association and concept formation. In: Proceedings of the 6th International Joint Conference on Artificial Intelligence (1979) Amari, S.: Information geometry of the EM and em algorithms for neural networks. Neural Networks 8(9), 1379–1408 (1995) Amari, S.: The EM algorithm and information geometry in neural network learning. Neural Comput. 7(1), 13–18 (1995) Amari, S.: Natural gradient works efficiently in learning. Neural computation 10(2), 251–276 (1998) Amari, S.: Information geometry of neural networks –new Bayesian duality theory–. In: International Conference on Neural Information Processing (ICONIP) (1996)

15

Wainwright, M.J., Jordan, M.I.: Graphical models, exponential families, and variational inference. Foundations and Trends in Machine Learning 1–2, 1–305 (2008) Barndorff-Nielsen, O.: Information and Exponential Families: In Statistical Theory. Wiley, Chichester, England (1978) Khan, M.E., Rue, H.: The Bayesian learning rule. Journal of Machine Learning Research 24(281), 1–46 (2023) Gibbs, J.W.: Elementary Principles in Statistical Mechanics. C. Scribner’s Sons, New York (1902) Jaynes, E.T.: Information Theory and Statistical Mechanics. Physical Review 106(4), 620–630 (1957) Donsker, M.D., Varadhan, S.S.: Asymptotic evaluation of certain Markov process expectations for large time, I. Communications on Pure and Applied Mathematics 28(1), 1–47 (1975) Donsker, M.D., Varadhan, S.R.S.: Asymptotic evaluation of certain Markov process expectations for large time, II. Communications on Pure and Applied Mathematics 28(2), 279–301 (1975) Donsker, M.D., Varadhan, S.R.S.: Asymptotic evaluation of certain Markov process expectations for large time, III. Communications on Pure and Applied Mathematics 29(4), 389–461 (1976) Donsker, M.D., Varadhan, S.R.S.: Asymptotic evaluation of certain Markov process expectations for large time. IV. Communications on Pure and Applied Mathematics 36(2), 183–212 (1983) Polyanskiy, Y., Wu, Y.: Information Theory: From Coding to Learning. Cambridge University Press, Cambridge, United Kingdom (2025) Kimeldorf, G.S., Wahba, G.: A correspondence between Bayesian estimation on stochastic processes and smoothing by splines. Annals of Mathematical Statistics 41(2), 495–502 (1970) Schölkopf, B., Herbrich, R., Smola, A.J.: A generalized representer theorem. In: Proceedings of the Workshop on Computational Learning Theory, vol. 2111, pp. 416–426 (2001) Csató, L., Opper, M.: Sparse on-line Gaussian processes. Neural computation 14(3), 641–668 (2002) Altun, Y., Smola, A.: Unifying divergence minimization and statistical inference via convex duality. In: International Conference on Computational Learning Theory (2006) 16

Dudı́k, M., Phillips, S.J., Schapire, R.E.: Maximum entropy density estimation with generalized regularization and an application to species distribution modeling. Journal of Machine Learning Research 8(44), 1217–1260 (2007) Zhu, J., Chen, N., Xing, E.P.: Bayesian inference with posterior regularization and applications to infinite latent SVMs. Journal of Machine Learning Research 15(1), 1799–1847 (2014) Robert, C.P.: The Bayesian Choice: from Decision-theoretic Foundations to Computational Implementation. Springer, New York (2007) Husain, H., Knoblauch, J.: Adversarial interpretation of Bayesian inference. In: International Conference on Artificial Intelligence and Statistics (2022) Möllenhoff, T., Khan, M.E.: SAM as an optimal relaxation of Bayes. In: The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023 (2023). https://openreview.net/forum?id=k4fevFqSQcX Rockafellar, R.: Duality and stability in extremum problems involving convex functions. Pacific Journal of Mathematics 21(1), 167–187 (1967) Khan, M.E., Aravkin, A.Y., Friedlander, M.P., Seeger, M.: Fast dual variational inference for non-conjugate latent Gaussian models. In: International Conference on Machine Learning (2013) Adam, V., Chang, P., Khan, M.E., Solin, A.: Dual parameterization of sparse variational Gaussian processes. In: Advances in Neural Information Processing Systems (2021) Khan, M.E.: Knowledge adaptation as posterior correction. arXiv:2506.14262 (2025) Möllenhoff, T., Swaroop, S., Doshi-Velez, F., Khan, M.E.: Federated ADMM from Bayesian duality. In: International Conference on Learning Representations (2026)

17

Record · ID 668025 · SHA-256 d3d100d3eea5467b
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.