NNT : 2026IPPAG003
On-Policy and Off-Policy Learning for Large Action Spaces Thèse de doctorat de l’Institut Polytechnique de Paris préparée à l’École nationale de la statistique et de l’administration économique École doctorale n◦ 574 École Doctorale de Mathématique Hadamard (EDMH) Spécialité de doctorat : Mathématiques appliquées
Thèse présentée et soutenue à Palaiseau, le 13 mars 2026, par
I MAD AOUALI Composition du Jury :
574
Vianney Perchet Professeur, CREST, ENSAE, IP Paris
Président
Olivier Cappé Directeur de recherche, CNRS
Rapporteur
Aurélien Garivier Professeur, Ecole Normale Superieure de Lyon
Rapporteur
Claire Vernade Professeure, University of Technology Nuremberg
Examinatrice
Victor-Emmanuel Brunel Professeur, CREST, ENSAE, IP Paris
Directeur de thèse
Anna Korba Professeure assistante, CREST, ENSAE, IP Paris
Invitée
David Rohde Chercheur, Criteo AI Lab
Invité
Acknowledgements
First and foremost, I wish to express my sincere and profound gratitude to my PhD supervisors, Victor-Emmanuel Brunel, Anna Korba, and David Rohde. It has been an immense privilege to learn from and work with them over these years. They shaped my research and personal growth in ways that will stay with me far beyond this thesis. Victor brought rigor, kindness, and openness to every discussion, profoundly shaping how I approach research and problem-solving. Anna’s brilliance, drive, and compassionate mentorship were central to the success of this PhD; her rare ability to combine deep technical insight with empathy and encouragement made her guidance invaluable. David was the best manager I could have hoped for, whose trust, patience, and human approach made all the difference when navigating both professional and personal challenges. I am profoundly grateful to the members of my thesis jury for the time, care, and expertise they devoted to evaluating this work. I would first like to express my deepest appreciation to Olivier Cappé and Aurélien Garivier for accepting the demanding role of rapporteurs, and for the considerable time and attention they dedicated to reading this manuscript in depth. I am sincerely grateful for their careful assessment, thoughtful comments, and constructive feedback. I would also like to warmly thank Vianney Perchet for the honor of presiding over the jury, and for his invaluable support. Finally, I am deeply grateful to Claire Vernade for serving as examinatrice, and for her generosity, encouragement, and support. It was a true privilege to have such distinguished researchers on my jury, and I deeply appreciate their scientific perspective, insightful remarks, and kindness. I would also like to thank Criteo, CAIL, and the Performance Science team, as well as CREST, ENSAE, Institut Polytechnique de Paris. Both the company and the laboratory provided outstanding scientific and institutional support throughout this thesis. I am deeply grateful to several mentors: to Branislav Kveton, whose insights as my first major co-author shaped my approach to research and publication; to Florian Strub, my DeepMind scholarship mentor, whose guidance inspired me to pursue a PhD; and to Flavian Vasile, and Michal Valko, for their invaluable support, both seen and unseen. My heartfelt thanks also go to all my co-authors, whose contributions have greatly enriched this work: A. Gilotte, B. Heymann, N. Nguyen, A. György, P. Alquier, N. Chopin, S. Katariya, AAS. Hammou, S. Ivanov, A. Benhalloum, M. Bompaire, M. Vono, M. Gartrell, V. Zaytsev, D. Legrand, and O. Jeunen. Special recognition goes to O. Sakhi, 2
M. Cherifa and A.B. Yahmed who were exceptional companions throughout this journey. Finally, to my family and friends: my deepest thanks to my mother and siblings for their unwavering love and support. This thesis is dedicated to my late father, whose dream was for me to pursue a PhD. To Basma, Hakim, Achraf, Ayman, Ismail, Tayeb, Anas, Nicolas, Youssef, Kini, Song, Issam, Charif, Yassine, Oussama, Abdellah, Ali, Hicham: thank you for your friendship and presence throughout this journey.
3
Abstract
Many interactive systems (e.g., recommender systems) can be modeled as contextual bandits. This framework captures the core challenge of decision-making under uncertainty: selecting actions based on context while learning from partial, noisy feedback. Learning in this setting follows two paradigms: on-policy learning, in which agents collect data and update their policy simultaneously in real time, and off-policy learning, in which the agent’s policy is learned offline from static logs collected under a different policy. Standard algorithms for both paradigms struggle to scale to large action spaces, facing either computational intractability or statistical inefficiency. This thesis develops principled and practical methods to make contextual bandit algorithms tractable in large action spaces, advancing both paradigms through novel algorithmic and theoretical contributions. For on-policy learning, we introduce structured Bayesian models that enable efficient exploration via information sharing. Our first contribution, mixed-effect Thompson sampling (meTS) (Chapter 3), couples √ action parameters through shared latent effects. This reduces Bayesian regret to Õ( T dKeff ), where Keff is an effective number of actions. When the number of shared effects is much smaller than the number of actions (L ≪ K), we have Keff ≪ K, yielding significant regret reduction. Moreover, meTS achieves dramatic improvements in both memory complexity (from O(K 2 d2 ) to O((L2 + K)d2 )) and runtime complexity (from O(K 3 d3 ) to O((L3 + K)d3 )). We extend this framework to diffusion Thompson sampling (dTS) (Chapter 4), which leverages deep generative models to capture complex action distributions. dTS further improves memory complexity to O((L + K)d2 ) and runtime complexity to O((L + K)d3 ), where L denotes the number of layers in the diffusion model. Both methods perform well empirically as analyzed without additional hyperparameter tuning, making them highly practical. For off-policy learning, we address fundamental bottlenecks through three complementary approaches. First, in Chapter 6, we develop the structured direct method (sDM), which models √ action parameters using a shared latent structure. We prove that sDM achieves O(1/ n) convergence in Bayesian suboptimality without requiring the restrictive full logging support assumption. sDM performs well in practice, and the performance gap between sDM and standard direct methods widens as the action space grows. Second, in Chapter 7, we challenge the conventional wisdom that better reward estimation yields better policies. We demonstrate that optimization intractability, rather than estimation accuracy, becomes the primary bottleneck in large action spaces, and advocate for policy-weighted 4
log-likelihood objectives that prioritize optimization tractability; these consistently outperform sophisticated estimators on datasets with up to one million actions. Third, in Chapter 8, we address importance sampling variance by combining variance-reducing regularization with principled pessimism. Our exponential smoothing estimators and unified PAC-Bayesian analysis yield tractable learning objectives amenable to stochastic optimization, providing concentration bounds and superior empirical performance. We validate these theoretical and algorithmic advances through extensive experiments on synthetic and real-world datasets. By developing scalable algorithms for both learning paradigms, this thesis enables the deployment of contextual bandits in modern applications where action spaces routinely exceed thousands or millions of actions.
5
Contents Résumé substantiel en français
12
1 Overview 1.1 Context and Scope . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1.2 Background . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1.3 Contributions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1.4 Related Work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
18 18 21 25 32
I
38
On-Policy Learning in Large Action Spaces
2 Introduction to Part I 2.1 Setting and Background . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2.2 Hierarchical Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2.3 Roadmap of Part I . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
39 39 40 40
3 Scaling Thompson Sampling with Mixed Effects 3.1 Setting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.2 Algorithm . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.3 Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.4 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.5 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
42 44 46 50 52 55
4 Scaling Thompson Sampling with Diffusion Models 4.1 Setting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4.2 Algorithm . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4.3 Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4.4 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4.5 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
57 58 59 63 66 69
II
70
Off-Policy Learning in Large Action Spaces
5 Introduction to Part II 5.1 Setting and Background . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5.2 Methodological Approaches . . . . . . . . . . . . . . . . . . . . . . . . . . 5.3 Roadmap of Part II . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
71 71 72 73
6 Scaling Direct Methods with Latent Parameters 6.1 Setting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6.2 Structured DM . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6.3 Linear-Gaussian Case . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6.4 Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6.5 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6.6 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
75 76 76 79 80 83 85
7 Optimization Matters More than Esimation 7.1 Analysis of IPS-Based Objectives . . . . . . . . . . . . . . . . . . . . . . . 7.2 Analysis of PWLL objectives . . . . . . . . . . . . . . . . . . . . . . . . . . 7.3 Empirical Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7.4 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
86 87 92 95 98
8 Principled Pessimism for Exponential Smoothing and Beyond 99 8.1 Background . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 100 8.2 Exponential Smoothing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 102 8.3 PAC-Bayes Analysis for Off-Policy Learning . . . . . . . . . . . . . . . . . 104 8.4 Discussion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 109 8.5 Experiments for Exponential Smoothing . . . . . . . . . . . . . . . . . . . 112 8.6 Extension to Other Regularizations . . . . . . . . . . . . . . . . . . . . . . 115 8.7 Experiments for Other Regularizations . . . . . . . . . . . . . . . . . . . . 118 8.8 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 121 9 Conclusions and Future Work
123
A Supplementary Materials for Chapter 3 125 A.1 Preliminaries . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 125 A.2 Posterior Derivations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 125 A.3 Regret Proofs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 129 A.4 Additional Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . 139 B Supplementary Materials for Chapter 4 143 B.1 Posterior for Linear Diffusion Models . . . . . . . . . . . . . . . . . . . . . 143 B.2 Posterior for Non-Linear Diffusion Models . . . . . . . . . . . . . . . . . . 145 B.3 Connection to Two-Level Hierarchies . . . . . . . . . . . . . . . . . . . . . 146 B.4 Formal Theory . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 147 B.5 Regret proof . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 148 B.6 Additional Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . 157 C Supplementary Materials for Chapter 6 161 C.1 Posterior Derivations Under Standard Priors . . . . . . . . . . . . . . . . . 161 C.2 Posterior Derivations Under Structured Priors . . . . . . . . . . . . . . . . 162 C.3 Proofs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 167 C.4 Additional Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . 174 D Supplementary Materials for Chapter 7 178 D.1 Proofs for Oracle Policies . . . . . . . . . . . . . . . . . . . . . . . . . . . . 178 7
D.2 Proofs for Optimization Properties . . . . . . . . . . . . . . . . . . . . . . 182 D.3 Stochastic Optimization Convergence Guarantees for PWLL . . . . . . . . 188 D.4 Additional Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . 191 E Supplementary Materials for Chapter 8 202 E.1 Bias and Variance Trade-Off . . . . . . . . . . . . . . . . . . . . . . . . . . 203 E.2 Proofs for Off-Policy Learning . . . . . . . . . . . . . . . . . . . . . . . . . 204 E.3 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 217
8
Notation
Notation General Mathematical Notation Symbol
Definition
[n] Rd Id ∆(A) O(·) ⊗
Set of first n positive integers: {1, 2, . . . , n} d-dimensional real vector space Identity matrix of dimension d × d Probability simplex over action set A Big-O notation for upper bounds Kronecker product
Probability conventions. Random variables are denoted with capital letters, and their realizations with the respective lowercase letters, except for Greek letters. With a slight abuse of notation, for random variables X, Y , the distribution (or density) of X | Y = y evaluated at x is denoted by p(x | y).
Norms and Inner Products Symbol
Definition
∥·∥ ∥a∥Σ
Euclidean norm (unless specified otherwise) √ ⊤ Weighted norm: a Σa for a ∈ Rd , Σ ≻ 0
Contextual Bandit Framework Symbol
Definition
X ⊂ Rd A = [K] T n
Context space (d-dimensional) Finite action set with K actions Number of interaction rounds Number of samples in logged dataset 9
Policies and Value Functions Symbol
Definition
π : X → ∆(A) π(· | x) πt π0 π∗ π̂ V (π) V̂ (π)
Stochastic policy mapping contexts to action distributions Probability distribution over actions given context x Policy at round t (on-policy setting) Logging policy (off-policy setting) Optimal policy (off-policy setting) Learned policy (off-policy setting) Value (expected reward) of policy π Estimated value of policy π
Environment Components Symbol
Definition
ν p(· | x, a) r(x, a) r̂(x, a) r(x, a; θ) r(x, a; θ̂)
Context distribution over X Conditional reward distribution given context x and action a Expected reward function: ER∼p(·|x,a) [R] Estimated reward function Parametric reward model Estimated reward function in parametric case
Random Variables and Data Symbol
Definition
Xt At Rt At,∗ Ht Dn
Context observed at round t Action taken at round t Reward received at round t Optimal action at round t: arg maxa∈A r(Xt , a) History of interactions: {(Xℓ , Aℓ , Rℓ )}ℓ<t Logged dataset: {(Xi , Ai , Ri )}ni=1
Performance Metrics Symbol Definition ∑︁T R(T ) = t=1 r(Xt , At,∗ ) − r(Xt , At ) Cumulative regret BR(T ) = E[R(T )] Bayesian cumulative regret so(π̂) = V (π∗ ) − V (π̂) Suboptimality gap of policy π̂ Bso(π̂) = E[so(π̂)] Bayesian suboptimality gap of policy π̂ 10
Matrix Operations and Concatenation Symbol
Definition
[a1 , a2 , . . . , an ] (ai )i∈[n] Vec(·) diag((Ai )i∈[n] ) (Ai )i∈[n] (Ai,j )(i,j)∈[n]×[m]
Horizontal concatenation of vectors into d × n matrix ⊤ ⊤ nd Vertical concatenation: (a⊤ 1 , . . . , an ) ∈ R Vectorization operator Block diagonal matrix with blocks A1 , . . . , An Vertical concatenation of matrices into nd × d matrix Block matrix where Ai,j is the (i, j)-th block
Eigenvalues Symbol
Definition
λ1 (A) λd (A)
Maximum eigenvalue of matrix A Minimum eigenvalue of matrix A
11
Résumé substantiel en français
Cette thèse étudie l’apprentissage séquentiel et contrefactuel dans des systèmes interactifs où l’espace des décisions est très grand. Ces systèmes sont aujourd’hui omniprésents : moteurs de recommandation, publicité computationnelle, places de marché, systèmes de tarification, robotique ou encore allocation de ressources. Leur fonctionnement repose sur une boucle d’interaction simple mais difficile à optimiser : à chaque étape, le système observe un contexte, choisit une action parmi un très grand nombre de possibilités, puis reçoit une récompense partielle et bruitée qui dépend conjointement du contexte et de l’action choisie. Dans un système de recommandation, le contexte peut représenter l’historique et les préférences d’un utilisateur, l’action correspond à l’item recommandé dans un catalogue contenant potentiellement des millions d’items, et la récompense mesure l’engagement de l’utilisateur, par exemple un clic ou un temps de visionnage. En publicité en ligne, l’action peut combiner le choix d’une annonce et d’un prix d’enchère, tandis que la récompense dépend d’événements successifs tels que le gain de l’enchère, le clic ou la conversion. Le cadre mathématique central de cette thèse est celui des bandits contextuels. On considère un espace de contextes X ⊂ Rd , un ensemble fini d’actions A = [K], une distribution inconnue de contextes ν, et des distributions conditionnelles de récompenses p(· | x, a). La fonction de récompense moyenne est définie par r(x, a) = E[R | X = x, A = a]. À chaque tour t, un contexte Xt ∼ ν est observé, l’agent choisit une action At selon une politique πt (· | Xt ), puis reçoit une récompense Rt ∼ p(· | Xt , At ). Ce formalisme capture l’essentiel de l’apprentissage interactif : l’agent ne voit que la récompense de l’action qu’il a effectivement choisie, et doit donc apprendre à partir d’un retour partiel. La difficulté principale analysée dans cette thèse est le passage à l’échelle lorsque K est très grand. La thèse traite deux paradigmes complémentaires. Le premier est l’apprentissage en ligne, ou on-policy learning, dans lequel l’agent interagit séquentiellement avec l’environnement et met à jour sa politique au fil des observations. Sa performance est mesurée par le regret cumulé, T ∑︂ R(T ) = (r(Xt , At,∗ ) − r(Xt , At )) , t=1
12
où At,∗ désigne l’action optimale dans le contexte Xt . L’enjeu fondamental est le compromis exploration-exploitation : l’agent doit explorer les actions incertaines pour apprendre, tout en exploitant les actions déjà estimées comme performantes afin de limiter la perte de récompense. Le second paradigme est l’apprentissage hors politique, ou off-policy learning, dans lequel l’agent ne peut plus interagir avec l’environnement et doit apprendre à partir d’un jeu de données journalisé Dn = {(Xi , Ai , Ri )}ni=1 collecté par une politique de logging π0 . L’objectif est alors d’apprendre une politique π̂ de grande valeur V (π) = EX∼ν EA∼π(·|X) [r(X, A)], et la performance est mesurée par l’écart de sous-optimalité V (π∗ )−V (π̂). Dans ce second cadre, l’apprentissage exige un raisonnement contrefactuel : il faut estimer ce qui se serait passé si une autre action avait été choisie, alors même que les données ne contiennent que les actions effectivement sélectionnées par π0 . Dans les deux paradigmes, la grande taille de l’espace d’actions amplifie les difficultés statistiques, computationnelles et d’optimisation. En ligne, explorer indépendamment des milliers ou millions d’actions devient prohibitif : chaque action reçoit peu d’observations, ce qui ralentit considérablement l’apprentissage et augmente le regret. Hors politique, la couverture des données se dégrade lorsque K croît : de nombreuses actions sont rarement ou jamais observées dans certaines régions de l’espace des contextes. Les méthodes fondées sur un modèle de récompense souffrent alors d’un fort biais d’extrapolation, tandis que les méthodes par pondération inverse des propensions peuvent avoir une variance très élevée lorsque π0 (a | x) est faible. Cette thèse montre également qu’un autre obstacle, souvent sous-estimé, devient dominant dans les grands espaces d’actions : l’optimisation des objectifs hors politique standards peut devenir intrinsèquement difficile, indépendamment de la qualité statistique de l’estimateur de valeur. La première partie de la thèse est consacrée à l’apprentissage en ligne dans les grands espaces d’actions. Les méthodes classiques telles que Upper Confidence Bound et Thompson Sampling reposent souvent sur des modèles disjoints, où chaque action a possède son propre paramètre θa et où la récompense est modélisée sous la forme r(x, a) = ϕ(x)⊤ θa . Ces modèles sont attractifs en pratique, notamment dans les systèmes de recommandation, car ils évitent de construire manuellement des caractéristiques conjointes contexteaction. Cependant, leur faiblesse est statistique : apprendre séparément un paramètre pour chaque action nécessite beaucoup de données par action, ce qui est incompatible avec de très grands catalogues. La contribution principale de cette partie consiste à conserver la flexibilité des modèles disjoints tout en introduisant des structures bayésiennes capables de partager l’information entre actions. Le premier algorithme proposé est mixed-effect Thompson Sampling, ou meTS. Il repose sur un modèle bayésien hiérarchique où les paramètres d’actions θa sont couplés par des effets latents partagés Ψ = (ψℓ )ℓ∈[L] . Ces effets peuvent représenter, par exemple, des catégories ou des facteurs communs entre items. Le modèle suppose que les paramètres d’actions sont conditionnellement indépendants sachant les effets latents, mais qu’ils partagent de l’information à travers ces effets. À chaque tour, meTS échantillonne d’abord les effets 13
latents depuis leur postérieur, puis échantillonne les paramètres d’actions conditionnellement à ces effets, avant de choisir l’action maximisant la récompense échantillonnée. Cette procédure conserve le principe de Thompson Sampling tout en rendant l’exploration statistiquement plus efficace. Dans le cas linéaire-gaussien, la thèse dérive des mises à jour exactes en forme fermée, et propose des approximations de Laplace tractables pour les modèles linéaires généralisés. L’analyse théorique établit une borne de regret bayésien de l’ordre )︃ (︃√︂ 2 2 ˜︁ T dKeff (σ0 + σΨ ) , O où Keff est un nombre effectif d’actions. Lorsque la structure latente est informative et que L ≪ K, on a Keff ≪ K, ce qui conduit à une amélioration multiplicative par rapport à Thompson Sampling standard. Sur le plan computationnel, la factorisation conditionnelle réduit fortement les coûts mémoire et temps par rapport à une modélisation bayésienne dense de toutes les actions. Les expériences montrent que les gains de meTS augmentent avec la taille de l’espace d’actions, confirmant que le partage d’information est essentiel pour l’exploration à grande échelle. La seconde contribution en ligne est diffusion Thompson Sampling, ou dTS. Cette méthode généralise meTS en remplaçant la hiérarchie à un niveau par une hiérarchie profonde inspirée des modèles de diffusion. Les paramètres d’actions sont générés au terme d’une chaîne de variables latentes reliées par des transformations non linéaires pré-entraînées. Cette structure permet de représenter des dépendances complexes entre actions, bien audelà des effets linéaires ou catégoriels. Le défi technique est que le postérieur exact devient intraitable à cause des non-linéarités des fonctions de lien et du modèle de récompense. La thèse introduit donc une procédure d’inférence en ligne fondée sur des mises à jour de type gaussien, où les précisions postérieures combinent précision a priori et précision issue des données, et où les moyennes sont obtenues par combinaison pondérée entre les prédictions du prior et les estimations de maximum de vraisemblance. L’intérêt de dTS est double. D’une part, l’algorithme exploite des priors riches appris hors ligne, par exemple à partir de représentations d’items, pour accélérer l’exploration en ligne. D’autre part, il conserve une structure de diffusion dans le postérieur, plutôt que de l’approximer par une simple gaussienne globale. Dans le cadre linéaire-gaussien, la thèse établit une borne de regret bayésien de l’ordre ⎞ ⎛⌜ ⃓ L+1 ⃓ ∑︂ ˜︁⎝⎷T dKeff σℓ2 ⎠ , O ℓ=1
et montre que la complexité peut être rendue linéaire en L + K. Empiriquement, dTS améliore les performances des méthodes de référence, y compris lorsque les priors de diffusion sont imparfaits ou appris à partir de données limitées. Cette contribution montre que les modèles génératifs profonds peuvent être utilisés non seulement pour représenter des actions, mais aussi pour structurer l’incertitude nécessaire à l’exploration. La seconde partie de la thèse porte sur l’apprentissage hors politique dans les grands espaces d’actions. Les méthodes classiques se divisent principalement en méthodes directes, 14
qui apprennent un modèle de récompense r̂(x, a), et méthodes par importance sampling, qui estiment directement la valeur d’une politique au moyen du ratio π(a | x)/π0 (a | x). Les méthodes directes sont sensibles au biais de modèle et à la rareté des observations par action. Les méthodes IPS sont non biaisées sous des hypothèses de support appropriées, mais leur variance peut exploser lorsque la politique cible attribue de la masse à des actions peu probables sous la politique de logging. En outre, la thèse montre que les objectifs IPS induisent souvent des paysages non concaves, plats et riches en maxima locaux lorsque l’espace d’actions est grand. La première contribution hors politique est la structured Direct Method, ou sDM. Elle transpose au cadre offline l’idée de structuration bayésienne introduite dans meTS. Au lieu d’estimer indépendamment un paramètre par action, sDM couple les paramètres θa à travers un vecteur latent partagé ψ. Après observation du jeu de données journalisé, l’algorithme calcule le postérieur de ψ, puis les postérieurs conditionnels des paramètres d’actions. La récompense estimée pour chaque action est obtenue en intégrant l’incertitude postérieure, et la politique apprise agit ensuite gloutonnement par rapport à cette récompense moyenne postérieure. L’analyse introduit une notion de sous-optimalité bayésienne √ adaptée au cadre hors politique. Elle montre que sDM atteint une convergence en O(1/ n) sans imposer l’hypothèse de support uniforme complet π0 (a | x) ≥ γ > 0 pour toutes les actions. La borne dépend plutôt de l’alignement entre la politique de logging et la politique optimale : plus les actions optimales sont couvertes par les données, plus l’apprentissage est efficace. Cette analyse révèle également un phénomène important : sous le critère bayésien considéré, les politiques gloutonnes sont optimales et peuvent surpasser les politiques pessimistes, contrairement au cadre fréquentiste où le pessimisme est souvent nécessaire pour se protéger contre les pires cas. Les expériences confirment que sDM améliore les méthodes directes standards, avec des gains croissants lorsque K augmente. La contribution suivante remet en question une hypothèse centrale de l’apprentissage hors politique : l’idée selon laquelle l’amélioration des estimateurs de valeur suffit à améliorer l’apprentissage de politiques. La thèse montre que, dans les grands espaces d’actions, l’erreur d’optimisation peut dominer l’erreur d’estimation. Même un estimateur statistiquement sophistiqué peut conduire à une mauvaise politique si l’objectif qu’il induit est difficile à optimiser. L’analyse des paysages d’optimisation montre que les objectifs fondés sur des estimateurs peuvent présenter des plateaux où les méthodes de gradient restent bloquées pendant O(K) itérations, ainsi qu’un nombre exponentiel de maxima locaux en fonction de K. Pour comprendre ces échecs, la thèse analyse les politiques oracle associées à différents estimateurs, c’est-à-dire les politiques qui maximiseraient ces estimateurs avec une quantité infinie de données. Cette analyse met en évidence que chaque estimateur impose un biais inductif spécifique : IPS favorise les politiques proches du support de logging, tandis que des méthodes clusterisées opèrent au niveau de groupes d’actions. Ces observations motivent des paramétrisations de politiques adaptées à l’objectif, qui réduisent l’espace de recherche effectif de K à une taille beaucoup plus petite, comme la taille du support de logging ou le nombre de clusters. La conclusion principale de cette partie est toutefois plus radicale : il peut être préférable 15
d’abandonner l’estimation explicite de valeur pour optimiser directement des objectifs de vraisemblance pondérée par la politique. La thèse introduit les objectifs policy-weighted log-likelihood, de la forme n
1 ∑︂ g(Ri , π0 (Ai | Xi )) log π(Ai | Xi ), Ûg (π) = n i=1 où g est une fonction de pondération positive. Ces objectifs ne sont pas des estimateurs de valeur, mais ils possèdent des paysages d’optimisation bien plus favorables : pour des politiques softmax linéaires, ils sont concaves, et deviennent fortement concaves avec régularisation ℓ2 . Ils admettent donc un optimum global unique accessible par optimisation stochastique standard. Les expériences à très grande échelle, incluant des espaces allant jusqu’à un million d’actions, montrent que ces objectifs simples et stables surpassent des méthodes hors politique fondées sur des estimateurs de valeur plus complexes. Cette contribution établit que, pour l’apprentissage de politiques dans les grands espaces d’actions, l’optimisabilité de l’objectif est un critère aussi fondamental que sa précision statistique. La dernière contribution principale de la thèse améliore les méthodes IPS régularisées en introduisant un pessimisme praticable et différentiable. Les ratios d’importance peuvent être très grands, ce qui augmente fortement la variance. Pour y remédier, la thèse étudie des estimateurs par exponential smoothing, notamment n
1 ∑︂ π(Ai | Xi ) Ri , V̂ (π) = n i=1 π0 (Ai | Xi )α α
et
n
1 ∑︂ Ṽ (π) = n i=1 β
(︃
π(Ai | Xi ) π0 (Ai | Xi )
)︃β Ri .
Ces estimateurs interpolent entre absence de régularisation et forte réduction de variance, tout en restant différentiables et donc compatibles avec l’optimisation par gradient. Pour apprendre de manière sûre avec ces estimateurs biaisés mais moins variables, la thèse dérive une borne PAC-bayésienne bilatérale contrôlant l’écart entre la valeur vraie et l’estimateur régularisé. Cette borne décompose l’erreur en plusieurs termes interprétables : une divergence entre la politique apprise et la politique de logging, un biais dû à la régularisation des poids, et une variance résiduelle. Elle conduit à un objectif pessimiste qui maximise une borne inférieure empirique de la valeur : V̂nα (πθ ) − pénalités de divergence − biais de régularisation − variance résiduelle. L’intérêt crucial de cette formulation est que tous les termes sont empiriques et différentiables. Contrairement à de nombreuses approches pessimistes dont les constantes théoriques sont inexploitables en pratique, cet objectif peut être optimisé à grande échelle par ascension de gradient stochastique. La thèse propose également un cadre PACbayésien unifié couvrant plusieurs familles de régularisation des poids d’importance, comme le clipping, l’exponential smoothing et l’implicit exploration. Ce cadre permet une comparaison cohérente de différentes formes de pessimisme et fournit des objectifs pratiques pour l’apprentissage offline sécurisé. 16
Enfin, la thèse présente plusieurs contributions additionnelles liées aux thèmes principaux. Dans le cadre en ligne, les idées de structuration bayésienne sont étendues au problème de best-arm identification à budget fixé. L’algorithme PI-BAI utilise l’information a priori pour répartir efficacement le budget d’exploration dans des bandits structurés. L’analyse fournit des garanties bayésiennes dépendant du prior sur la probabilité d’erreur, et montre que des allocations non adaptatives bien informées peuvent surpasser des stratégies adaptatives classiques. Dans le cadre hors politique, la thèse contribue également au développement du logarithmic smoothing, un estimateur pessimiste de la forme λ V̂LS (π) =
(︃ )︃ n π(Ai | Xi ) 1 ∑︂ log 1 + λ Ri . nλ i=1 π0 (Ai | Xi )
Cet estimateur agit comme une alternative douce et différentiable au clipping, bénéficie de garanties de concentration serrées, et permet d’obtenir des bornes de sous-optimalité plus fines que les approches précédentes. Enfin, cette thèse CIFRE maintient un lien constant avec les systèmes de recommandation industriels à grande échelle. Les applications développées autour de la recommandation optimisant la récompense, de l’évaluation offline, de la simulation contrefactuelle et des modèles de recommandation par ardoise ont fourni à la fois un terrain expérimental et une source de questions théoriques. Dans son ensemble, cette thèse défend une idée centrale : pour apprendre efficacement dans de grands espaces d’actions, il ne suffit pas d’appliquer directement les algorithmes classiques de bandits contextuels ou d’apprentissage hors politique. Il faut exploiter la structure entre actions, contrôler explicitement l’incertitude et la couverture des données, et concevoir des objectifs dont le paysage d’optimisation reste favorable à grande échelle. Les contributions proposées répondent à ces exigences selon deux axes complémentaires. En apprentissage en ligne, des modèles bayésiens hiérarchiques et diffusionnels permettent de partager l’information entre actions et de réduire le regret. En apprentissage hors politique, des méthodes directes structurées, des objectifs de vraisemblance pondérée et des principes de pessimisme différentiable rendent l’apprentissage statistiquement robuste et computationnellement réalisable. Ces résultats contribuent à rapprocher la théorie des bandits contextuels des contraintes réelles des systèmes interactifs modernes, où les décisions doivent être prises parmi des catalogues massifs, à partir de signaux partiels, bruités et parfois fortement biaisés.
17
Chapter 1
Overview
1.1
Context and Scope
Interactive machine learning systems are a cornerstone of modern technology, optimizing decision-making in applications ranging from recommender systems and financial markets to robotics. These systems operate in a sequential loop: they process contextual information, select an action from a wide range of possibilities, and receive feedback that depends on both the context and the chosen action. A fundamental challenge in designing these systems is the scale of the decision space; in many real-world settings, the number of potential actions can be huge. Consider recommender systems, where streaming platforms or e-commerce sites must select an item to present to a user. The context comprises rich data, such as user preferences and history. The action is the selection of a specific item from a catalog containing thousands or millions of options. The system’s objective is to learn a policy that maximizes user engagement (the reward ), measured by metrics such as watch time or clicks. Similarly, in computational advertising, a platform selects which ad to display via real-time bidding. The context includes user attributes and page details. The action is composite: selecting an ad from a large inventory and determining a bid price. The feedback (reward) arrives in stages, from winning the auction to subsequent user clicks or conversions. The system must maximize advertiser value while adhering to budget constraints. We model these interactive systems using the contextual bandit framework. This framework captures the essential characteristics of interactive learning while maintaining the tractability required for theoretical analysis and practical implementation.
1.1.1
Contextual Bandits
Figure 1.1a visualizes the interaction loop. The environment consists of a Context Generator and a Reward Generator, both of which are assumed to be fixed but unknown to the agent. At each round, the environment emits a context. The agent observes this con18
Contextual bandit environment Context
Context Generator
Context
Contextual bandit environment at round
Reward Generator
Action
Reward
Agent
(b) Graphical representation in round t ∈ [T ].
(a) Single interaction loop.
Figure 1.1: Contextual bandit framework. text and selects an action from a finite set1 . Finally, the environment generates a scalar reward based on the context-action pair. By repeating this loop, the agent accumulates experience to refine its policy. Formally, let X ⊂ Rd denote the context space and A = [K] the finite action set. A stochastic policy π : X → ∆(A) maps each context x ∈ X to a probability distribution π(· | x) over actions. The environment is specified by: • A context distribution ν over X ; • A family of conditional reward distributions {p(· | x, a)}(x,a)∈X ×A . The expected reward function for any context-action pair (x, a) is defined as: r(x, a) = ER∼p(·|x,a) [R].
(1.1)
The interaction unfolds over T rounds. In each round t ∈ [T ]: 1. The environment draws a context Xt ∼ ν and reveals it to the agent. 2. The agent selects an action At ∼ πt (· | Xt ) according to its current policy πt . 3. The environment samples a reward Rt ∼ p(· | Xt , At ) and returns it to the agent. A graphical representation of this interaction in round t ∈ [T ] is visualized in Figure 1.1b.
1.1.2
Learning Paradigms
We address learning in this framework via two complementary paradigms: on-policy (online) and off-policy (offline). With a slight abuse of terminology, we use learning loosely to encompass the full agent behavior, including both reward estimation and action selection. On-Policy (Online) Learning In this setting, the agent updates its policy πt sequentially. Let Ht = {(Xℓ , Aℓ , Rℓ )}ℓ<t denote the history available at the start of round t. The agent uses Ht to construct 1
The action space can technically be infinite, but this thesis focuses on large finite action spaces.
19
the policy πt . Following the interaction, the agent augments the history with the new observation (Xt , At , Rt ) to form Ht+1 and repeats the process. Performance is measured by the cumulative regret: R(T ) =
T ∑︂
(r(Xt , At,∗ ) − r(Xt , At )) ,
(1.2)
t=1
where At,∗ = arg maxa∈A r(Xt , a) is the optimal action in round t. While minimizing regret is equivalent to maximizing cumulative reward, the literature prioritizes regret as it normalizes performance against the optimal oracle, facilitating theoretical comparisons across environments. The agent faces the exploration-exploitation dilemma: it must balance exploration of poorly understood actions with exploitation of actions believed to yield high rewards. Feedback is partial (only the reward for the chosen action is observed) and noisy, and exploration may be constrained by safety or budget requirements. Remark 1 (Beyond regret minimization). While this thesis focuses on regret minimization, other objectives exist, most notably Best-Arm Identification (BAI). BAI aims to identify the optimal action rather than maximize cumulative reward. This is important for applications like A/B testing and clinical trials. Although we focus on regret, our core modeling contributions for scaling Thompson sampling extend to the BAI setting, as demonstrated in our related work (Nguyen et al., 2025) (not included in this manuscript). Off-Policy (Offline) Learning In this setting, the agent learns a policy π̂ from a static logged dataset Dn = {(Xi , Ai , Ri )}ni=1 collected by a logging policy π0 as Xi ∼ ν, Ai ∼ π0 (· | Xi ) and Ri ∼ p(· | Xi , Ai ). No additional interactions with the environment are allowed. The objective is to find a policy maximizing the expected value: V (π) = EX∼ν EA∼π(·|X) [r(X, A)].
(1.3)
Performance is measured by the suboptimality gap of the learned policy π̂: so(π̂) = V (π∗ ) − V (π̂),
(1.4)
where π∗ = argmaxπ∈Π V (π) is the unknown optimal policy in a class of policies Π. Since the agent learns solely from data generated by π0 , it must perform counterfactual reasoning. This introduces several challenges. First, support mismatch occurs when specific actions are rarely or never selected by π0 in certain regions of the context space; consequently, the dataset provides little to no information about the rewards in these regions. Second, re-weighting instability arises since standard techniques, such as inverse propensity scoring (importance sampling), re-weight logged samples using the density ratio π(a | x)/π0 (a | x). When π0 (a | x) is small, this ratio 20
can explode, leading to high-variance estimates. Third, high bias can affect methods that rely on parametric models to estimate the reward. Extrapolating rewards for unobserved context-action pairs introduces errors if the model is misspecified. The difficulties in both paradigms are amplified when the action space is large (K in the thousands or millions). In on-policy settings, independent exploration of every action becomes infeasible, yielding high regret. In off-policy settings, data sparsity worsens: reward models risk significant extrapolation error, and importance weights suffer from extreme variance as the probability of observing any specific action vanishes. Developing methods that scale gracefully (both statistically and computationally) with the size of the action space is the central theme of this thesis.
1.2
Background
Most on-policy and off-policy learning algorithms2 are fundamentally constructed from two core components: reward estimation, which approximates the expected reward function r(x, a) for any context-action pair from data, and decision-making, which leverages these estimates and their associated uncertainty to select actions.
1.2.1
Reward Estimation
The central task in reward estimation is to learn an approximate function r̂(x, a) that estimates the true expected reward r(x, a) defined in Equation (1.1). This function is trained on a dataset, denoted by Data, whose structure depends on the learning paradigm: On-policy data. The agent collects data sequentially. The dataset at round t is the history of interaction up to round t: Data = Ht = {(Xi , Ai , Ri )}t−1 i=1 .
(1.5)
Off-policy data. The agent learns from a static dataset logged by a logging policy π0 : Data = Dn = {(Xi , Ai , Ri )}ni=1 ,
with Ai ∼ π0 (· | Xi ).
(1.6)
We denote the learned model as r̂(x, a) = r(x, a; θ̂), where θ̂ are parameters obtained via a statistical objective. Common approaches include: Maximum Likelihood Estimation (MLE) MLE seeks parameters θ that maximize the probability of observing the collected rewards. The objective is: ∑︂ θ̂MLE = arg max log p(Ri | Xi , Ai ; θ). θ
(Xi ,Ai ,Ri )∈Data
For Gaussian rewards p(· | x, a; θ) = N (r(x, a; θ), σ 2 ), this simplifies to minimizing the sum of squared errors (ordinary least squares). For Bernoulli rewards, it corresponds to minimizing binary cross-entropy (e.g., logistic regression). 2
Recall that we use learning loosely to encompass both reward estimation and action selection (Section 1.1.2).
21
Maximum A Posteriori (MAP) MAP estimation incorporates a prior distribution p0 (θ) to regularize the objective: ⎛ ⎞ ∑︂ θ̂MAP = arg max ⎝ log p(Ri | Xi , Ai ; θ) + log p0 (θ)⎠ . θ
(Xi ,Ai ,Ri )∈Data
For example, combining a Gaussian likelihood with a zero-mean Gaussian prior p0 (·) = N (0, λId ) is equivalent to ridge regression (ℓ2 regularization). Full Bayesian Inference Rather than a point estimate, Bayesian inference characterizes uncertainty by computing the full posterior distribution p(θ | Data): ∏︂
p(θ | Data) ∝ p(θ)
p(Ri | Xi , Ai ; θ).
(Xi ,Ai ,Ri )∈Data
This approach is powerful for guiding exploration. A common implementation assumes conjugate Gaussian distributions: p0 (·) = N (µ0 , Σ0 ) and p(· | x, a; θ) = N (r(x, a; θ), σ 2 ). When r(x, a; θ) is linear in θ (Bayesian linear regression), the posterior is Gaussian and analytically tractable. For non-linear models, posterior approximation methods are required. The functional form of the reward model, r(x, a; θ), is an important design choice that dictates the balance between computational tractability, data efficiency, and expressive power. While numerous function classes exist, linear models remain a cornerstone of the field due to their simplicity and strong theoretical guarantees. Even within this linear family, there is an important distinction between joint and disjoint formulations. The joint linear model defines r(x, a; θ) = ϕ(x, a)⊤ θ, sharing a single parameter θ ∈ Rd across all actions. While data-efficient, this approach relies on designing a feature map ϕ(x, a) capable of capturing complex context-action interactions. Designing ϕ(x, a) is very hard in practice, and inadequate feature engineering leads to poor performance. This is why in practice, it is often more common to adopt a disjoint linear model that learns an independent parameter θa ∈ Rd for each action, yielding r(x, a; θ) = ϕ(x)⊤ θa . This is common in recommender systems, where predictions are inner products of user and item embeddings. This formulation is robust and avoids complex feature engineering. However, its primary drawback is poor statistical scalability: while the computational overhead of maintaining K independent embeddings can be addressed with sufficient compute, the statistical challenge remains. This is because learning each embedding independently requires substantial data per action, which becomes prohibitive as K grows. Other non-linear models exist for capturing complex reward functions. Generalized linear models (GLMs), for instance, employ a link function to model non-Gaussian rewards (e.g., logistic regression for binary outcomes: p(R = 1 | x, a) = σ(ϕ(x)⊤ θa )). These models align closely with the linear setting, and the algorithms proposed in this thesis are suitable for them. A significant portion of this thesis focuses on scaling these disjoint reward models to large action spaces. 22
Remark 2 (Scope of parametric reward modeling). The reward estimation framework presented above assumes parametric reward models r(x, a; θ). This assumption underlies Chapters 3, 4 and 6, where we develop structured parametric models that share information across actions to improve statistical efficiency. The remaining two chapters (Chapters 7 and 8) take a slightly different approach: rather than explicitly estimating rewards, they optimize policies directly using inverse propensity scoring and policy-weighted objectives, thereby making them agnostic to the choice of reward model.
1.2.2
Decision-Making
We now turn to the second component: decision-making. This stage (often) relies on the reward model r̂ derived using the estimation techniques discussed previously. While the difference in reward estimation between on-policy and off-policy settings is primarily driven by how data is accumulated, the principles guiding action selection in these two paradigms are fundamentally distinct. Decision-Making in On-Policy Learning Recall that in the on-policy setting, the agent must balance exploration and exploitation to minimize regret. The two dominant paradigms are Upper Confidence Bound (UCB) and Thompson Sampling (TS). Upper Confidence Bound (UCB) UCB drives exploration via a bonus term added to the reward estimate, selecting: At = arg max (r̂(Xt , a) + bonust (Xt , a)) . a∈A
√︂ −1 ϕ(x), where Vt,a For linear models (e.g., LinUCB), the bonus scales with ϕ(x)⊤ Vt,a is the design matrix for action a, encouraging the selection of less-certain actions. Thompson Sampling (TS) TS (or posterior sampling) implements randomized exploration. The agent samples a reward function r̃t from the posterior and acts greedily with respect to it: At = arg max r̃t (Xt , a), a∈A
where r̃t (x, a) = r(x, a; θt ), θt ∼ p(θ | Ht ).
Exploration is implicit: high posterior uncertainty yields diverse samples θt , leading to varied actions. As data accumulates, the posterior contracts, and behavior naturally becomes exploitative3 . In large action spaces, standard UCB and TS struggle. Treating actions independently gathers information too slowly, leading to prohibitive regret. This necessitates structured models that share information across actions. 3
While we describe sampling parameters θt , the general principle involves sampling from the posterior predictive distribution of rewards.
23
Decision-Making in Off-Policy Learning Recall that the off-policy setting requires identifying an optimal policy from a static dataset collected under a distinct logging policy. The two dominant philosophies for tackling this are Greedy Policies and Pessimistic Policies. Greedy Policies Greedy methods select the policy π̂g that maximizes a point estimate of value, V̂ (π): π̂g = arg max V̂ (π). π∈Π
(1.7)
We distinguish between two primary classes of estimators. The direct method (DM) relies on the reward model r̂ derived in the previous section. In contrast, inverse propensity scoring (IPS) bypasses reward modeling to estimate the policy value directly using importance weighting4 : n
1 ∑︂ ∑︂ π(a | Xi ) r̂(Xi , a), V̂dm (π) = n i=1 a∈A n
V̂ips (π) =
1 ∑︂ π(Ai | Xi ) Ri . n i=1 π0 (Ai | Xi )
The optimization procedures for these objectives differ significantly. The policy maximizing V̂dm (π) is simply the one that acts greedily with respect to the learned reward model, π̂gdm (x) = arg max r̂(x, a) . a∈A
(1.8)
For IPS, the optimization is typically performed numerically over a class of parameterized policies πθ : π̂gips = arg max V̂ips (πθ ) . θ∈Rd
(1.9)
Greedy policies are effective when the underlying estimator is accurate. Specifically, when the reward model r̂ is well-specified (for DM), or the importance weights have low variance (for IPS). However, the arg max operator can amplify estimation errors, leading to a policy that over-exploits optimistic model inaccuracies or high-variance weight estimates. Pessimistic Policies Pessimistic methods mitigate error amplification by penalizing the objective with a quantified uncertainty term: [︂ ]︂ π̂p = arg max V̂ (π) − pen(π) . (1.10) π∈Π
4
Importance weighting allows IPS to be unbiased under the common support assumption (i.e., π0 (a|x) = 0 =⇒ π(a|x) = 0)
24
This principle applies to both DM and IPS as: [︂ ]︂ Pess-DM: π̂pdm (x) = arg max r̂(x, a) − β σ̂r (x, a) , a∈A
(1.11)
and Pess-IPS:
[︂ ]︂ π̂pips = arg max V̂ips (πθ ) − β σ̂ips (πθ ) . θ
(1.12)
Here, σ̂r captures the uncertainty of the reward model, while σ̂ips captures the uncertainty of the IPS estimator itself (e.g., its variance). The penalty term pen(π) prevents the maximization operator from selecting overestimated policies. It also regulates distributional shift by penalizing policies that place probability mass on context-action pairs with low coverage under π0 (support mismatch). Standard methods, whether relying on DM or IPS, and whether adopting greedy or pessimistic policies, face severe limitations in large action spaces. For DM, the prevailing practice of modeling action parameters independently prevents information sharing, making learning statistically inefficient. For IPS, the well-recognized issue is variance: importance weights can be large. However, we also demonstrate in this thesis that optimization can be an even greater bottleneck in large action spaces. This is because standard IPSbased objectives (whether Equation (1.9) or Equation (1.12)) induce highly non-concave landscapes with flat plateaus that trap gradient-based optimizers. Finally, for pessimistic methods, existing formulations often rely on intractable bounds that are incompatible with modern stochastic optimization techniques. This thesis addresses these specific pathologies: we introduce structured models for DM to enforce information sharing; we propose new policy-weighted log-likelihood objectives that yield superior optimization landscapes compared with IPS-based objectives; and we develop variance-reduced, tractable pessimistic objectives that are amenable to stochastic optimization at scale.
1.3
Contributions
Part I: On-Policy Learning in Large Action Spaces In the on-policy setting, exploration strategies that adopt disjoint reward models5 and learn each action parameter independently gather information slowly, resulting in high regret and failure to converge to an optimal policy within practical time horizons. Part I addresses this challenge by scaling Thompson Sampling to large action spaces while retaining the disjoint reward model parameterization. Our primary contribution is the introduction of structured Bayesian models with informative priors that share statistical strength across actions, enabling efficient exploration without sacrificing the robustness and flexibility of disjoint reward models. 5
Recall that disjoint models offer greater flexibility and are widely used in industrial settings such as large-scale recommender systems that use a separate embedding for each item, whereas joint reward models require careful feature engineering and are not, to the best of our knowledge, widely deployed in practical recommendation systems.
25
(Chapter 3) Scaling Thompson Sampling with Mixed-Effects We propose a hierarchical Bayesian framework that couples action parameters through L shared latent effect parameters Ψ = (ψℓ )ℓ∈[L] ∈ RdL , where each effect ψℓ can represent, for example, a category of items: Ψ ∼ q0 , θa | Ψ ∼ p0,a (· | Ψ) , Rt | θ, Ψ, Xt , At ∼ p(· | Xt ; θAt ) ,
∀a ∈ [K] , ∀t ∈ [T ] .
Upon this model, we build mixed-effect Thompson sampling (meTS). meTS maintains a posterior over effects qt (Ψ) = p(Ψ | Ht ) and K conditional posteriors over actions pt,a (θa | Ψ) = p(θa | Ψ, Ht ). In round t ∈ [T ], parameters are sampled hierarchically: Ψt ∼ qt (·) , θt,a ∼ pt,a (· | Ψt ) ,
∀a ∈ [K] .
Actions are then selected via the standard TS rule: At = arg maxa∈[K] r(Xt ; θt,a ). We derive exact closed-form updates for linear models and tractable Laplace approximations for generalized linear models. Theoretically, we establish a Bayesian regret bound in the linear-Gaussian case: (︃√︂ )︃ 2 2 BR(T ) = Õ T dKeff (σ0 + σΨ ) , 2 where σ02 and σΨ are the prior variances of the action and effect parameters, respectively, and Keff is the effective number of actions. When L ≪√︁K, we have Keff ≪ K, yielding a multiplicative Bayesian regret improvement of K/Keff over standard TS. Computationally, meTS exploits the conditional independence of action parameters given the latent effects, reducing memory complexity from O(K 2 d2 ) to O((L2 + K)d2 ) and runtime from O(K 3 d3 ) to O((L3 + K)d3 ). Empirically, meTS consistently outperforms baselines, with gains increasing with K.
AISTATS 2023 (Poster) - Aouali et al. (2023b): • I. Aouali, B. Kveton, and S. Katariya. Mixed-effect Thompson sampling. In International Conference on Artificial Intelligence and Statistics, pages 2087–2115. PMLR, 2023b. (Chapter 4) Scaling Thompson Sampling with Diffusion Models This chapter extends the hierarchical framework of meTS by introducing diffusion Thompson Sampling (dTS), which replaces the single-layer prior with a deep hierarchy of latent variables governed by a diffusion model : ψL ∼ N (0, ΣL+1 ) , ψℓ−1 | ψℓ ∼ N (fℓ (ψℓ ), Σℓ ) , θa | ψ1 ∼ N (f1 (ψ1 ), Σ1 ) , Rt | θ, (ψℓ )ℓ∈[L] , Xt , At ∼ p(· | Xt ; θAt ) , 26
∀ℓ ∈ [L] \ {1} , ∀a ∈ [K] , ∀t ∈ [T ] .
The link functions fℓ are pre-trained non-linear transformations (e.g., neural networks), enabling rich representations of inter-action structure. As in meTS, exploration proceeds by sampling parameters top-down through the hierarchy and selecting the reward-maximizing action. The key technical challenge is that the exact posterior is intractable due to nonlinearities in both the reward and link functions. To enable fast online updates without expensive MCMC, we derive a tractable inference procedure based on Gaussianlike updates6 . The resulting posterior preserves the diffusion structure: it remains a hierarchy of conditional Gaussians, but with fine-tuned link functions and precisions: Σ̄−1 t,ℓ−1 =
Σ−1 + Ḡt,ℓ−1 , ℓ ⏞⏟⏟⏞ ⏞ ⏟⏟ ⏞ prior precision data precision (︂ B̄t,ℓ−1 fˆt,ℓ (ψℓ ) = Σ̄t,ℓ−1 Σ−1 f (ψ ) + ⏞ ℓ ⏟⏟ℓ ℓ⏞ ⏞ ⏟⏟ ⏞ prior contribution
)︂
.
data contribution
Here, Ḡt,ℓ−1 and B̄t,ℓ−1 are sufficient statistics propagated upward through the hierarchy. As data accumulates, covariances contract and means shift from prior toward MLE. Crucially, this formulation preserves the expressiveness of diffusion models since the posterior is not a single Gaussian but a posterior diffusion model that retains the generative structure of the prior. Theoretically, we analyze dTS in the fully linear-Gaussian setting to gain analytical insight, deriving a Bayes regret bound: ⎞ ⎛⌜ ⃓ L+1 ⃓ ∑︂ σℓ2 ⎠ , BR(T ) = Õ ⎝⎷T dKeff ℓ=1
where Σℓ = σℓ2 Id and Keff ≪ K is the effective number of actions. Computationally, dTS exploits hierarchical conditional independence to reduce memory and time complexity further from O(K 2 d2 ) and O(K 3 d3 ) to O((L + K)d2 ) and O((L + K)d3 ), respectively: linear scaling in L that improves upon meTS. Empirically, dTS consistently outperforms baselines by leveraging pre-trained diffusion priors, even when these priors are imperfect or trained on limited data. NeurIPS 2025 (Poster) - (Aouali, 2025, 2023): • I. Aouali. Diffusion models meet contextual bandits. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. • I. Aouali. Linear diffusion models meet contextual bandits with large action spaces. In NeurIPS 2023 Foundation Models for Decision Making Workshop, 2023. 6
By Gaussian-like updates, we mean that the posterior precision is the sum of prior and evidence precisions, and the posterior mean is the precision-weighted combination of prior mean and maximum likelihood estimate.
27
Part II: Off-Policy Learning in Large Action Spaces In the off-policy setting, both DM and IPS face severe limitations as the action space grows. DM suffers from high model bias and statistical inefficiency due to sparse data coverage. IPS exhibits high variance and potential bias due to insufficient support; moreover, as we demonstrate in this thesis, optimizing IPS-based objectives becomes intractable in large action spaces. Part II addresses these failure modes through three complementary approaches. (Chapter 6) Scaling Direct Methods with Latent Parameters Standard DMs estimate independent d-dimensional parameters for each action, which becomes statistically inefficient when actions are rarely observed. We introduce the structured direct method (sDM), which couples action parameters through a shared latent vector ψ (analogous to Chapter 3): ψ ∼ q, θa | ψ ∼ pa (·; fa (ψ)), R | X, A, θ, ψ ∼ p(· | X; θA ). sDM computes the posterior over latent effects p(ψ | Dn ) and conditional posteriors p(θa | ψ, Dn ). The marginal posterior p(θa | Dn ) is obtained by integrating out ψ, yielding the reward estimate r̂(x, a) = E[r(x, a; θ) | Dn ], which is used in a greedy policy as: π̂g (a | x) = 1{a = arg max r̂(x, b)} . b∈A
To analyze performance, we introduce Bayesian suboptimality (BSO) and prove that √ sDM achieves O(1/ n) convergence. The result avoids the restrictive full support assumption, which requires π0 (a | x) ≥ γ > 0 for all actions. Instead, the bound depends on the alignment between the logging policy π0 and the optimal policy π∗ : performance improves smoothly as coverage of optimal actions increases. We also prove that under BSO, greedy policies are optimal and outperform pessimistic ones. This contrasts with the frequentist setting, where pessimism hedges against worstcase scenarios and is generally preferred. Experiments on synthetic and real-world data confirm that sDM outperforms existing methods, with gains increasing with K. AISTATS 2025 (Poster) - Aouali et al. (2025): • I. Aouali, V.-E. Brunel, D. Rohde, and A. Korba. Bayesian off-policy evaluation and learning for large action spaces. In International Conference on Artificial Intelligence and Statistics, pages 136–144. PMLR, 2025. (Chapter 7) Optimization Matters More Than Estimation This chapter challenges a common paradigm in off-policy learning. The field has traditionally focused on developing sophisticated value estimators with improved statistical properties, assuming that maximizing a more accurate estimator yields a better policy. We demonstrate that this emphasis is misplaced in large action spaces, where optimization error dominates estimation error, making even advanced estimators ineffective for policy learning. 28
Our key insight is that estimator-based objectives, despite their statistical appeal, induce highly non-concave landscapes when paired with standard policy classes. We show that gradient-based optimization can remain trapped in suboptimal plateaus for O(K) iterations, and that the landscape contains exponentially many local maxima in K. These pathologies make global optimization intractable for large K. To characterize these failures, we analyze the oracle policies of various estimators: the policies that maximize the estimators with infinite data. This analysis reveals that each estimator induces a distinct inductive bias. For instance, standard IPS searches within the logging policy’s support, while cluster-based methods such as MIPS (Saito and Joachims, 2022) operate at the cluster level. These insights motivate objective-aware policy parametrizations: by aligning the policy parametrization with the estimator’s bias, we reduce the effective search space from K to the significantly smaller logging support size k0 or cluster count C, partially alleviating the optimization challenges. Ultimately, we advocate for a fundamental shift: abandoning value estimation in favor of policy-weighted log-likelihood (PWLL) objectives: n
Ûg (π) =
1 ∑︂ g(Ri , π0 (Ai | Xi )) log π(Ai | Xi ) , n i=1
where g is a positive weighting function. Although PWLL objectives are not value estimators, we prove they are concave (and strongly concave with ℓ2 regularization) for linear softmax policies, while achieving oracle policies comparable to estimatorbased objectives. This guarantees efficient convergence to a unique global maximum, while eliminating the optimization pathologies. Large-scale experiments on datasets with up to one million actions validate this approach. Simple PWLL methods consistently outperform state-of-the-art estimatorbased objectives, with the performance gap widening as action spaces grow. Moreover, PWLL objectives exhibit remarkable robustness to optimization hyperparameters, whereas estimator-based methods require careful tuning and often fail under minor configuration changes. CONSEQUENCES, RecSys 2025 (Poster) - Submitted to ICLR 2026 (Aouali and Sakhi, 2025): • I. Aouali and O. Sakhi. Off-policy learning in large action spaces: Optimization matters more than estimation. Under review at ICLR, 2026. (Chapter 8) Principled Pessimism for Exponential Smoothing and Beyond While sDM and PWLL offer alternative paradigms, this chapter improves the widely used family of IPS-based methods by combining variance-reducing regularization with principled pessimism for safe policy learning. The variance of IPS scales with importance weights, which can explode in large action spaces. To control this, we introduce differentiable exponential smoothing 29
(ES) estimators that regularize these weights: n
IPS-α : IPS-β :
1 ∑︂ π(Ai | Xi ) Ri , n i=1 π0 (Ai | Xi )α )︃β n (︃ 1 ∑︂ π(Ai | Xi ) β Ri . Ṽ (π) = n i=1 π0 (Ai | Xi )
V̂ α (π) =
These estimators smoothly trade variance for bias while preserving differentiability. To learn safely with these regularized estimators, we derive a two-sided PAC-Bayes generalization bound that will be used in a pessimistic objective: ⃓ √︃ kl (θ, θ ) ⃓ ⃓ ⃓ 1 0 + ⃓V (πθ ) − V̂nα (πθ )⃓ ≤ 2n
B α (π ) ⏞ n⏟⏟ θ⏞
Regularization bias
+
kl2 (θ, θ0 ) λ + nλ 2
Varαn (πθ ) ⏞ ⏟⏟ ⏞
,
Remaining variance
where kl1,2 measure divergence between the current policy πθ and the logging policy πθ0 , Bnα captures the regularization bias, and Varαn captures the remaining variance. The exact expressions of these quantities are given in Chapter 8. The pessimistic learning objective maximizes the lower bound as: [︄ π̂ = arg max V̂nα (πθ ) − πθ
√︃
]︄ kl1 (θ, θ0 ) kl (θ, θ ) λ 2 0 − Bnα (πθ ) − − Varαn (πθ ) . 2n nλ 2
This objective penalizes policies with high bias or variance, steering optimization toward reliable regions. What is important for scalability is that all terms in the objective are empirical and differentiable, enabling end-to-end optimization via standard stochastic gradient ascent. This contrasts with prior pessimistic objectives that relied on intractable theoretical constants or were incompatible with stochastic optimization. We further present a unified PAC-Bayes framework that generalizes this approach to all importance-weight regularizers in the literature (clipping, ES, implicit exploration), enabling fair comparison through a universal set of practical pessimistic objectives. This work also laid the foundation for logarithmic smoothing (see Additional Contributions below), which refines the analysis to achieve significantly tighter bounds and sharp suboptimality guarantees. ICML 2023 (Oral) - UAI 2024 (Poster) - (Aouali et al., 2023a, 2024) • I. Aouali, V.-E. Brunel, D. Rohde, and A. Korba. Exponential smoothing for off-policy learning. In Proceedings of the 40th International Conference on Machine Learning, pages 984–1017. PMLR, 2023a. • I. Aouali, V.-E. Brunel, D. Rohde, and A. Korba. Unified PAC-Bayesian study of pessimism for offline policy learning with regularized importance sampling. In Uncertainty in Artificial Intelligence, pages 88–109. PMLR, 2024. 30
Additional Contributions This section outlines additional research conducted during this thesis. While these contributions are not included in the main manuscript, they correspond to published works involving equal or significant contributions from the author. On-Policy: Extension to Best-Arm Identification We extend hierarchical and structured modeling from regret minimization to fixedbudget best-arm identification (BAI), introducing prior-informed best-arm identification (PI-BAI): a non-adaptive algorithm that leverages prior knowledge for efficient budget allocation. We provide a fully Bayesian analysis for structured settings (e.g., linear and hierarchical bandits), departing from classical frequentist approaches. This yields the first prior-dependent guarantees on Bayesian error probability in fixed-budget BAI. PI-BAI is robust to prior misspecification and consistently outperforms baselines, including adaptive strategies, challenging the prevailing assumption that adaptivity is essential for fixed-budget exploration. AISTATS 2025 (Poster) - (Nguyen et al., 2025): • N. Nguyen, I. Aouali, A. György, and C. Vernade. Prior-dependent allocations for Bayesian fixed-budget best-arm identification in structured bandits. In Proceedings of the 28th International Conference on Artificial Intelligence and Statistics (AISTATS), 2025. Off-Policy: Extension to Logarithmic Smoothing We refine the theoretical framework for regularized IPS estimators from Chapter 8. While the bounds developed there were useful for learning in practice, they can be loose for suboptimality guarantees in certain cases. To address this, we derive a general high-order moment concentration bound for regularized estimators and identify the estimator that minimizes this bound. This analysis yields a novel pessimistic estimator, logarithmic smoothing (LS): V̂lsλ (π) =
(︃ )︃ n π(Ai | Xi ) 1 ∑︂ log 1 + λ Ri . nλ i=1 π0 (Ai | Xi )
Similar to ES, LS acts as a soft, differentiable alternative to clipping, concentrates at a sub-Gaussian rate, and achieves finite variance without requiring bounded importance weights. The resulting high-probability risk bound is provably tighter than state-of-the-art alternatives, enabling sharp suboptimality guarantees. NeurIPS 2024 (Spotlight) - (Sakhi et al., 2024): • O. Sakhi, I. Aouali, P. Alquier, and N. Chopin. Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and Learning. In Advances in Neural Information Processing Systems (NeurIPS), 2024. Off-Policy: Applications to Large-Scale Recommender Systems 31
This CIFRE thesis maintained a continuous feedback loop between theory and practice. Our work on large-scale industrial recommender systems served a dual purpose: it provided a testing ground for the off-policy methods developed in this thesis, while the real-world challenges encountered in these systems directly motivated the theoretical questions addressed in the main chapters. This resulted in several workshop publications and tutorials shared with the community (Gilotte et al., 2025; Aouali et al., 2022a,b,c, 2021). • A. Gilotte, O. Sakhi, I. Aouali, and B. Heymann. Offline contextual bandit with counterfactual sample identification. arXiv preprint arXiv:2509.10520, 2025. • I. Aouali, A. Benhalloum, M. Bompaire, A. Ait Sidi Hammou, S. Ivanov, B. Heymann, D. Rohde, O. Sakhi, F. Vasile, and M. Vono. Reward optimizing recommendation using deep learning and fast maximum inner product search. In proceedings of the 28th ACM SIGKDD conference on knowledge discovery and data mining, pages 4772–4773, 2022a.. • I. Aouali, A. Benhalloum, M. Bompaire, B. Heymann, O. Jeunen, D. Rohde, O. Sakhi, and F. Vasile. Offline evaluation of reward-optimizing recommender systems: The case of simulation. arXiv preprint arXiv:2209.08642, 2022b. • I. Aouali, A. A. S. Hammou, O. Sakhi, D. Rohde, and F. Vasile. Probabilistic rank and reward: A scalable model for slate recommendation. arXiv preprint arXiv:2208.06263, 2022c. • I. Aouali, S. Ivanov, M. Gartrell, D. Rohde, F. Vasile, V. Zaytsev, and D. Legrand. Combining reward and rank signals for slate recommendation. arXiv preprint arXiv:2107.12455, 2021.
1.4
Related Work
1.4.1
On-Policy Learning in Contextual Bandits
In the on-policy (online) setting (Slivkins, 2019; Lattimore and Szepesvari, 2019; Bubeck et al., 2012; Li et al., 2010; Chu et al., 2011), the agent must balance choosing actions that maximize current reward estimates (exploitation) with exploring other actions to improve these estimates (exploration). This trade-off is often addressed using upper confidence bounds (UCBs) (Auer et al., 2002) or Thompson sampling (TS) (Thompson, 1933). Upper confidence bound (UCB) algorithms handle the exploration-exploitation trade-off by constructing high-probability confidence intervals around reward estimates and selecting the action with the largest upper bound (Auer et al., 2002; Chu et al., 2011; Abbasi-Yadkori et al., 2011; Dani et al., 2008). The intuition is optimism under uncertainty: poorly explored actions have wide confidence intervals and thus large upper bounds, which encourages exploration. Despite strong theoretical guarantees, UCB methods are often less practical than TS due to sensitivity to confidence-parameter tuning and lack of inherent randomization. Nevertheless, UCB remains a cornerstone of bandit theory and continues to inspire new exploration strategies. Part I focuses on TS, but the 32
hierarchical principles developed there naturally extend to UCB-based exploration. Thompson sampling (TS) operates in a Bayesian framework: a prior and likelihood are specified, the agent samples rewards from the posterior at each round, and then chooses the action with the highest sampled reward. TS is randomized by construction, easy to implement, and exhibits strong empirical performance in both simulated and real-world problems (Russo and Van Roy, 2014; Chapelle and Li, 2012; Russo et al., 2018). It also enjoys strong theoretical guarantees, including optimal or near-optimal regret in a variety of models (Kaufmann et al., 2012; Agrawal and Goyal, 2013b; Korda et al., 2013; Russo and Van Roy, 2014; Agrawal and Goyal, 2017; Abeille and Lazaric, 2017; Russo and Van Roy, 2016; Lu and Van Roy, 2019). Part I advances TS by integrating informative hierarchical priors that enable efficient learning in large action spaces. Hierarchical Bayesian bandits (Bastani et al., 2019; Kveton et al., 2021; Basu et al., 2021; Simchowitz et al., 2021; Wan et al., 2021; Hong et al., 2022b; Peleg et al., 2022; Wan et al., 2022; Tomkins et al., 2021; Urteaga and Wiggins, 2018) apply TS to simple graphical models in which action parameters are typically drawn from Gaussian distributions centered at a small number of latent parameters. These works primarily address metaand multi-task learning in multi-armed bandits, transferring information across tasks or arms. Our mixed-effect Thompson sampling (Chapter 3) extends this line of work by introducing a hierarchical structure with multiple latent effect parameters in the contextual bandit setting. It also provides Bayes regret bounds and computational guarantees in the large action space regime. Our diffusion Thompson sampling (Chapter 4) further generalizes these approaches by replacing simple Gaussian hierarchies with deep, nonlinear diffusion models that capture complex inter-action dependencies through flexible link functions fℓ . Approximate Thompson sampling is a central challenge in Bayesian bandits because most posteriors are intractable and require approximate inference. Prior work (Riquelme et al., 2018; Chapelle and Li, 2012; Kveton et al., 2020) highlights the strong empirical performance of approximate TS in complex models. For mixed-effect TS (Chapter 3), we exploit the Gaussian structure of the hierarchy. We first apply a Laplace-like approximation to the reward likelihood at the action level, obtaining a Gaussian pseudo-observation on each parameter θa . Because both the priors on latent effects and on action parameters are Gaussian and the hierarchy is linear, these pseudo-observations can then be propagated exactly in closed form through the hierarchy, yielding an approximate posterior that preserves the original mixed-effect structure. For diffusion TS (Chapter 4), the prior hierarchy is defined by non-linear link functions fℓ , so Gaussian propagation is no longer exact. To retain a hierarchical diffusion model, we make an additional approximation: at each update we locally linearize the link-function updates. Combined with the Laplacelike approximation on the likelihood, this yields a chain of conditional Gaussians with updated means and precisions, i.e., a posterior diffusion model that preserves the prior hierarchy while remaining computationally tractable. Bandits with underlying structure are closely related to our setting, where we assume structured relationships among actions. In latent bandits (Maillard and Mannor, 2014; Hong et al., 2020), a single latent variable indexes multiple candidate models. In structured finite-armed bandits (Lattimore and Munos, 2014; Gupta et al., 2018), each action 33
is linked to a known mean function parameterized by a common latent parameter that is learned online. TS has also been applied to more complex structures such as graphical and combinatorial bandits (Gopalan et al., 2014; Yu et al., 2020). However, these methods do not simultaneously guarantee computational and statistical efficiency in large action spaces. Meta- and multi-task learning with UCB-style methods also has a long history (Azar et al., 2013; Gentile et al., 2014; Deshmukh et al., 2017; Cella et al., 2020; Hu et al., 2021; Cella et al., 2022; Yang et al., 2020), but these works typically adopt a frequentist perspective, analyze stronger notions of regret, and often yield conservative algorithms. In contrast, our mixed-effect and diffusion TS algorithms are Bayesian, come with Bayes regret guarantees, and are explicitly designed to exploit pre-learned structure to achieve both statistical efficiency and scalable online inference. Large action spaces. Our work directly addresses the challenge of learning with per-action parameters θa (disjoint reward models) rather than a single shared parameter θ (joint reward models). The disjoint formulation, while more expressive and widely used in practice, faces severe scalability issues that we address through hierarchical structure. Our analysis shows that both mixed-effect TS (Chapter 3) and diffusion TS (Chapter 4) achieve regret bounds that scale with an effective number of actions Keff ≪ K. The expression of Keff depends however on K. Some prior works (Foster et al., 2020; Xu and Zeevi, 2020; Zhu et al., 2022) propose bandit algorithms whose regret is independent of K. However, their setting differs substantially from ours: they assume a reward function r(x, a) = ϕ(x, a)⊤ θ with a single shared parameter θ ∈ Rd and a known mapping ϕ, whereas we consider r(x, a) = ϕ(x)⊤ θa (or simply r(x, a) = x⊤ θa ) with K separate d-dimensional action parameters. The dependence on K in our setting reflects the inherent complexity of learning individual action parameters, which is the price paid for expressiveness. Obtaining a rich, known mapping ϕ that captures complex context-action dependencies can be challenging in practice, whereas our setting mirrors common scenarios such as recommender systems where each product has its own embedding learned from data. Note that both algorithms can be applied to the joint reward-model case; in that setting, our analysis would yield a K-independent regret bound.
1.4.2
Off-Policy Learning in Contextual Bandits
The challenges of large action spaces extend beyond on-policy learning to the equally important off-policy setting (Li et al., 2011; Bottou et al., 2013; Swaminathan and Joachims, 2015a), where decisions must be made using historical data collected under different policies. This section surveys the off-policy contextual bandit literature and positions the methods developed in Part II. Off-policy learning fundamentally relies on off-policy evaluation, which estimates the value of a target policy π using data collected under a logging policy π0 . Given logged data Dn = {(Xi , Ai , Ri )}ni=1 with Ai ∼ π0 (· | Xi ), off-policy evaluation seeks to estimate V (π) without deploying π. Off-policy learning then optimizes over a policy class using this estimated value. Consequently, the prevailing paradigm is estimator-centric: first design an estimator V̂ (π) with good statistical properties, then maximize it. As we show in Chapter 7, this estimate-then-optimize approach breaks down in large action spaces because optimization error, rather than estimation error, becomes the bottleneck. 34
Inverse propensity scoring (IPS) (Horvitz and Thompson, 1952; Dudík et al., 2012) corrects for distribution shift via importance weights: n
1 ∑︂ π(Ai | Xi ) Ri . V̂ips (π) = n i=1 π0 (Ai | Xi ) While unbiased under the common support assumption, IPS suffers from extremely high variance. It can also incur substantial bias when the logging policy has deficient support (Sachdeva et al., 2020), especially in large action spaces where the logging policy can only cover a small fraction of the actions. To mitigate variance, numerous importanceweight regularization techniques have been proposed, such as weight clipping (Ionides, 2008; Bottou et al., 2013) and others (Su et al., 2020; Metelli et al., 2021; Gabbianelli et al., 2024; Swaminathan and Joachims, 2015b; Gilotte et al., 2018). Chapter 8 introduces differentiable exponential-smoothing (ES) estimators that smoothly trade bias for variance and enable gradient-based optimization. Direct methods (DM) (Jeunen and Goethals, 2021; Aouali et al., 2025) avoid importance weighting by modeling the expected reward for any context–action pair and evaluating policies using: n 1 ∑︂ ∑︂ π(a | Xi ) r̂(Xi , a). V̂dm (π) = n i=1 a∈A DM is particularly attractive in large-scale recommender systems where IPS struggles due to its high variance (Sakhi et al., 2020; Jeunen and Goethals, 2021; Aouali et al., 2022c). However, standard implementations typically rely on a disjoint model that estimates one parameter vector θa ∈ Rd per action. In large action spaces with sparse logging, this leads to severe statistical inefficiency: many actions are rarely observed and thus poorly estimated. Chapter 6 addresses this via the structured direct method (sDM), which leverages hierarchical Bayesian modeling to share statistical strength across actions. Direct methods (DM) (Jeunen and Goethals, 2021; Aouali et al., 2025) build a model of the expected reward for any context–action pairs and evaluate policies using n
1 ∑︂ ∑︂ π(a | Xi ) r̂(Xi , a). V̂dm (π) = n i=1 a∈A DM is particularly attractive in large-scale recommender systems, where IPS struggles (Sakhi et al., 2020; Jeunen and Goethals, 2021; Aouali et al., 2022c). Standard implementations, however, typically use a disjoint model that estimates one parameter vector θa per action. In large action spaces with sparse logging, this leads to severe statistical inefficiency: many actions are rarely observed and thus poorly estimated. Chapter 6 introduces the structured direct method (sDM), which addresses this limitation via hierarchical Bayesian modeling. Doubly robust (DR) estimators (Robins and Rotnitzky, 1995; Bang and Robins, 2005; Dudík et al., 2011; Dudik et al., 2014; Farajtabar et al., 2018) combine DM and IPS to achieve robustness: n )︁ 1 ∑︂ π(Ai | Xi ) (︁ V̂dr (π) = V̂dm (π) + Ri − r̂(Xi , Ai ) . n i=1 π0 (Ai | Xi ) 35
DR has become a default choice in many off-policy evaluation studies (Dudík et al., 2011; Dudik et al., 2014; Farajtabar et al., 2018; Su et al., 2020). Both of our contributions in Chapter 8 (regularized importance weights) and Chapter 6 (structured reward models) can be used to enhance the components of DR. Large-scale IPS variants. Importance-weight regularization alone is often insufficient when the action space is very large. Structural assumptions can dramatically reduce the variance. For instance, marginalized IPS (MIPS) (Saito and Joachims, 2022) clusters actions via a mapping h(a) and works with cluster-level importance weights: n
1 ∑︂ π(h(Ai ) | Xi ) Ri . V̂mips (π) = n i=1 π0 (h(Ai ) | Xi ) This reduces variance by operating over a smaller cluster space instead of the full action set. This dimensionality reduction principle has inspired numerous extensions (Peng et al., 2023; Sachdeva et al., 2024; Cief et al., 2024; Taufiq et al., 2024; Saito et al., 2023). Our sDM (Chapter 6) provides a complementary structural approach designed for DM instead of IPS; it can be viewed as a Bayesian latent-structure counterpart to MIPS, replacing hard clustering with soft probabilistic coupling. Pessimistic off-policy learning. Maximizing a point estimator (IPS, DM, or DR) can be unsafe when the estimator deviates from the true value. Pessimistic approaches instead construct lower confidence bounds on V (π) and optimize those, following the principle of pessimism in the face of uncertainty (Jin et al., 2021). Asymptotic and finite-sample lower bounds have been developed for various estimators (Bottou et al., 2013; Kuzborskij et al., 2021; Gabbianelli et al., 2024), providing worst-case guarantees on policy performance. Many pessimistic learning methods are directly motivated by such bounds (Swaminathan and Joachims, 2015a; London and Sandler, 2019; Kuzborskij et al., 2021; Aouali et al., 2023a; Wang et al., 2023). For example, Swaminathan and Joachims (2015a) combine empirical-Bernstein inequalities with clipped IPS, leading to variance-penalized learning objectives. Recently, the PAC-Bayesian paradigm (McAllester, 1998; Catoni, 2007; Alquier, 2021) has been increasingly applied to off-policy learning, offering a flexible toolkit for deriving data-dependent generalization bounds. London and Sandler (2019) introduced a scalable PAC-Bayesian perspective, which has been further developed by Flynn et al. (2023); Sakhi et al. (2022); Aouali et al. (2023a, 2024); Gabbianelli et al. (2024) to yield tight, directly optimizable bounds. Chapter 8 advances this direction by deriving a two-sided PAC-Bayes bound for exponentially smoothed IPS estimators, resulting in a fully differentiable pessimistic learning objective that jointly controls reward, bias, variance, and divergence from the logging policy. Moreover, this framework generalizes to encompass other estimators in the literature, establishing a unified set of principles for pessimistic learning. While our subsequent work on logarithmic smoothing (LS) (Sakhi et al., 2024) further tightens these guarantees, Chapter 8 lays the foundational theoretical groundwork on which LS builds. Optimization-centric learning. All methods above follow an estimator-centric philosophy: find a good value estimator and maximize it. In large action spaces, these estimatorbased objectives typically induce highly non-concave landscapes with flat plateaus and 36
many local maxima (Chapter 7), ensuring that optimization error dominates estimation error. Chapter 7 proposes a paradigm shift toward optimization-centric off-policy learning. Rather than insisting on high-quality estimators of the value, we advocate for policy-weighted log-likelihood objectives whose optimization landscape is benign, ensuring convergence to effective policies even in massive action spaces.
37
Part I On-Policy Learning in Large Action Spaces
38
Chapter 2
Introduction to Part I
This first part of the thesis addresses the following fundamental question: How can we design exploration-exploitation algorithms that remain both statistically efficient and computationally feasible when the number of actions is large?
2.1
Setting and Background
In this part, we consider the on-policy (online) contextual bandit setting where an agent interacts with an environment over T rounds. At each round t ∈ [T ]: 1. The agent observes a context Xt ∈ X ⊆ Rd drawn from a distribution ν; 2. The agent selects an action At ∈ A = [K] based on the history Ht = {(Xs , As , Rs )}t−1 s=1 ; 3. The agent receives a stochastic reward Rt ∼ p(· | Xt ; θ∗,At ). Each action a ∈ [K] is associated with an unknown true parameter θ∗,a ∈ Rd . The true expected reward is given by r(x, a; θ∗ ) = r(x; θ∗,a ), where θ∗ = (θ∗,a )a∈[K] ∈ RKd denotes the concatenation of all true action parameters. Throughout this part, we assume the reward distribution is a generalized linear model (GLM) (McCullagh and Nelder, 1989). For any context x ∈ X and action a ∈ A, p(· | x; θ∗,a ) is an exponential-family distribution with mean g(ϕ(x)⊤ θ∗,a ), where g is the mean function. This formulation recovers linear bandits (Auer, 2002) when p(· | x; θ∗,a ) = N (·; ϕ(x)⊤ θ∗,a , σ 2 ) (with identity link g(u) = u), and logistic bandits (Filippi et al., 2010) when p(· | x; θ∗,a ) = Ber(g(ϕ(x)⊤ θ∗,a )) with the sigmoid link g(u) = (1 + exp(−u))−1 . We adopt a Bayesian perspective where the unknown true parameters θ∗ are assumed to be drawn from a prior distribution p0 . Our objective is to minimize the Bayes regret: [︄ T ]︄ )︂ ∑︂ (︂ (2.1) BR(T ) = E r(Xt , At,∗ ; θ∗ ) − r(Xt , At ; θ∗ ) , t=1
where At,∗ = arg maxa∈[K] r(Xt , a; θ∗ ) is the optimal action at round t. The expectation is taken over the prior p0 , the stochastic rewards, contexts, and the agent’s policy. 39
2.1.1
Scalability Challenges
Standard Thompson sampling (TS) maintains a posterior distribution p(θ | Ht ) over the parameter space and sampling θt ∼ p(· | Ht ) at each round to select the action At = arg maxa∈[K] r(Xt , a; θt ). However, when the number of actions K is large, there is trade-off between statistical and computational efficiency: Statistical inefficiency of disjoint priors. A common simplification ∏︁ is to learn each action’s parameter θa independently using a factorized prior p0 (θ) = K a=1 p0,a (θa ). While this makes posterior updates computationally cheap, it prevents information sharing. The agent must learn about every action from scratch, which is prohibitive in large action spaces. Computational intractability of joint priors. Conversely, modeling dependencies via a full joint posterior over RdK allows for information sharing but is computationally intractable. Storing the covariance requires O(K 2 d2 ) memory, and a single update requires O(K 3 d3 ) time, making it intractable for online interaction.
2.2
Hierarchical Models
To address this dilemma, we propose a general hierarchical Bayesian framework where action parameters are coupled through a set of latent parameters Ψ = (ψℓ )ℓ∈[L] ∈ RLd , with L ≪ K. The generative process is defined as: Ψ ∼ q0 (·) θa | Ψ ∼ p0,a (· | Ψ) , Rt | θ, Ψ, Xt , At ∼ p(· | Xt ; θAt ) ,
∀a ∈ [K] ∀t ∈ [T ]
(Prior over latent structure), (Conditional prior per action), (Reward observation).
(2.2) (2.3) (2.4)
Here, q0 encodes global uncertainty, while p0,i specifies how individual actions deviate from the shared structure. The resulting marginal prior p0 (θ) naturally couples all action parameters. The corresponding posterior preserves this hierarchy: p(θ, Ψ | Ht ) = p(Ψ | Ht )
K ∏︂
p(θa | Ψ, Ht,a ).
(2.5)
a=1
The latent posterior p(Ψ | Ht ) aggregates evidence from all actions, enabling global information sharing, while the action posteriors p(θa | Ψ, Ht,a ) allow for efficient sampling of action parameters θa independently given Ψ. Thompson sampling then proceeds by first sampling the global structure Ψt , then sampling action parameters θt,a conditioned on Ψt .
2.3
Roadmap of Part I
The following chapters present two concrete instantiations of this hierarchical framework. Chapter 3: Mixed-Effects Thompson Sampling. We begin by investigating a linear instantiation of the hierarchy where each∑︁action parameter is modeled as a linear combination of L shared effects, θa | Ψ ∼ N ( Lℓ=1 ba,ℓ ψℓ , Σ0,a ), with known weights ba,ℓ . 40
We derive mixed-effect Thompson sampling (meTS), an algorithm that exploits the conjugacy of linear-Gaussian models to perform exact, closed-form posterior updates. For non-linear GLM rewards, we introduce a tractable Laplace approximation that we propagate through the hierarchy. We provide theoretical guarantees showing that meTS achieves a Bayes regret bound scaling with the effective number of actions Keff ≪ K. Chapter 4: Diffusion Thompson Sampling. We then extend the framework to support deep, non-linear hierarchical structures using diffusion models. In this setting, the priors form a Markov chain of latent variables ψL → · · · → ψ1 → θa , connected by potentially non-linear link functions fℓ (e.g., neural networks) learned from offline data. We propose diffusion Thompson sampling (dTS) and develop a posterior approximation that updates the link functions and covariances to match observed data while preserving the generative diffusion. This chapter demonstrates how to leverage powerful generative priors for exploration while remaining computationally feasible for online deployment.
41
Chapter 3
Scaling Thompson Sampling with Mixed Effects
Contents Setting and Background . . . . . . . . . . . . . . . . . . . . . . . . . .
39
2.1.1
Scalability Challenges . . . . . . . . . . . . . . . . . . . . . . .
40
2.2
Hierarchical Models . . . . . . . . . . . . . . . . . . . . . . . . . . . .
40
2.3
Roadmap of Part I . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
40
2.1
This chapter begins with the fundamental observation that the expected rewards of actions in real-world problems are often correlated. To model this phenomenon, we study a structured mixed-effect bandit environment in which each action parameter depends on one or more effect parameters that are shared across actions. Therefore, taking an action teaches the agent about its effect parameters, thereby informing it about other actions that share the same effect parameters. We present three motivating examples for this. Movie recommendation. Here, we want to recommend a movie to a user with the highest expected rating. User j and movie a are represented by vectors xj (context) and θa (action parameter), respectively. The expected rating that user j gives to movie a is x⊤ j θa . We assume that the vector xj is observed. Then the most natural idea is to learn all θa individually using standard bandit methods (Li et al., 2010; Chu et al., 2011). This is statistically inefficient when the number of movies is high. Fortunately, the movies could be organized into L categories and such information can be leveraged to explore efficiently. We present three approaches (A), (B) and (C) that do this next. (A) For each category ℓ ∈ [L], a parameter ψℓ is learned online using all interactions with the movies in category ℓ. The parameter ψℓ represents all the movies in category ℓ and is used instead of their individual θa . Therefore, this approach has a high bias, as all movies in the same category are assumed to have the same expected rating. This issue can be addressed by a better model. (B) We model each movie parameter θa as a random variable centered in its category parameter ψℓ . Now movies in the same category no longer have 42
the same expected rating due to the additional uncertainty. Both the category parameters ψℓ and movie parameters θa are learned online. The former is learned using all interactions with the movies in category ℓ, while the latter is learned using all interactions with movie a conditioned on ψℓ . The category parameter ψℓ is learned using more data, which helps to learn θa more efficiently. This is a special case of our setting. (C) The shortcoming of (B) is that each movie belongs to a single category, which is unrealistic. To address this issue, we allow movies to belong to multiple categories and then proceed as in (B). To connect with our terminology, the categories ℓ ∈ [L] denote effects; their parameters ψℓ are the effect parameters; and the movie parameters θa are the action parameters. Ad placement: Here, the agent selects a list (or slate) of M items from a catalog of L items with the objective of maximizing the click-through-rate. We assume that the agent receives only binary bandit feedback indicating whether the user clicked one of the items in the slate (Dimakopoulou et al., 2019; Rejwan and Mansour, 2020). Again, user j and slate a are represented by xj (context) and θa (action parameter), respectively. The corresponding click-through-rate is g(x⊤ j θa ), where g is the sigmoid function. The set of slates (of size K ≈ LM ) is exponentially large, which makes learning θa individually difficult. Fortunately, the slates are related through a much smaller set of items (of size L). Therefore, slates containing common items can teach the agent about one another, enabling efficient exploration. ∑︁ Efficient exploration is achieved by decomposing slate a’s parameter as θa = ℓ∈[L] ba,ℓ ψℓ + ϵa . Here ψℓ ∈ Rd is the parameter of item ℓ and ba,ℓ ∈ R is a mixing weight that captures position biases. That is, ba,ℓ = 0 if item ℓ is not in slate a, and ba,ℓ is high if item ℓ is ranked high in slate a. This captures the fact that the probability of a click on an item is influenced by its position on the slate, and this bias can be estimated offline. Finally, ϵa is a random noise that can incorporate uncertainty due to model misspecification, for instance due to an estimation error of ba,ℓ . The benefit of this decomposition is that the parameter of item ℓ, ψℓ , is learned using all interactions with the slates with item ℓ. The slate parameter θa is learned using all interactions with slate a conditioned on ψℓ . This is more statistically efficient than learning θa individually, which only uses the interactions with slate a. Drug design: Here, the goal is to find the optimal drug design in clinical trials (Durand et al., 2018). Subject j and drug a are represented by vectors xj and θa , respectively, and the expected efficacy of drug a for subject j is x⊤ j θa . Again, the most natural idea is to learn all drug parameters θa individually. This leads to statistical inefficiency when the number of candidate drugs is high. Fortunately, we can leverage the fact that drug candidates in the same trial often share components to explore efficiently. Precisely, a drug is a combination of multiple components, each with a specific dosage. Each component ℓ is represented by a parameter ψℓ , and the drug parameter θa is a known ∑︁ combination of the component parameters ψℓ weighted by their dosage. That is, θa = ℓ∈[L] ba,ℓ ψℓ + ϵa , where ba,ℓ is the dosage of component ℓ in drug a and ϵa is a random noise to incorporate uncertainty due to model misspecification. The efficacy of each component has an effect on the overall efficacy of the drug and is boosted by the dosage. In all examples, we assume an underlying structure among the actions, that they are affected by multiple effects. In some problems, it is known how the effect arises. For 43
instance, in the drug design, the actions are the drugs and the effects are their components. The mixing weight that relates an action (drug) to an effect (component) is the dosage of that component in the drug. In other problems, it may not be apparent how the effect arises and this has to be learned. We discuss this in detail in Section 3.1.3. We make the following contributions. 1) We formalize a general mixed-effect bandit framework represented by a two-level graphical model where each action is associated with a d-dimensional parameter that depends on one or multiple effect parameters. 2) We design mixed-effect Thompson sampling (meTS), which leverages this structure to be both statistically and computationally tractable. We show that closed-form posteriors can be derived for Gaussian instances and efficient approximations exist in more general cases. 3) We prove that the Bayes regret of meTS is bounded by a sum of two terms: one is associated with learning the action parameters and the other quantifies the cost of learning the effect parameters. Both terms reflect the structure of the environment and the quality of priors. 4) We show empirically that meTS and its variants perform extremely well, and are computationally efficient in both synthetic and real-world problems.
3.1
Setting
We consider the contextual bandit setting in Section 2.1. Each action a ∈ A = [K] is associated with an unknown d-dimensional action parameter θa ∈ Rd . The correlations between the action parameters arise because they are derived from L shared unknown ddimensional effect parameters, ψℓ ∈ Rd for ℓ ∈ [L]. Specifically, we assume that the action parameter θa is sampled from the action prior distribution p0,a as θa | Ψ ∼ p0,a (· | Ψ), where Ψ = (ψℓ )ℓ∈[L] ∈ RLd is a concatenation of the effect parameters. The distribution p0,a can capture sparsity, when θa depends only on a subset of Ψ; and also incorporate uncertainty due to model misspecification, when θa is not a deterministic function of Ψ. Finally, the effect parameters Ψ are sampled from a joint effect prior q0 , which is known by the agent and represents its initial uncertainty about Ψ. In summary, all variables in our environment are generated as Ψ ∼ q0 , θa | Ψ ∼ p0,a (· | Ψ) , Rt | Xt , At , θ, Ψ ∼ p(· | Xt ; θAt ) ,
(3.1) ∀a ∈ A , ∀t ∈ [T ] ,
where p(· | x; θa ) is the reward distribution of action a in context x, which only depends on parameter θa and the context x. The terminology of effect parameters arises from the fact that ψℓ affect the model parameters θa , which in turn define Rt . The effects are mixed through the action prior p0,a and hence the name mixed-effect. Our setting can be viewed as a two-level graphical model, where ψ1 , . . . , ψL are parent nodes and θ1 , . . . , θK are child nodes (Figure 3.1). The structure is represented by missing arrows from parent (effect parameters) to child (action parameters) nodes. A missing arrow from parent ψℓ to child θa means that action a is independent of the ℓ-th effect. Our model can capture all examples provided in the introduction of this chapter. For instance, in movie recommendation, the categories ℓ ∈ [L] and movies a ∈ A would 44
: taken action at round
Figure 3.1: Example of a graphical model induced by Equation (3.1). be represented by the effect parameters ψℓ and action parameters θa , respectively. The weight ba,ℓ is the relevance of movie a to category ℓ. Linearity in effects. A simple yet powerful assumption is that the action prior p0,a is parametrized by a weighted sum of effect parameters L )︂ (︂ ⃓ ∑︂ ⃓ ba,ℓ ψℓ , θa | Ψ ∼ p0,a · ⃓
∀a ∈ A ,
ℓ=1
where ba = (ba,ℓ )ℓ∈[L] ∈ RL are L known mixing weights for action a. The effect ℓ on action a is determined by ba,ℓ . As an example, ba,ℓ = 0 when action a is independent of effect ℓ. This is an important special case of our setting, since additive models are widely used in both theory and practice, as they often yield closed-form posteriors that are computationally tractable. Next we present two instances of this setting, where p0,a ∑︁ is a multivariate Gaussian with mean Lℓ=1 ba,ℓ ψℓ and covariance Σ0,a .
3.1.1
Mixed-Effect Linear Bandit
A natural joint effect prior q0 for d-dimensional effect parameters ψℓ is a multivariate Gaussian with mean µΨ ∈ RLd and covariance ΣΨ ∈ RLd×Ld . The action prior p0,a is a ∑︁ Gaussian with mean Lℓ=1 ba,ℓ ψℓ ∈ Rd and covariance Σ0,a ∈ Rd×d : Ψ ∼ N (µΨ , ΣΨ ) , L (︂ ∑︂ )︂ ba,ℓ ψℓ , Σ0,a , θa | Ψ ∼ N ℓ=1 Rt | Xt , At , θ, Ψ ∼ N (Xt⊤ θAt , σ 2 ) ,
where σ 2 > 0 is the variance of the observation noise. 45
(3.2) ∀a ∈ A , ∀t ∈ [T ] ,
3.1.2
Mixed-Effect Generalized Linear Bandit
Here the effect and action parameters are generated as in Equation (3.2) but the reward Rt is sampled from a generalized linear model (GLM) (McCullagh and Nelder, 1989), which is non-linear. In particular, p(· | Xt ; θa ) is an exponential-family distribution with mean g(Xt⊤ θa ) and the whole model is Ψ ∼ N (µΨ , ΣΨ ) , L (︂ ∑︂ )︂ θa | Ψ ∼ N ba,ℓ ψℓ , Σ0,a ,
(3.3) ∀a ∈ A ,
ℓ=1
Rt | Xt , At , θ, Ψ ∼ p(· | Xt ; θAt ) ,
∀t ∈ [T ] .
Let Ber(p) be a Bernoulli distribution with mean p. One particular choice of a GLM is g(u) = 1/(1 + exp(−u)) and p(· | Xt ; θ) = Ber(g(Xt⊤ θ)), which corresponds to a logistic bandit (Filippi et al., 2010). Remark 3. Note that in both settings, we use x⊤ θ instead of ϕ(x)⊤ θ for some featuremap ϕ, but this is just for ease of exposition, and everything generalizes smoothly to when using ϕ.
3.1.3
Structure Learning
The structures in Equations (3.2) and (3.3) may be intrinsic in some problems, such as drug design. When this is not the case, we propose the following approach to learning a proxy structure. For any a ∈ A, let θ̂a represent an offline estimate of action parameter θa (e.g., learned offline using interactions from previous bandit tasks). To learn, we fit a Gaussian mixture model (GMM) (Reynolds et al., 2009) with L clusters to θ̂a . Each cluster ℓ ∈ [L] is represented by its center µψℓ ∈ Rd and covariance Σψℓ ∈ Rd×d . These correspond to the mean of the effect parameter ψℓ and its uncertainty. The GMM also outputs the probability that θ̂a belongs to cluster ℓ, for all combinations of a ∈ A and ℓ ∈ [L]. This probability is the mixing weight ba,ℓ . The proposed procedure is general and adaptable to a wide range of use cases. The primary challenge lies in deriving the offline estimates θ̂a . A straightforward approach involves learning these parameters from historical data collected in previous bandit tasks. Broadly, this can be formulated as an offline representation-learning problem (Tripuraneni et al., 2021), for which numerous techniques exist. For instance, in our MovieLens experiments (Section 3.4.2), we employ a low-rank factorization of the rating matrix to obtain these estimates. A key strength of our approach is its flexibility; it integrates seamlessly with standard offline learning tools, thereby taking a step toward bridging the gap between offline and online learning.
3.2
Algorithm
We propose a Thompson sampling algorithm (Thompson, 1933; Russo and Van Roy, 2014; Scott, 2010), which is a natural Bayesian solution to our problem. The algorithm 46
Algorithm 1 meTS: Mixed-Effect Thompson Sampling. Input: Joint effect prior q0 , action priors p0,· Initialize q1 ← q0 and p1,· ← p0,· for t = 1, . . . , T do Sample Ψt ∼ qt for a = 1, . . . , K do Sample θt,a ∼ pt,a (· | Ψt ) θt ← (θt,a )a∈A At ← argmaxa∈A r(Xt , a; θt ) Receive reward Rt ∼ p(· | Xt ; θ∗,At ) Compute new posteriors qt+1 and pt+1,·
is based on hierarchical sampling (Lindley and Smith, 1972), which reflects the structure in our model. Before we present it, we need to introduce additional notation. We denote by Ht = (Xi , Ai , Ri )i∈[t−1] the history of all interactions of the agent up to round t, by St,a = {i ∈ [t − 1] : Ai = a} the rounds where the agent takes action a up to round t, and by Ht,a = (Xi , Ai , Ri )i∈St,a the corresponding history. Our algorithm meTS is presented in Algorithm 1. Since effect parameters are shared across actions, their posteriors exhibit dependencies. To handle this, we maintain two types of posterior densities: • A joint effect posterior qt (Ψ) = p(Ψ | Ht ) for all effect parameters Ψ in round t; • An action posterior pt,a (θ | Ψ) = p(θa | Ht,a , Ψ) for each action a ∈ A, conditioned on the effect parameters. meTS employs hierarchical sampling in each round t: 1. Sample effect parameters: Ψt ∼ qt (·) 2. Sample action parameters: θt,a ∼ pt,a (· | Ψt ) for each a ∈ A 3. Select action: At = argmaxa∈A r(Xt , a; θt ) where θt = (θt,a )a∈A . This hierarchical sampling scheme is equivalent to sampling from the exact marginal posterior p(θa | Ht ). To see this, observe that marginalizing over Ψ yields: ∫︂ p(θa | Ht ) =
p(θa , Ψ | Ht ) dΨ , ∫︂Ψ p(θa | Ψ, Ht ) p(Ψ | Ht ) dΨ ,
= ∫︂Ψ
pt,a (θa | Ψ) qt (Ψ) dΨ .
= Ψ
47
(3.4)
3.2.1
Posterior Derivations
The posteriors are computed as follows. We first express the joint effect posterior qt as qt (Ψ) ∝
K ∫︂ ∏︂ a=1
(3.5)
Lt,a (θa )p0,a (θa | Ψ) dθa q0 (Ψ) ,
θa
∏︁ where Lt,a (θa ) = (x,a,r)∈Ht,a p(r | x; θa ) is the likelihood of all observations of action a up to round t given θa . Next, for any action a ∈ A, the action posterior pt,a is expressed as (3.6)
pt,a (θa | Ψ) ∝ Lt,a (θa )p0,a (θa | Ψ) .
pt,a is similarly sparse to p0,a . Specifically, in any round t, pt,a and p0,a are parameterized by the same subset of effect parameters Ψ, since Lt,a (θa ) does not depend on Ψ. The joint effect posterior qt and action posteriors pt,a have closed forms in Gaussian models, which allows efficient sampling and theoretical analysis. Beyond these, MCMC and variational inference can be used to approximate qt and pt,a . Next we derive closedform posteriors for the mixed-effect model with linear rewards in Equation (3.2) and provide an efficient approximation for the mixed-effect model with non-linear rewards in Equation (3.3).
3.2.2
Mixed-Effect Linear Bandit
Let Gt,a = σ −2
∑︂
Xi Xi⊤ ,
Bt,a = σ −2
i∈St,a
∑︂
Ri Xi .
(3.7)
i∈St,a
be the outer product of contexts corresponding to action a up to round t, and their sum weighted by rewards, respectively. Both are scaled by the observation noise variance σ 2 . Using these quantities, the effect posterior is defined as follows. Proposition 1. For any round t ∈ [T ], the joint effect posterior is a multivariate Gaussian qt = N (µ̄t , Σ̄t ), where −1 Σ̄−1 t = ΣΨ +
∑︂
)︁ (︁ −1 −1 −1 −1 −1 b a b⊤ ⊗ Σ − Σ (G + Σ ) Σ t,a 0,a 0,a 0,a , a 0,a
(3.8)
a∈A
(︂ )︂ ∑︂ −1 −1 −1 µ̄t = Σ̄t Σ−1 µ + b ⊗ (Σ (G + Σ ) B ) . a t,a t,a 0,a 0,a Ψ Ψ a∈A
The effect posterior is additive in individual actions and each action contributes to the effect posterior mean and covariance proportionally to ba,ℓ , which is the mixture weight for θa in Equation (3.2). Proposition 1 is proved in Section A.2.1. Now we present the action posterior. 48
Proposition 2. For any round t ∈ [T ], action a ∈ A, and effect parameters Ψt , the action posterior is a multivariate Gaussian pt,a (· | Ψt ) = N (·; µ̃t,a , Σ̃t,a ), where −1 Σ̃−1 t,a = Σ0,a + Gt,a , L (︂ )︂ ∑︂ −1 µ̃t,a = Σ̃t,a Σ0,a ba,ℓ ψt,ℓ + Bt,a .
(3.9)
ℓ=1
The action posterior in Equation (3.9) is a standard multivariate Gaussian posterior whose prior depends on Ψt , which is sampled by meTS. Proposition 2 is proved in Section A.2.2.
3.2.3
Mixed-Effect Generalized Linear Bandit
Closed-form posteriors are unavailable in this setting, so approximations are required. We use a Laplace-style scheme that approximates the likelihood Lt,a (·) by a Gaussian, rather than applying Laplace to the full posterior. This choice preserves a Gaussian form that can be propagated analytically through the hierarchical updates. Let µlap t,a denote the MLE (see remark below for a discussion about the computation of 1 the MLE in practice), and let Glap t,a be the Hessian of − log Lt,a (·): µlap t,a = argmax log Lt,a (θa ), θa
Glap t,a =
∑︂
⊤ g(X ̇ i⊤ µlap t,a ) Xi Xi .
i∈St,a
We then approximate the likelihood (not the posterior) by )︁ (︁ ⊤ lap lap Lt,a (θa ) ∝ exp − 12 (θa − µlap t,a ) Gt,a (θa − µt,a ) ,
(3.10)
Substituting Equation (3.10) into Equation (3.5) yields qt (·) ≈ N (·; µ̄t , Σ̄t ), where µ̄t and Σ̄t are computed as in Proposition 1, except for the replacements Gt,a ← Glap t,a ,
lap Bt,a ← Glap t,a µt,a .
Similarly, substituting Equation (3.10) into Equation (3.6) gives pt,a (· | Ψ) ≈ N (·; µ̃t,a , Σ̃t,a ), with Gt,a ← Glap t,a ,
lap Bt,a ← Glap t,a µt,a .
1
Note that we are assuming a generalized linear [︁ where the log-likelihood of ]︁the data associated ∑︁ model with action a can be written as log Lt,a (θa ) = i∈St,a Ri Xi⊤ θa − A(Xi⊤ θa ) + C(Ri ) where C is a realvalued function and A is twice continuously differentiable, with derivative Ȧ = g representing the mean function.
49
Although these expressions follow mechanically from substituting the Gaussian likelihood approximation, the intuition is straightforward. The replacement Gt,a ← Glap t,a reflects lap lap the curvature induced by the nonlinear mean function g, while Bt,a ← Gt,a µt,a mirrors mle the linear-Gaussian case, where the MLE θ̂t,a is characterized by the normal equations mle Gt,a θ̂t,a = Bt,a , and its generalized-linear counterpart is µlap t,a . Remark 4. The MLE µlap t,a = argmaxθa ∈Rd log Lt,a (θa ) may be ill-posed. In practice, we λ 2 use a small ℓ2 -regularized estimator: µlap t,a ∈ argmaxθa ∈Rd log Lt,a (θa )− 2 ∥θa ∥2 , where λ > 0 to fix this.
3.2.4
Computational Complexity
The benefit of modeling the effect parameters is not immediately clear. Thus, it is tempting to marginalize them out, and only maintain a single joint posterior of all action parameters θ ∈ RKd . Posterior updates in this case would be complex and computationally inefficient when K ≫ L, which is common in practice. The main advantage of meTS is that the sampling of effect parameters Ψt ∼ qt allows us to use the conditional independence of actions given Ψ, and model θa | Ht,a , Ψt independently. This is more computationally efficient than modeling the joint θ | Ht when K ≫ L. To see this, suppose that all posteriors are multivariate Gaussians (Section 3.2.2). In this case, θ | Ht requires O(K 2 d2 ) space, due to storing a Kd × Kd covariance matrix; while meTS requires only O((L2 + K)d2 ) space, due to storing the covariances of qt and pt,a . Since the sampling relies on covariance inverses, the time complexity also improves. For the joint posterior, it is O(K 3 d3 ), while it is only O((L3 + K)d3 ) for meTS. One can also marginalize out the effect parameters Ψ and have K separate posteriors, one for each action parameter θa . While this improves computational efficiency, it does not model that the actions are correlated, since θa | Ht,a is modeled instead of θa | Ht . This leads to a statistical inefficiency due to the loss of information as the histories of other actions Ht,a′ are discarded. We validate this through theory (Section 3.3.2) and experiments (Section 3.4).
3.3
Analysis
This section is organized as follows. First, we state our regret bound. Then, we discuss how it captures the structure of our problem. We use Õ for the big O notation up to polylogarithmic factors.
3.3.1
Main Result
We analyze meTS in the linear setting in Section 3.1.1. Throughout, we assume that the true action parameters and rewards are generated according to the same hierarchical model used by meTS (Equation (3.2)), i.e., we operate in the fully well-specified setting. For ease of exposition, we further assume the existence of constants σ0 , σΨ , κx > 0 such 50
that Σ0,a = σ02 Id
for all a ∈ A,
2 ΣΨ = σΨ ILd ,
∥Xt ∥22 ≤ κx
for all t ∈ [T ].
The bound on ∥Xt ∥2 is standard, and we relax the other two assumptions in Section A.3. Theorem 1. For any δ ∈ (0, 1), the Bayes regret of meTS in the mixed-effect model in Section 3.1.1 is bounded as BR(T ) ≤ where c =
√︂
√︁
2 2 )K , κ (σ 2 + κb σΨ π x 0
(3.11)
2T (Ra (T ) + Re (T )) log(1/δ) + cT δ , κb = maxa∈A ∥ba ∥22 ,
T κx σ02 )︁ κx σ02 R (T ) = dKca log 1 + , , ca = (︁ κ σ 2 )︁ dσ 2 log 1 + σx 2 0 (︁ κx σ02 )︁ 2 2 )︁ (︁ κ κ σ 1 + Kκ σ x b b Ψ Ψ σ2 Re (T ) = dLce log 1 + 2 , ce = (︁ 2 )︁ . σ2 κx κb σΨ σ0 + T κx log 1 + σ2 a
(︁
The second term in Equation (3.11) is constant for δ = 1/T , in which case the above bound √ is Õ( T ). The main quantities of interest are Ra (T ) and Re (T ), and they have natural interpretations. Ra (T ) corresponds to the action regression problem: with K parameters √ of dimension d, prior width σ0 , maximum context length κx , and T observations with noise σ. The dependence of Ra (T ) on these quantities is identical to a corresponding linear bandit (Lu and Van Roy, 2019). On the other hand, Re (T ) corresponds to the effect regression problem: with L parameters of dimension d, prior width σΨ , maximum √ mixing-weight length κb , and K actions that can be viewed as observations with noise σ0 (Section 3.2.2). The dependence of Re (T ) on these quantities mimics those in Ra (T ). To simplify exposition, let κx = κb = σ = 1. Then )︃ (︃√︂ (︁ 2 )︁ 2 2 T d Kσ0 + LσΨ (1 + σ0 ) . BR(T ) = Õ This can be re-written as BR(T ) = Õ
(3.12)
)︂ (︂√︁ 2 (1+σ 2 ) Kσ 2 +LσΨ 2 0 T dKeff (σ02 + σΨ ) , where Keff = 0 σ2 +σ 2 0
Ψ
2 is the effective number of actions. When L ≪ K and σ02 ≪ σΨ , we have Keff ≪ K, yielding significant regret reduction over standard Thompson Sampling. 2 The dependence on σ02 and σΨ is natural: since Bayesian regret measures performance under the prior, smaller prior variances correspond to more informative beliefs about the true parameters, which makes learning easier and reduces regret. Conversely, larger variances reflect greater prior uncertainty and increase the difficulty of identifying the optimal action. The scaling with K, L, and d is also intuitive: fewer parameters to estimate lead to lower regret. These trends are consistent with our empirical observations in Section A.4.
51
3.3.2
Benefits of Structure
Note that we do not provide a matching lower bound. To argue that our upper bound reflects the intrinsic structure of the problem, we compare meTS to agents that either have access to more information or exploit less structure. We start with the former. Consider meTS with known effect parameters Ψ. Setting σΨ = 0 in Equation (3.12) yields the reduced regret √︂ BR(T ) = Õ( T dKσ02 ),
which no longer depends ∑︁ on L. Likewise, consider meTS under a perfectly specified linear model, in which θa = ℓ∈[L] ba,ℓ ψℓ for all a ∈ A. This corresponds to σ0 = 0, giving √︂ 2 BR(T ) = Õ( T dLσΨ ), which is independent of K. In particular, the K-dependence in our regret bound arises precisely from modeling the variability of action parameters around the effect parameters via Σ0,a . Without it, the regret of sDM is independent of K We now turn to an agent that neither knows Ψ nor models it explicitly. This agent learns only θ (Section 3.2.4) by marginalizing out Ψ in Equation (3.2): θa ∼ N
L (︂ ∑︂
)︂ ba,ℓ µψℓ , Σ̆0,a ,
∀a ∈ A,
ℓ=1 2 )Id is the marginal prior covariance and (µψℓ )ℓ∈[L] is the prior where Σ̆0,a = (σ02 + ∥ba ∥22 σΨ mean of the effects, so that µΨ = (µψℓ )ℓ∈[L] (Section 3.1.1). Importantly, marginalizing out Ψ and treating each action parameter independently discards action correlations, even though the true generative model induces correlations via the shared effects. This agent therefore uses a less structured and less informative prior than meTS.
Using the definition of Σ̆0,a and κb = maxa∈A ∥ba ∥22 = 1, the regret of this agent scales as in Equation (3.12) with σΨ = 0, except that the maximum prior variance σ02 is replaced 2 . Hence, by σ02 + σΨ (︁√︂ )︁ 2 BR(T ) = Õ T dK(σ02 + σΨ ) . When K > L (up to constants), this regret can be substantially larger √︁ than the bound for meTS in Equation (3.12). The improvement is on the order of K/L in regimes where the effects are far more uncertain than the actions, i.e., σΨ ≫ σ0 . For example, in our ad-placement setting, L is the number of catalog items, while K ≈ LM is the number of slates of size M . Thus K/L ≈ LM −1 , with typical scales such as L ≈ 106 and M ≈ 10. Our empirical results in Sections A.4 and 3.4.1 support this: meTS significantly outperforms classical methods when the effect parameters are more uncertain than the action parameters.
3.4
Experiments
We evaluate meTS on both synthetic and real-world problems. In each plot, we report the average values and their standard errors. Additional experiments are conducted in 52
Section A.4. The code is provided in this Github repository.
3.4.1
Synthetic Experiments
We start with two synthetic problems: the linear and logistic bandit settings in Equations (3.2) and (3.3), respectively. The effect prior is parameterized by µΨ = 0Ld and ΣΨ = 3ILd , the action covariance is Σ0,a = Id for all a ∈ A, and the observation noise is σ = 1. We use this setting since modeling of the effect parameters is the most beneficial when they are more uncertain than the action ones (Section 3.3.2). The context Xt is sampled uniformly from [−1, 1]d . We run 50 simulations and sample the mixing weights ba,ℓ from [−1, 1] in each run.
We consider the following baselines. For the linear setting, we compare meTS-Lin (Section 3.2.2), LinUCB (Abbasi-Yadkori et al., 2011), LinTS (Agrawal and Goyal, 2013a) and HierTS (Hong et al., 2022b). For the logistic setting, we compare meTS-GLM (Section 3.2.3), meTS-Lin (Section 3.2.2), UCB-GLM (Li et al., 2017), GLM-TS (Chapelle and Li, 2012) and HierTS (Hong et al., 2022b). GLM-UCB (Filippi et al., 2010) is excluded because it exhibits very high regret. We also include variational mean-field approximations of meTS (meTS-Lin-Fa and ∏︁ meTS-GLM-Fa), where the full Gaussian effect posterior qt is approximated as qt (Ψ) ≈ Lℓ=1 qt,ℓ (ψℓ ). This factorization enables sampling each ψℓ ∈ Rd independently and replaces operations on a full Ld × Ld covariance with blockwise updates. This improves the time and space complexities of meTS by L2 and L, respectively.
All baselines but HierTS ignore the structure. HierTS incorporates the structure similarly to meTS-Lin but only has a single effect parameter with prior N (0d , 3Id ), with the same mean and covariance as the effect parameters of meTS. To compare fairly with LinTS and GLM-TS, their marginal prior mean and covariance are chosen as 0d and Σ̆0,a = ⊤ Σ0,a + Γa ΣΨ Γ⊤ a , where Γa = ba ⊗ Id . This is to account for the uncertainty of the effect parameters despite marginalizing them out.
In Figure 3.2, we plot the regret in both problems for T = 5000, K = 100, L = 3, and d = 2 (higher values of K up to 100, 000 are tested in our additional experiment in Figure 3.3 below). meTS and its factored variant outperform all baselines that ignore the structure or incorporate it partially. Moreover, meTS-GLM outperforms meTS-Lin in the logistic bandit, which shows the benefit of the approximation in Section 3.2.3. This attests to the generality and flexibility of meTS and the posterior derivations in Section 3.2. We also show in Section A.4.1 that a higher K, L, or d leads to a higher regret due to learning more parameters, which is captured by our regret bounds. 53
Linear bandit: K = 100, L = 3, d = 2
3500
meTS-Lin meTS-Lin-Fa HierTS LinTS LinUCB
3000
Regret
2500 2000 1500
meTS-GLM meTS-GLM-Fa meTS-Lin HierTS GLM-TS UCB-GLM
1000 800 600 400
1000
200
500 0
Logistic bandit: K = 100, L = 3, d = 2
1200
0
1000
2000 3000 Round n
4000
5000
0
0
1000
2000 3000 Round n
4000
5000
Figure 3.2: Evaluation on synthetic problems. In Figure 3.3, we examine how the final cumulative regret scales with the number of actions K in the linear bandit setting, fixing L = 3 and d = 2. As K increases from 100 to 100, 000, meTS-Lin consistently achieves substantially lower regret than LinTS, with the gap widening as K grows. This demonstrates that meTS-Lin scales more favorably with the action space size by leveraging the shared effect structure, which aligns with our theoretical analysis. When K ≫ L, this structural advantage becomes increasingly pronounced. Varying K, L=3, d=2
Final cumulative regret
10000 8000
meTS-Lin LinTS
6000 4000 2000 0 100
500 1000
5000
50000100000
Number of actions K
Figure 3.3: Final cumulative regret as a function of the number of actions K in the linear bandit setting with L = 3 and d = 2.
3.4.2
MovieLens Experiments
We study the problem of movie recommendation using the MovieLens 1M dataset (Lam and Herlocker, 2016). This dataset contains one million ratings given by 6040 users to 3952 movies. We apply low-rank factorization to the rating matrix to obtain 5-dimensional representations: xj ∈ R5 for user j ∈ [6040] and θa ∈ R5 for movie a ∈ [3952]. We use 54
the movies as actions and the context Xt is sampled uniformly from user vectors xj . We consider both linear and logistic rewards. Given a user xj , the linear reward for movie θa is 2 ⊤ sampled from N (x⊤ j θa , σ ) while the logistic reward is sampled from Ber(g(xj θa )), where g is the sigmoid function. We run 50 simulations with K = 100 randomly sampled movies in each run. We compare meTS to most baselines in Section 3.4.1. We do not include UCB-GLM and GLM-UCB because their regret is very high. In LinTS and GLM-TS, the prior mean of action a is µ and its covariance is Σ̆0 = diag(v) ∈ Rd×d , where µ ∈ Rd and v ∈ Rd are the mean and variance of the movie vectors along all dimensions, respectively. The mixed-effect structure in Equations (3.2) and (3.3) is not available in this problem. Therefore, we use the approach in Section 3.1.3 to learn it. More precisely, we cluster the movies into L = 5 mixture components by training a GMM on the offline action vectors θa (Section 3.1.3). Each cluster center corresponds to an effect parameter mean µψℓ ∈ Rd and the mixing weight ba,ℓ is the probability that movie a belongs to cluster ℓ, as given by the GMM. We set the effect prior covariance as ΣΨ = 0.75 diag((Σ̆0 )ℓ∈[L] ) ∈ RLd×Ld and the prior covariance of action a as Σ0,a = 0.25 Σ̆0 ∈ Rd×d , where Σ̆0 is the same as in both LinTS and GLM-TS. This means that the marginal covariance of action a in meTS ⊤ is 0.25 Σ̆0 + 0.75 Γa ΣΨ Γ⊤ a , where Γa = ba ⊗ Id . Therefore, it is on the same order as Σ̆0,a when ∥ba ∥22 ≈ 1, and meTS is parameterized comparably to LinTS and GLM-TS. At the same time, we also model that the effect parameters are more uncertain than the action ones, since Σ0,a = 0.25 Σ̆0 while ΣΨ = 0.75 diag((Σ̆0 )ℓ∈[L] ). In Figure 3.4, we plot the regret for T = 5000 rounds. We observe that meTS has the lowest regret, even if the true rewards are not generated from a mixed-effect model. This shows the robustness of meTS to model misspecification, which we further validate in Section A.4.3. It also highlights the flexibility of our framework, where a proxy structure is learned from offline data. Linear bandit: K = 100, L = 5, d = 5
1400
300
1000
250
800
200
Regret
1200
600
200 0
0
1000
2000 3000 Round n
4000
meTS-GLM meTS-GLM-Fa meTS-Lin HierTS GLM-TS
150
meTS-Lin meTS-Lin-Fa HierTS LinTS
400
Logistic bandit: K = 100, L = 5, d = 5
350
100 50
5000
0
0
1000
2000 3000 Round n
4000
5000
Figure 3.4: Evaluation on MovieLens problems.
3.5
Conclusion
In this chapter, we introduced a mixed-effect bandit framework based on a two-level graphical model in which each action may depend on multiple underlying effects. This 55
structure enables more efficient exploration, and we designed meTS to leverage it effectively. When implemented as analyzed, meTS performs strongly on both synthetic and real-world benchmarks. Although our presentation focused on the concrete models in Equations (3.2) and (3.3), the underlying algorithmic ideas extend seamlessly to the general mixed-effect model in Equation (3.1). The methodological and theoretical tools developed here lay the groundwork for richer formulations, one of which we explore in detail in the next chapter. Our work has several limitations. First, the regret analysis assumes a well-specified prior: the true parameters must be generated from the same hierarchical model used by meTS. While some experiments suggest robustness to misspecification, formal guarantees under prior mismatch remain open. Second, closed-form posteriors are available only for linear-Gaussian rewards; generalized linear models require Laplace approximations (Section 3.2.3), which are not analyzed theoretically. Third, the mixed-effect structure must be known or learned offline, adding an additional modeling step. Finally, meTS models only two-level hierarchies; deeper latent structures, which may better capture complex action correlations, require the diffusion-based approach developed in Chapter 4.
56
Chapter 4
Scaling Thompson Sampling with Diffusion Models
Contents 3.1
3.2
3.3
3.4
3.5
Setting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
44
3.1.1
Mixed-Effect Linear Bandit . . . . . . . . . . . . . . . . . . .
45
3.1.2
Mixed-Effect Generalized Linear Bandit . . . . . . . . . . . .
46
3.1.3
Structure Learning . . . . . . . . . . . . . . . . . . . . . . . .
46
Algorithm . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
46
3.2.1
Posterior Derivations . . . . . . . . . . . . . . . . . . . . . . .
48
3.2.2
Mixed-Effect Linear Bandit . . . . . . . . . . . . . . . . . . .
48
3.2.3
Mixed-Effect Generalized Linear Bandit . . . . . . . . . . . .
49
3.2.4
Computational Complexity . . . . . . . . . . . . . . . . . . . .
50
Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
50
3.3.1
Main Result . . . . . . . . . . . . . . . . . . . . . . . . . . . .
50
3.3.2
Benefits of Structure . . . . . . . . . . . . . . . . . . . . . . .
52
Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
52
3.4.1
Synthetic Experiments . . . . . . . . . . . . . . . . . . . . . .
53
3.4.2
MovieLens Experiments . . . . . . . . . . . . . . . . . . . . .
54
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
55
In the previous chapter, we explored how action correlations can be captured using mixedeffects models, in which actions share a set of effect parameters. This approach proved effective when the underlying structure, such as categories in movie recommendation or components in drug design, can be learned through clustering. However, real-world action correlations might exhibit more complex patterns. Thus, this chapter presents an alternative approach inspired by the remarkable success of diffusion models in approximating complex distributions (Sohl-Dickstein et al., 2015; Ho et al., 2020; Dhariwal and Nichol, 57
2021; Rombach et al., 2022). Rather than explicitly modeling shared effects, we leverage pre-trained diffusion models to capture the rich structure of action parameters and use them as priors in contextual Thompson sampling. We make the following contributions. 1) We introduce a framework for contextual bandits with diffusion-derived priors and develop diffusion Thompson sampling (dTS), which is both statistically efficient and computationally tractable. dTS enables fast posterior updates and sampling via an efficient approximation inspired by exact Gaussian posteriors. 2) Beyond applying pre-trained diffusion models to contextual bandits, a key contribution is enabling efficient posterior computation and sampling for a d-dimensional parameter θ | D under a diffusion model prior, without updating the diffusion model parameters (i.e., without backpropagating through the neural network). This is relevant not only to bandits and RL but also to broader applications (Chung et al., 2022). Our approximations are motivated by exact closed-form solutions available when the diffusion model is fully linear; these solutions form the basis for our nonlinear approximations, which achieve strong empirical performance while avoiding the computational burden of standard approximate posterior sampling techniques.
4.1
Setting
We consider the contextual bandit setting in Section 2.1. Then, we define the prior distribution using a diffusion model, with a set of L consecutive unknown latent parameters ψℓ ∈ Rd for ℓ ∈ [L]. Precisely, the action parameter θa depends on the 1-st latent parameter ψL as θa | ψ1 ∼ N (f1 (ψ1 ), Σ1 ), where the link function f1 and covariance Σ1 are known. Also, the ℓ − 1-th latent parameter ψℓ−1 depends on the ℓ-th latent parameter ψℓ as ψℓ−1 | ψℓ ∼ N (fℓ (ψℓ ), Σℓ ), where fℓ and Σℓ are known. Finally, the L-th latent parameter ψL is sampled as ψL ∼ N (0, ΣL+1 ), where ΣL+1 is known. We summarize this model in Equation (4.1) below: ψL ∼ N (0, ΣL+1 ) , ψℓ−1 | ψℓ ∼ N (fℓ (ψℓ ), Σℓ ) , θa | ψ1 ∼ N (f1 (ψ1 ), Σ1 ) , Rt | θ, (ψℓ )ℓ∈[L] , Xt , At ∼ p(· | Xt ; θAt ) ,
(4.1) ∀ℓ ∈ [L]/{1} , ∀a ∈ A , ∀t ∈ [T ] .
In practice, this model can be built by pre-training a diffusion model on offline estimates of the action parameters θa . Remark 5 (Joint models). Our algorithm and analysis also apply to the case where all actions share a single unknown parameter θ ∈ Rd . Let ϕ : X (︁× [K] →)︁ Rd be a known feature map, and assume the reward distribution mean is g ϕ(x, a)⊤ θ . Then, the diffusion prior in Equation (4.1) specializes by replacing the per-action parameters (θa )a∈[K] with a single shared parameter θ: ψL ∼ N (0, ΣL+1 ), ψℓ−1 | ψℓ ∼ N (fℓ (ψℓ ), Σℓ ) , θ | ψ1 ∼ N (f1 (ψ1 ), Σ1 ) , (︁ ⃓ )︁ Rt | θ, (ψℓ )ℓ∈[L] , Xt , At ∼ p · ⃓ ϕ(Xt , At )⊤ θ , 58
(4.2) ∀ ℓ ∈ [L] \ {1}, ∀ t ∈ [T ].
This formulation is useful when a shared feature map ϕ is available. In that case, the diffusion model can be pre-trained on parameters {θs }Ss=1 from previous tasks, and dTS can then be applied to a new task S+1 using the pre-trained prior. To avoid clutter, our main exposition focuses on the model in Equation (4.1), but all theoretical results and algorithmic components extend naturally to this shared-parameter case, which we also include in some experiments (explicitly noted when applicable).
4.2
Algorithm
We design a Thompson sampling algorithm that samples the latent and action parameters hierarchically (Lindley and Smith, 1972). Let Ht = (Xi , Ai , Ri )i∈[t−1] denote the history of all interactions up to round t, and let Ht,a = (Xi , Ai , Ri ){i∈[t−1];Ai =a} be the history of interactions with action a up to round t. To motivate our algorithm, we decompose the posterior density p(θa | Ht ) recursively as ∫︂ p(ψL | Ht )
p(θa | Ht ) = ψ1:L
L ∏︂
p(ψℓ−1 | ψℓ , Ht )p(θa | ψ1 , Ht,a ) dψ1:L .
(4.3)
ℓ=2
Hierarchical sampling. This decomposition induces the following sampling procedure. First, draw a sample ψt,L according to the posterior density p(ψL | Ht ). Then, for each ℓ ∈ [L] \ {1}, draw ψt,ℓ−1 from the conditional posterior p(ψℓ−1 | ψt,ℓ , Ht ). Finally, given ψt,1 , draw each action parameter independently from p(θa | ψt,1 , Ht,a ) (the θa are conditionally independent given ψ1 ). This defines Algorithm 2, diffusion Thompson Sampling (dTS). Posterior components via recursion. To implement dTS, we provide a recursive scheme to express the required posteriors using known quantities. These expressions may not always admit closed forms and often require approximation. The conditional actionposterior can be written as ∏︂ p(θa | ψ1 , Ht,a ) ∝ p(Ri | Xi ; θa ) N (θa ; f1 (ψ1 ), Σ1 ), (4.4) i∈St,a
where St,a = {ℓ ∈ [t − 1] : Aℓ = a} is the set of rounds in which action a was selected. Now, we characterize the conditional latent-posteriors. Before we do so, we make the following notation clarification. With slight abuse of notation, p(Ht | ψℓ ) denotes the likelihood of the observations up to round t given ψℓ : p(Ht | ψℓ ) = p((Ri )i<t | (Xi )i<t , (Ai )i<t , ψℓ ) With this notation in mind, for any ℓ ∈ [L] \ {1}, the conditional latent-posterior is p(ψℓ−1 | ψℓ , Ht ) ∝ p(Ht | ψℓ−1 ) N (ψℓ−1 ; fℓ (ψℓ ), Σℓ ), and the top-layer posterior is p(ψL | Ht ) ∝ p(Ht | ψL ) N (ψL ; 0, ΣL+1 ). 59
All terms above are known except the likelihoods p(Ht | ψℓ ), which are computed recursively. The recursion starts with [︄ ]︄ K ∫︂ ∏︂ ∏︂ p(Ht | ψ1 ) = p(Ri | Xi ; θa ) N (θa ; f1 (ψ1 ), Σ1 ) dθa , (4.5) a=1
θa
i∈St,a
and for ℓ ∈ [L] \ {1}, proceeds as ∫︂ p(Ht | ψℓ ) = p(Ht | ψℓ−1 ) N (ψℓ−1 ; fℓ (ψℓ ), Σℓ ) dψℓ−1 .
(4.6)
ψℓ−1
Algorithm 2 dTS: diffusion Thompson Sampling Input: Prior components {fℓ , Σℓ }L+1 ℓ=1 and reward model p. for t = 1, . . . , T do Draw ψt,L according to the posterior density p(ψL | Ht ) for ℓ = L, . . . , 2 do Draw ψt,ℓ−1 according to p(ψℓ−1 | ψt,ℓ , Ht ) for a = 1, . . . , K do Draw θt,a according to p(θa | ψt,1 , Ht,a ) Select action At = argmaxa∈[K] r(Xt , a; θt ), where θt = (θt,a )a∈[K] Observe reward Rt ∼ p(· | Xt ; θ∗,At ) and update the posteriors. All posterior expressions above use known quantities (fℓ , Σℓ , p(r | x; θ)). However, these expressions typically need to be approximated, except when the link functions fℓ are linear and the reward distribution p(· | x; θ) is linear-Gaussian, where closed-form solutions can be obtained with careful derivations. These approximations are not trivial, and prior studies often rely on computationally intensive approximate sampling algorithms. In the following sections, we explain how we derive our efficient approximations which are motivated by the closed-form solutions of linear instances.
4.2.1
Posterior Approximation
The reward distribution is parameterized as a generalized linear model (GLM) (McCullagh and Nelder, 1989), which allows for non-linear rewards. In addition, the diffusion model itself is highly non-linear due to the link functions fℓ . These two sources of non-linearity make the posterior intractable, so we apply two layers of approximation: (i) a likelihood approximation to linearize the reward model, and (ii) a diffusion approximation to handle the non-linear hierarchy induced by the diffusion model prior. (i) Likelihood approximation. We use an approach similar to the Laplace approximation, but instead of approximating the entire posterior, we approximate only the likelihood by a Gaussian. Precisely, the reward distribution p(· | x; θa ) belongs to the exponential family with mean function g. Thus ∏︂ (︁ )︁ p(Ri | Xi ; θa ) ≈ N θa ; B̂t,a , Ĝ−1 (4.7) t,a , i∈St,a
60
where B̂t,a is the maximum likelihood estimate and Ĝt,a is the Hessian of the negative log-likelihood: ∑︂ ∑︂ (︁ )︁ (4.8) log p(Ri | Xi ; θa ), Ĝt,a = ġ Xi⊤ B̂t,a Xi Xi⊤ , B̂t,a = argmax θa ∈Rd
i∈St,a
i∈St,a
and St,a = {ℓ ∈ [t − 1] : Aℓ = a} is the set of rounds in which action a was selected. Of course, Ĝt,a might not be invertivle and thus we replace it by Ĝt,a + 10−3 Id in practice. Unlike Laplace, which fits a global Gaussian to the full posterior, this step linearizes only the likelihood, thereby preserving the hierarchical diffusion structure of the prior. (ii) Diffusion approximation. Plugging the Gaussian likelihood approximation (4.7) into the posterior expressions p(θa | ψ1 , Ht,a ) and p(ψℓ−1 | ψℓ , Ht ) removes the non-linearity of the reward model. However, the diffusion hierarchy remains non-linear through fℓ . To handle this, we build on the closed-form posteriors of the linear diffusion case (where fℓ (ψℓ ) = Wℓ ψℓ ; see Section B.1) and generalize them by replacing the linear terms Wℓ ψℓ with their non-linear counterparts fℓ (ψℓ ). This substitution yields a posterior diffusion model that retains the same hierarchical form as the prior but with data-dependent means and covariances for the conditional Gaussians. Details on how we transition from the linear to the general non-linear setting are provided in Sections B.1 and B.2. The resulting approximate posteriors admit the following closed-form expressions. Approximate action posterior. We approximate the conditional action posterior as (︁ )︁ p(θa | ψ1 , Ht,a ) ≈ N θa ; µ̂t,a , Σ̂t,a , where Σ̂−1 t,a =
Σ−1 1
+
⏞⏟⏟⏞
prior precision
Ĝt,a ⏞⏟⏟⏞
,
µ̂t,a = Σ̂t,a
(︂
Σ−1 1 f1 (ψ1 ) ⏞
⏟⏟
⏞
+
prior contribution
data precision
Ĝ B̂ ⏞ t,a⏟⏟ t,a⏞
)︂
.
data contribution
(4.9) This posterior update has a clear interpretation. The posterior precision Σ̂−1 t,a is the sum of the prior precision and the data precision. The posterior mean µ̂t,a is the precisionweighted average of the prior mean and the MLE B̂t,a . As more data are observed, the covariance shrinks and the mean moves from the prior mean f1 (ψ1 ) toward the MLE B̂t,a . When no data are available (Ĝt,a = 0), the posterior reduces to the prior N (f1 (ψ1 ), Σ1 ); in the limit of infinite data (Ĝt,a → ∞), the posterior collapses to the MLE B̂t,a , with µ̂t,a → B̂t,a and Σ̂t,a → 0. Approximate latent posteriors. For each ℓ ∈ [L + 1] \ {1}, we approximate the latent posterior as (︁ )︁ p(ψℓ−1 | ψℓ , Ht ) ≈ N ψℓ−1 ; µ̄t,ℓ−1 , Σ̄t,ℓ−1 , with Σ̄−1 t,ℓ−1 =
Σ−1 ℓ ⏞⏟⏟⏞
prior precision
+
Ḡt,ℓ−1 , µ̄t,ℓ−1 = Σ̄t,ℓ−1 ⏞ ⏟⏟ ⏞
data precision
61
(︂
Σ−1 f (ψ ) + ⏞ ℓ ⏟⏟ℓ ℓ⏞
prior contribution
B̄t,ℓ−1 ⏞ ⏟⏟ ⏞
data contribution
)︂
, (4.10)
where, by convention, fL+1 (ψL+1 ) = 0 since the top layer ψL has no parent ψL+1 . The quantities Ḡt,ℓ and B̄t,ℓ are computed recursively. The base recursion is K ∑︂ (︁ −1 )︁ −1 Ḡt,1 = Σ1 − Σ−1 , 1 Σ̂t,a Σ1
B̄t,1 = Σ−1 1
a=1
K ∑︂
Σ̂t,a Ĝt,a B̂t,a ,
(4.11)
a=1
and for each ℓ ∈ [L] \ {1}, −1 −1 Ḡt,ℓ = Σ−1 ℓ − Σℓ Σ̄t,ℓ−1 Σℓ ,
B̄t,ℓ = Σ−1 ℓ Σ̄t,ℓ−1 B̄t,ℓ−1 .
(4.12)
The latent posterior update in Equation (4.10) has the same structure as the action posterior. The posterior precision Σ̄−1 t,ℓ−1 is the sum of the prior and data precisions , and the posterior mean is their precision-weighted combination. The data terms Ḡt,ℓ−1 and B̄t,ℓ−1 are computed recursively (Equations (4.11) and (4.12)), so information collected at the action level propagates upward through the hierarchy. Interpretation. The resulting approximate posterior remains a diffusion model whose conditional Gaussians have updated, data-dependent means and covariances. The latentposterior means can be viewed as refined link functions: (︁ )︁ fˆt,ℓ (ψℓ ) = µ̄t,ℓ−1 = Σ̄t,ℓ−1 Σ−1 fℓ (ψℓ ) + B̄t,ℓ−1 , ℓ
and Σ̄t,ℓ represents their updated uncertainty. Both are updated with data: covariances contract as uncertainty decreases, and means move from the prior toward the MLE. Unlike a full Laplace approximation, this formulation preserves the expressiveness of the posterior rather than replacing it globally with a single Gaussian, while also avoiding the heavy computation required by other approximate inference methods.
4.2.2
Extension to Joint Reward Models
For the shared-parameter model in Remark 5, dTS’s posterior approximations are similar. The action posterior is p(θ | ψ1 , Ht ) ≈ N (µ̂t , Σ̂t ), where (︁ )︁ −1 Σ̂−1 µ̂t = Σ̂t Σ−1 (4.13) t = Σ1 + Ĝt , 1 f1 (ψ1 ) + Ĝt B̂t . where B̂t = argmax θ∈Rd
∑︂
∑︂ (︁ (︁ )︁ )︁ log p Ri | ϕ(Xi , Ai )⊤ θ , Ĝt = ġ ϕ(Xi , Ai )⊤ B̂t ϕ(Xi , Ai )ϕ(Xi , Ai )⊤ .
i<t
i<t
Similarly, for ℓ ∈ [L + 1] \ {1}, the latent posterior is p(ψℓ−1 | ψℓ , Ht ) ≈ N (µ̄t,ℓ−1 , Σ̄t,ℓ−1 ), where (︁ −1 )︁ −1 f (ψ ) + B̄ , (4.14) Σ̄−1 = Σ + Ḡ , µ̄ = Σ̄ Σ t,ℓ−1 t,ℓ−1 t,ℓ−1 ℓ ℓ t,ℓ−1 t,ℓ−1 ℓ ℓ where, by convention, fL+1 (ψL+1 ) = 0 and the quantities Ḡt,ℓ and B̄t,ℓ are computed recursively as Base case: Recursive case:
−1 −1 B̄t,1 = Σ−1 Ḡt,1 = Σ−1 1 Σ̂t Ĝt B̂t . 1 − Σ1 Σ̂t Σ1 , −1 −1 −1 −1 B̄t,ℓ = Σℓ Σ̄t,ℓ−1 B̄t,ℓ−1 . Ḡt,ℓ = Σℓ − Σℓ Σ̄t,ℓ−1 Σℓ ,
(4.15) (4.16)
Again, this shared-parameter variant of dTS is presented for completeness and to illustrate the generality of our posterior derivations; the main focus of the chapter remains on the per-action disjoint formulation in Equation (4.1). Unless stated otherwise, all theoretical results and experiments use the main version of dTS described in Algorithm 2. 62
4.3
Analysis
In this section, we present an informal Bayes regret analysis of dTS to build intuition around dTS’s Bayesian regret scaling with problem parameters d, K, L, etc. This analysis is informal for two reasons. First, we analyze a simplified linear-Gaussian setting rather than the general nonlinear case on which we focus in this chapter: the reward distribution is linear-Gaussian and each link function fℓ (ψℓ ) = Wℓ ψℓ is a known linear mapping, inducing a hierarchy of L linear-Gaussian layers from the latent root to the action parameters. Second, we assume the model is well-specified (similar to Chapter 7): the true action parameters are generated according to the diffusion prior used by dTS. Under these assumptions, the posterior becomes exact, enabling an analysis analogous to that used in Chapter 3. However, our recursive hierarchical structure introduces technical differences: posteriors must be derived inductively using total covariance decompositions, and regret bounds require tracking information flow across all latent layers. We emphasize that this regret bound does not extend to the general nonlinear case studied in experiments; it is included here solely to provide theoretical intuition under simplifying assumptions. Formal statements and derivations are provided in Sections B.4 and B.5. Bayes regret bound. The bound of dTS in this case is ⌜ L )︂ (︂⃓ ⃓ ∑︂ 2 2ℓ ) , σℓ+1 σmax BR(T ) = Õ ⎷T (dKσ12 + d ℓ=1
√︂
∑︁ σ2 2 2 where σmax = maxℓ∈[L+1] 1+ σℓ2 . This can be re-written as BR(T ) = Õ( T dKeff L+1 ℓ=1 σℓ ), ∑︁ 2 σ 2ℓ Kσ 2 + L σℓ+1 max where Keff = 1 ∑︁ℓ=1 is the effective number of actions. This dependence on the L+1 2 σ ℓ=1 ℓ
horizon T aligns with prior Bayes regret bounds scaling with T . However, the bound comprises L + 1 main terms. First, one relates to action parameters learning, conforming to a standard form (Lu and Van Roy, 2019), while the L remaining terms are associated with learning each of the latent parameters. Sparsity refinement. If each mixing matrix exhibits column sparsity, that, Wℓ = (W̄ℓ , 0d,d−dℓ ) with dℓ ≪ d active columns, then the bound becomes ⌜ L )︂ (︂⃓ ⃓ ∑︂ ⎷ 2 2 2ℓ T (dKσ1 + dℓ σℓ+1 σmax ) . BR(T ) = Õ ℓ=1
Hence, informative, sparse priors can cut the cost of learning deep latent chains down from d to dℓ . As in Chapter 3, a less informative prior (such as high variance) leads to a more challenging problem and thus a higher bound. Therefore, smaller values of K, L, d, dℓ translate to fewer parameters to learn, leading to lower regret. The regret also decreases when the initial variances σℓ2 decrease. These dependencies are common in Bayesian analysis, and empirical results match them. Dependence on K. The reader may question why our bound depends on K. This dependence arises from two modeling choices. First, we study the disjoint (per-action) setting 63
r(x, a; θ) = x⊤ θa , where θ = (θa )a∈[K] ∈ RdK , requiring the learning of Kd parameters. Second, we model the relationship between θa and ψ1 stochastically as N (W1 ψ1 , σ12 Id ) to accommodate potential nonlinearity. While this choice confers robustness to model misspecification, it introduces additional uncertainty and requires learning both the action parameters θa and the latent parameters ψℓ , resulting in a bound that depends on both K and L. Despite this dependence, ∑︁ 2 dTS enjoys two key advantages. First, the regret scales with 2 Kσ1 rather than K ℓ σℓ , which is particularly beneficial when σ1 is small, as is often the case with diffusion model priors. Second, thanks to informative priors, our bound has significantly smaller constants compared to both the Bayesian and frequentist regret bounds for LinTS. We demonstrate this empirically in Section B.6.5 and provide a theoretical comparison in Section 4.3.1. Both analyses confirm that dTS’s advantage over LinTS increases as the action space grows. Can regret be independent of K? Prior works (Foster et al., 2020; Xu and Zeevi, 2020; Zhu et al., 2022) have proposed bandit algorithms whose regret does not scale with K. However, these results apply to the shared-parameter setting r(x, a; θ) = ϕ(x, a)⊤ θ, where only a single d-dimensional parameter must be learned, but this formulation requires access to a suitable feature map ϕ. dTS is compatible with this setting (Section 4.2.2), in which case its regret would indeed be independent of K. Alternatively, even in the disjoint per-action case considered in this chapter, setting σ1 = 0 would yield a K-independent regret bound. However, we believe this assumption is unrealistic in practice and would compromise the robustness of dTS to model misspecification.
4.3.1
Benefits
Computational benefits. Action correlations prompt an intuitive approach: marginalize all latent parameters and maintain a joint posterior of (θa )a∈[K] | Ht . Unfortunately, this is computationally inefficient for large action spaces. To illustrate, suppose that all posteriors are multivariate Gaussians. Then maintaining the joint posterior (θa )a∈[K] | Ht necessitates converting and storing its dK × dK-dimensional covariance matrix, leading to O(K 3 d3 ) and O(K 2 d2 )(︁(︁time and complexities. )︁ space )︁ (︁(︁ )︁ 2In )︁ contrast, the time and space 3 complexities of dTS are O L + K d and O L + K d . This is because dTS requires converting and storing L + K covariance matrices, each being d × d-dimensional. The improvement is huge when K ≫ L, which is common in practice. Certainly, a more straightforward way to enhance computational efficiency is to discard latent parameters and maintain K individual posteriors, each relating to (︁an action parameter θa ∈ Rd )︁ (︁ )︁ (LinTS). This improves time and space complexity to O Kd3 and O Kd2 . However, LinTS maintains independent posteriors and fails to capture the correlations among actions; it only models θa | Ht,a rather than θa | Ht as done by dTS. Consequently, LinTS incurs higher regret due to the information loss caused by unused interactions of similar actions. Our regret bound and empirical results reflect this aspect. Statistical benefits. We argue that our bound reflects the overall structure of the problem by comparing dTS to algorithms that only partially use the structure or do not use it at all as follows. Precisely, when the link functions are linear, we can transform the diffusion prior into a Bayesian linear model (LinTS) by marginalizing out the latent 64
parameters; in which case the prior on action parameters becomes θa ∼ N (0, Σ), with the θa being not necessarily independent, and initial ∑︁L Σ is2 the marginal ∏︁ℓ covariance of 2 ⊤ action parameters and it writes Σ = σ1 Id + ℓ=1 σℓ+1 Bℓ Bℓ with Bℓ = i=1 Wi . Then, it is tempting to directly apply LinTS to solve our problem. This approach will induce higher regret because the additional uncertainty of the latent parameters is accounted for in Σ despite integrating them. This causes the marginal action uncertainty Σ to be much than the conditional action uncertainty σ12 Id , since we have Σ = σ12 Id + ∑︁L higher 2 ⊤ 2 ℓ=1 σℓ+1 Bℓ Bℓ ≽ σ1 Id . This discrepancy leads to higher regret, especially when K is large. This is due to LinTS needing to learn K independent d-dimensional parameters, each with a considerably higher initial covariance Σ. This is also reflected by our regret 2 bound. To simply comparisons, suppose that σ ≥ maxℓ∈[L+1] σℓ so that σmax ≤ 2. Then ℓ 2ℓ the regret bounds of dTS (where we bound σmax by 2 ) and LinTS read ⌜ ⃓ L ∑︂ (︁⃓ )︁ ⎷ 2 2 dTS : Õ 2ℓ ) , T (dKσ1 + dℓ σℓ+1
⌜ ⃓ L ∑︂ (︁⃓ )︁ ⎷ 2 2 LinTS : Õ ) . T dK(σ1 + σℓ+1
ℓ=1
ℓ=1
Then regret improvements are captured by the variances σℓ and the sparsity dimensions dℓ , and we proceed to illustrate this through the following scenarios. (I) Decreasing variances. Assume that σℓ = 2ℓ for any ℓ ∈ [L + 1]. Then, the regrets become ⌜ ⃓ L ∑︂ (︁⃓ )︁ (︁√︁ )︁ ⎷ dTS : Õ T (dK + dℓ 4ℓ )) , LinTS : Õ T dK2L ) ℓ=1
Now to see the order of gain, assume the problem is high-dimensional (d ≫)︁ 1), and set (︁√︁ d L = log2 (d) and dℓ = ⌊ 2ℓ ⌋. Then the regret of dTS becomes Õ nd(K + L)) , and hence the multiplicative factor 2L in LinTS is removed and replaced with a smaller additive factor L. (II) Constant variances. Assume that σℓ = 1 for any ℓ ∈ [L + 1]. Then, the regrets become ⌜ ⃓ L ∑︂ )︁ (︁√︁ )︁ (︁⃓ ⎷ T (dK + dℓ 2ℓ )) , LinTS : Õ T dKL) dTS : Õ ℓ=1
(︁√︁ )︁ Similarly, let L = log2 (d), and dℓ = ⌊ 2dℓ ⌋. Then dTS’s regret is Õ T d(K + L) . Thus the multiplicative factor L in LinTS is removed and replaced with the additive factor L. By comparing this to (I), the gain with decreasing variances is greater than with constant ones. In general, diffusion models use decreasing variances (Ho et al., 2020) and hence we expect great gains in practice. All observed improvements in this section could become even more pronounced when employing non-linear diffusion models. In our theory, we used linear diffusion models, and yet we can already discern substantial differences. Moreover, under non-linear diffusion Equation (4.1), the latent parameters cannot be analytically marginalized, making LinTS with exact marginalization inapplicable. 65
Regret
Regret
Linear diffusion, linear reward 1e3 K=100, L=2, d=5
9 8 7 6 5 4 3 2 1 0
4.0 3.5 3.0 2.5 2.0 1.5 1.0 0.5 0.0
dTS-LL HierTS LinTS LinUCB
0
1000 2000 3000 4000 5000
1e5 K=10000, L=4, d=20 dTS-LL HierTS LinTS LinUCB
Linear diffusion, nonlinear reward 1e3 K=100, L=2, d=5
1.8 1.6 1.4 1.2 1.0 0.8 0.6 0.4 0.2 0.0
3.0
dTS-LN dTS-LL GLM-TS UCB-GLM
1000 2000 3000 4000 5000
0.6 0.4 0.2
0
0.0
1000 2000 3000 4000 5000
1e3 K=10000, L=4, d=20
4.0 3.5 3.0 2.5 2.0 1.5 1.0 0.5 0.0
dTS-LN dTS-LL GLM-TS UCB-GLM
2.5 2.0 1.5 1.0 0.0
dTS-NL LinTS LinUCB
0.8
0.5 0
Nonlinear diffusion, linear reward 1e4 K=100, L=2, d=5
1.0
0
1000 2000 3000 4000 5000
Round t 2 [n]
0
1000 2000 3000 4000 5000
1e7 K=10000, L=4, d=20 dTS-NL LinTS LinUCB
0
1000 2000 3000 4000 5000
Round t 2 [n]
Round t 2 [n]
Nonlinear diffusion, nonlinear reward 1e3 K=100, L=2, d=5 1.6 1.4 dTS-NN 1.2 dTS-NL 1.0 GLM-TS 0.8 UCB-GLM 0.6 0.4 0.2 0.0
8 7 6 5 4 3 2 1 0
0
1000 2000 3000 4000 5000
1e2 K=10000, L=4, d=20 dTS-NN dTS-NL GLM-TS UCB-GLM
0
1000 2000 3000 4000 5000 Round t 2 [n]
Figure 4.1: Regret of dTS with varying diffusion and reward models and varying parameters d, K, L.
4.4
Experiments
Experimental setup. We evaluate dTS using both synthetic and MovieLens problems. In our experiments, we run 50 random simulations and plot the average regret with standard error. Our main contribution is to demonstrate that pretraining a diffusion model offline enables the construction of expressive and informative priors that substantially improve exploration efficiency in contextual bandits. We first evaluate dTS in a setting where the prior matches the true generative process (Section 4.4.1 to isolate the benefit of informative priors), and then consider a misspecified regime (Section 4.4.2 and Section B.6) where the prior is either trained on out-of-distribution data or intentionally perturbed. These experiments show that even when the prior is imperfect, dTS maintains strong performance: highlighting its robustness and practical relevance.
4.4.1
True Prior is a Diffusion Model
Synthetic bandit problems are generated from the diffusion model in Equation (4.1) with both linear and non-linear rewards. Linear rewards follow p(· | x; θa ) = N (x⊤ θa , 1), while non-linear rewards are binary from p(· | x; θa ) = Ber(g(x⊤ θa )), with g as the sigmoid function. Covariances are Σℓ = Id , and contexts Xt are uniformly drawn from [−1, 1]d . We vary d ∈ {5, 20}, L ∈ {2, 4}, K ∈ {102 , 104 }, and set the horizon to T = 5000, considering both linear and non-linear models. Linear diffusion. We consider Equation (4.1) with fℓ (ψ) = Wℓ ψ, where Wℓ uniformly drawn from [−1, 1]d×d . Sparsity is introduced by zeroing the last dℓ columns of Wℓ as Wℓ = (W̄ℓ , 0d,d−dℓ ). For d = 5 and L = 2, (d1 , d2 ) = (5, 2); for d = 20 and L = 4, (d1 , d2 , d3 , d4 ) = (20, 10, 5, 2). Non-linear diffusion. We consider Equation (4.1) where fℓ are 2-layer neural networks with random weights in [−1, 1], ReLU activation, and hidden layers of size h = 20 for d = 5, and h = 60 for d = 20. 66
Baselines. For linear rewards, we use LinUCB (Abbasi-Yadkori et al., 2011), LinTS (Agrawal and Goyal, 2013a), and HierTS (Hong et al., 2022b), marginalizing out all latent parameters except ψL , which corresponds to HierTS-1 in Section B.3. For nonlinear rewards, we include UCB-GLM (Li et al., 2017) and GLM-TS (Chapelle and Li, 2012). We exclude GLM-UCB (Filippi et al., 2010) due to high regret and HierTS as it’s designed for linear rewards. We name dTS as dTS-dr, where d refers to diffusion type (L for linear, N for non-linear) and r indicates reward type (L for linear, N for non-linear). For example, dTS-LL signifies dTS in linear diffusion with linear rewards. Results and interpretations. Results are shown in Figure 4.1 and we make the following observations: 1) dTS demonstrates superior performance (Figure 4.1). dTS consistently outperforms the baselines across all settings, including the four combinations of linear/non-linear diffusion and reward (columns in Figure 4.1) and both bandit settings with varying K, L, and d (rows in Figure 4.1). 2) Latent diffusion structure may be more important than the reward distribution. When rewards are non-linear (second and fourth columns in Figure 4.1), we include variants of dTS that use the correct diffusion prior but the wrong reward distribution, applying linear-Gaussian instead of logistic-Bernoulli (dTS-LL in the second column and dTS-NL in the fourth). Despite the reward misspecification, these variants outperform models using the correct reward distribution but ignoring the latent diffusion structure, such as GLM-TS and UCB-GLM. This highlights the importance of accounting for latent structure, which can be more critical than an accurate reward distribution. 3) Performance gap between dTS and LinTS widens as K increases (Figure 4.2a). To show dTS’s improved scalability, we evaluate its performance with varying values of K ∈ [10, 5 × 104 ], in the linear diffusion and rewards setting. Figure 4.2a shows the final cumulative regret for varying K values for both dTS-LL and LinTS, revealing a widening performance gap as K increases. 4) Regret scaling with K, d and L matches our theory (Figure 4.2b). We assess the effect of the number of actions K, context dimension d, and diffusion depth L on dTS’s regret. Using the linear diffusion and rewards setting, for which we have derived a Bayes regret upper bound, we plot dTS-LL’s regret across varying values of K ∈ {10, 100, 500, 1000}, d ∈ {5, 10, 15, 20}, and L ∈ {2, 4, 5, 6} in Figure 4.2b. As predicted by our theory, the empirical regret increases with larger values of K, d, or L, as these make the learning problem more challenging, leading to higher regret. 5) Diffusion prior misspecification (Figure 4.2c). Here, dTS’s diffusion prior parameters differ from the true diffusion prior. In the linear diffusion and reward setting, we replace the true parameters Wℓ and Σℓ with misspecified ones, Wℓ + ϵ1 and Σℓ + ϵ2 , where ϵ1 and ϵ2 are uniformly sampled from [v, v +0.5]d×d , with v controlling the misspecification level. We vary v ∈ {0.5, 1, 1.5} and assess dTS’s performance, comparing it to the wellspecified dTS-LL and the strongest baseline in this fully-linear setting, HierTS. As shown in Figure 4.2c, dTS’s performance decreases with increasing misspecification but remains superior to the baseline, except at v = 1.5, where their performances are comparable. Additional misspecification experiments are presented in Section 4.4.2, where the bandit 67
Effect of prior misspecification
Effect of K; d; L on the regret of LindTS 3500
2500
3000 2000
2000 1500
Increasing K
1000
Increasing d
500
Increasing L
0 0.0
50000
0.5
1.0
1.5
2.0
2.5
1500
LindTS (v=0.5) LindTS (v=1) LindTS (v=1.5) LindTS HierTS
1000 500 0
3.0
0
1000
Increasing values of either K; d or L
Number of actions K
(a) Perf. gap w.r.t. K.
Regret
2500 Regret
Regret in round n
1e4 Regret as a function of K 1.8 dTS-LL 1.6 LinTS 1.4 1.2 1.0 0.8 0.6 0.4 0.2 0.0 10 100 500 5000
(b) Scaling w.r.t. K, d, L.
2000
3000
4000
5000
Round t 2 [n]
(c) Prior misspecification.
Figure 4.2: Effect of various factors on dTS’s performance. 4.0
3.5 3.0 2.5 2.0 1.5 1.0 0.5 0.0 10
50
100 2500 5000 1000050000
Number of pre-training samples
(a) Ratio of LinTS/dTS cumulative regret in the last round with varying pre-training sample size in [10, 5×104 ]. Higher values mean a bigger performance gap.
8
3.5
Cumulative regret
LinTS regret / dTS regret
LinTS regret / dTS regret
4.0
3.0 2.5 2.0 1.5 1.0
7 6 5 4 3 1 0
2
10
40
70
100
Difusion depth L
(b) Ratio of LinTS/dTS cumulative regret in the last round with varying diffusion depth L in [2, 100]. Higher values mean a bigger performance gap.
dTS LinTS
2
0
20
40
60
80
100
Round t 2 [n]
(c) Regret of dTS in MovieLens. The diffusion model with L = 40 is pre-trained on embeddings obtained by lowrank factorization of MovieLens rating matrix.
Figure 4.3: (a) and (b): Impact of pre-training sample size and diffusion depth L for the Swiss roll data. (c): Regret of dTS in MovieLens. environment is not sampled from a diffusion model.
4.4.2
True Prior is Not a Diffusion Model
Swiss roll data. Unlike previous experiments, the true action parameters are now sampled from the Swiss roll distribution (see Figure B.1 in Section B.6.1), rather than from a diffusion model. The diffusion model used by dTS is pre-trained on samples from this distribution, with the offline pre-training procedure described in Section B.6.2. Figure 4.3a shows that larger sample sizes increase the performance gap between dTS and LinTS. More samples improve the estimation of the diffusion prior (see Figure B.1 in Section B.6.1), leading to better dTS performance. Notably, comparable performance was achieved with as few as 10 samples, and dTS outperformed LinTS by a factor of 1.5 with just 50 samples. While more samples may be required for more complex problems, LinTS would also struggle in such cases. Therefore, we expect these gains to be even more significant in more challenging settings. We studied the effect of the pre-trained diffusion model depth L and found that L ≈ 40 68
yields the best performance, with a drop beyond that point (Figure 4.3b). While our theory doesn’t apply directly here, as it assumes a linear diffusion model, it still offers some intuition on the decreased performance for L > 40. The theorem shows dTS’s regret bound increases with L when the true distribution is a diffusion model. For small L, the pre-trained model doesn’t fully capture the true distribution, making the theorem inapplicable, but at L ≈ 40, the distribution is nearly captured, and further increases in L lead to higher regret, consistent with our theory. MovieLens data. We also evaluate dTS using the standard MovieLens (Lam and Herlocker, 2016) setting. In this semi-synthetic experiment, a user is sampled from the rating matrix in each interaction round, and the reward is the rating the user gives to a movie (see Clavier et al. (2023, Section 5) for details about this setting). Here, the true distribution of action parameters is unknown and not a diffusion model. The diffusion model is pre-trained on offline estimates of action parameters obtained through low-rank factorization of the rating matrix. Figure 4.3c demonstrates that dTS outperforms LinTS in this setting. Additional CIFAR ablations are provided in Section B.6.4 where similar strong improvements are observed.
4.5
Conclusion
We use a pre-trained diffusion model as a strong and flexible prior for dTS. Diffusion pre-training leverages abundant offline data, which is then fine-tuned through online interactions via our tractable posterior approximation. This approximation enables efficient posterior sampling and updates while maintaining strong empirical performance. Moreover, dTS admits a simple Bayesian regret bound in the linear–Gaussian setting. Our work has several limitations. First, our Bayes regret analysis applies only to the linearGaussian setting with a well-specified prior; extending formal guarantees to nonlinear diffusion models remains open. Second, our posterior approximation, while motivated by exact solutions in the linear case, lacks theoretical justification for general nonlinear link functions: its strong empirical performance does not come with formal approximation error bounds. Finally, dTS requires offline pre-training of the diffusion model, which assumes access to historical estimates of action parameters; in domains where such data is unavailable or expensive to obtain, the benefits of diffusion priors may not be realized.
69
Part II Off-Policy Learning in Large Action Spaces
70
Chapter 5
Introduction to Part II
This second part of the thesis addresses the following fundamental question: How can we reliably learn high-performing policies from static logged data when the number of actions is large?
5.1
Setting and Background
In this part, we consider the off-policy (offline) setting where an agent is provided with a static logged dataset Dn = {(Xi , Ai , Ri )}ni=1 collected by a logging policy π0 . The data collection process proceeds as follows: for each round i ∈ [n]: 1. The environment draws a context Xi ∼ ν, where ν is a distribution with support X forming a compact subset of Rd ; 2. The logging policy selects an action Ai ∼ π0 (· | Xi ) from the action set A = [K]; 3. The environment generates a stochastic reward Ri ∼ p(· | Xi , Ai ), where Ri ∈ [0, 1]. Unlike the on-policy setting of Part I, no further interaction with the environment is permitted. The objective is to learn a new policy π̂ from this static dataset that maximizes the true (but unknown) expected value: [︁ ]︁ V (π) = EX∼ν EA∼π(·|X) [r(X, A)] , (5.1) where r(x, a) = ER∼p(·|x,a) [R] is the expected reward function. Performance is measured by the suboptimality gap of the learned policy: so(π̂) = V (π∗ ) − V (π̂),
(5.2)
where π∗ = arg maxπ∈Π V (π) is the optimal policy within a class Π. Since V (π) cannot be computed directly, off-policy learning algorithms often rely on an empirical estimate V̂ (π) constructed from Dn . This estimation task is known as off-policy evaluation (OPE) in the literature. The two dominant estimation approaches are: direct 71
method (DM) and inverse propensity scoring (IPS). DM employ a learned reward model r̂(x, a) to estimate the value as: n
1 ∑︂ ∑︂ π(a | Xi ) r̂(Xi , a) . V̂dm (π) = n i=1 a∈A
(5.3)
IPS re-weights observed rewards using importance sampling as: n
1 ∑︂ π(Ai | Xi ) V̂ips (π) = Ri . n i=1 π0 (Ai | Xi )
(5.4)
Given an estimator V̂ (π), the agent must then select a policy. This step distinguishes between greedy policies, which directly maximize π̂ = arg maxπ∈Π V̂ (π), and pessimistic policies, which incorporate an uncertainty penalty π̂ = arg maxπ∈Π [V̂ (π) − pen(π)].
5.1.1
Scalability Challenges
When the number of actions K is large, both estimation paradigms and their associated optimization procedures encounter fundamental obstacles: Statistical inefficiency of DM. Standard DM approaches model each action’s reward function independently. As K grows, the data available per action diminishes, leading to poorly estimated reward functions and lower performance. High variance of importance sampling. IPS’s variance grows with the importance weights π(a|x)/π0 (a|x). These weights explode in large action spaces, producing estimates too noisy for reliable optimization. Intractable optimization landscapes. Beyond estimation challenges, the optimization problem arg maxπ∈Π V̂ (π) itself becomes computationally intractable in large action spaces. IPS-based objectives induce highly non-concave landscapes with exponentially many local maxima and flat plateaus that trap gradient-based optimizers. As we show in Chapter 7, this optimization bottleneck often dominates estimation error, making even statistically superior estimators ineffective in practice.
5.2
Methodological Approaches
To address these challenges, the methods developed in this part pursue three complementary strategies: structured reward modeling for sample-efficient DM, surrogate objectives that prioritize optimization tractability over estimation accuracy, and principled regularization and pessimism for importance-weighted estimators. Structured Bayesian models. Drawing inspiration from the hierarchical framework of Part I, we introduce latent structure into reward modeling. Action parameters are coupled through shared latent variables ψ: ψ ∼ q(·), θa | ψ ∼ pa (·; fa (ψ)), R | X, A, θ, ψ ∼ p(· | X; θA ). 72
(5.5) ∀a ∈ A,
This formulation enables information sharing across actions: observations from frequently selected actions inform the posterior over ψ, which in turn improves reward estimates for rarely observed actions. Optimization-aware objectives. Rather than designing sophisticated value estimators and then optimizing them, we propose objectives designed primarily for favorable optimization landscapes. The policy-weighted log-likelihood (PWLL) family: n
Ûg (π) =
1 ∑︂ g(Ri , π0 (Ai | Xi )) log π(Ai | Xi ) , n i=1
(5.6)
where g is a positive weighting function, yields concave objectives for linear-softmax policies π. This guarantees efficient convergence to a unique global maximum, bypassing the optimization pathologies of value estimation altogether. Regularized importance weighting. For practitioners committed to IPS-based methods, we develop variance-controlled estimators through importance weight regularization. The exponential smoothing estimator: n
1 ∑︂ π(Ai | Xi ) Ri , V̂ (π) = n i=1 π0 (Ai | Xi )α α
α ∈ [0, 1] ,
(5.7)
smoothly trades variance for bias while preserving differentiability. Combined with pessimistic optimization and PAC-Bayes generalization bounds, this yields principled, tractable objectives that are amenable to stochastic gradient ascent for safe off-policy learning.
5.3
Roadmap of Part II
The following chapters develop these methodological approaches. Chapter 6: Scaling Direct Methods with Latent Parameters. We begin by addressing the statistical inefficiency of DM through structured Bayesian modeling. Building on the hierarchical framework of Part I, we introduce the structured direct method (sDM), which couples action parameters through a shared latent vector. The posterior over these latent variables aggregates evidence across all actions, enabling effective generalization to actions with sparse data coverage. We analyze performance through √ Bayesian suboptimality and prove that greedy policies paired with sDM achieve O(1/ n) convergence under mild assumptions on the alignment between logging and optimal policies. Chapter 7: Optimization Matters More Than Estimation. We then challenge the conventional paradigm of off-policy learning. Through theoretical analysis and largescale experiments, we demonstrate that optimization error dominates estimation error in large action spaces. Specifically, we prove that for any IPS-based estimator, gradient ascent can remain trapped in suboptimal regions for O(K) iterations, and that the optimization landscape contains exponentially many local maxima in K. We then propose objective-aware policy parametrizations: by aligning the policy class with the estimator’s inductive bias, we can partially mitigate these optimization challenges. However, for a 73
more complete solution, we propose policy-weighted log-likelihood (PWLL) objectives as an alternative to IPS-based objectives. These objectives are provably concave for linear softmax policies, guaranteeing efficient convergence to a global optimum. Experiments on datasets with up to one million actions validate that PWLL consistently outperforms state-of-the-art estimator-based methods. Chapter 8: Principled Pessimism for Exponential Smoothing and Beyond. Finally, for practitioners committed to importance weighting methods, we develop a theoretically grounded framework for variance control and pessimistic policy learning. We propose exponential smoothing estimators that regularize importance weights, trading controlled bias for reduced variance. To leverage these regularized estimators for safe policy learning, we derive two-sided PAC-Bayes generalization bounds where all quantities are empirical and differentiable. The pessimistic learning objective maximizes the lower bound on policy value, penalizing policies with high bias or variance and steering optimization toward reliable regions. This also yields tractable objectives amenable to stochastic gradient optimization. We further present a unified PAC-Bayes framework covering the major importance weight regularization techniques in the literature (clipping, exponential smoothing, implicit exploration), enabling principled comparison and demonstrating that the choice of pessimistic objective often matters more than the specific regularizer.
74
Chapter 6
Scaling Direct Methods with Latent Parameters
Contents Setting and Background . . . . . . . . . . . . . . . . . . . . . . . . . .
71
5.1.1
Scalability Challenges . . . . . . . . . . . . . . . . . . . . . . .
72
5.2
Methodological Approaches . . . . . . . . . . . . . . . . . . . . . . . .
72
5.3
Roadmap of Part II . . . . . . . . . . . . . . . . . . . . . . . . . . . .
73
5.1
In this chapter, we address the statistical inefficiency of direct methods (DMs) in large action spaces: the first challenge highlighted in Chapter 5. Standard DMs estimate independent d-dimensional parameters for each action. This approach becomes statistically inefficient in large action spaces, where data are collected by a logging policy that explores only a small subset of available actions, leaving many actions rarely or never observed. To overcome this limitation, we analyze DMs through a Bayesian lens and propose making them sample-efficient by incorporating informative priors. We make the following contributions. 1) We introduce the structured direct method (sDM), a Bayesian approach that uses informative priors to share reward information across actions. By updating beliefs about similar actions based on observed data, sDM improves statistical efficiency without compromising computational scalability. 2) To evaluate sDM, we propose Bayesian metrics that assess the average performance across problem instances sampled from the prior. This departs from the standard frequentist focus on worst-case scenarios. These metrics formally quantify the benefits of informative priors. 3) Our theoretical analysis of Bayesian suboptimality (BSO) reveals two key insights: (1) performance degrades gracefully even without the standard assumption of full logging support, and (2) greedy policies are provably optimal under the BSO metric, standing in contrast to the pessimistic policies typically favored in frequentist settings. 4) We empirically validate sDM and our theoretical findings using both synthetic and real-world datasets. 75
6.1
Setting
We consider the setting described in Section 5.1. The only additional assumption is the existence of unknown true parameters θ∗,a ∈ Rd for each action a, such that rewards are distributed as Ri ∼ p(· | Xi ; θ∗,Ai ). Let θ∗ = (θ∗,a )a∈A ∈ RdK denote the concatenation of all action parameters. The reward function r(x, a; θ∗ ) = ER∼p(·|x;θ∗,a ) [R] gives the expected reward of action a in context x. The goal is to find a policy π ∈ Π that maximizes: V (π; θ∗ ) = EX∼ν EA∼π(·|X) [r(X, A; θ∗ )] . This chapter focuses on DM that estimates the value V (π; θ∗ ) as: V̂dm (π) =
1 ∑︂ ∑︂ π(a | Xi )r̂(Xi , a) , n a∈A
(6.1)
i∈[n]
where r̂(x, a) is an estimation of r(x, a; θ∗ ). DM estimators may exhibit modeling bias, but they generally have lower variance than IPS (Saito and Joachims, 2022). Another advantage of DM is its practical utility without assuming access to the logging policy π0 (Jeunen and Goethals, 2021; Aouali et al., 2022c; Hong et al., 2023). Also, DMs can be incorporated into a Bayesian framework, where informative priors can be used to enhance statistical efficiency. This allows for the development of scalable methods suitable for large action spaces, as shown in our work.
6.2
Structured DM
6.2.1
Structured Priors
Pitfalls of non-structured priors. Before presenting sDM, we first describe the pitfalls of using the following widely used standard prior, θa ∼ N (µa , Σa ) , ⊤
∀a ∈ A ,
(6.2)
2
R | θ, X, A ∼ N (ϕ(X) θA , σ ) , where ϕ(x) provides a d-dimensional representation of the context x ∈ X , and N (µa , Σa ) represents the prior density of θa , with σ 2 being the reward noise variance. Under this prior, each action a has an associated parameter θa . Given the prior in Equation (6.2), the posterior distribution of an action parameter follows a multivariate Gaussian: θa | Dn ∼ N (µ̂a , Σ̂a ), where −1 Σ̂−1 a = Σa + Ga ,
−1 Σ̂−1 a µ̂a = Σa µa + Ba .
with Ga = σ −2
∑︂
1{Ai =a} ϕ(Xi )ϕ(Xi )⊤ ,
Ba = σ −2
i∈[n]
∑︂
1{Ai =a} Ri ϕ(Xi )
i∈[n]
Note that Ga and Ba only use the subset of samples Dn where action a was observed, meaning data from other actions b ̸= a do not contribute to the posterior inference for 76
action a. This results in statistical inefficiency, especially if the logged data Dn doesn’t cover all actions. In particular, the posterior for an unseen action a, θa | Dn , would simply revert to the prior N (µa , Σa ), since we would have Ga = 0d×d and Ba = 0d in such case. Structured priors. To address the above issue, we assume that action rewards correlate and embed this knowledge into the prior. While one could model these correlations by considering the joint posterior distribution of (θa )a∈A | Dn , this becomes computationally burdensome when the number of actions K is large. Instead, we introduce an unknown d′ ′ dimensional latent parameter ψ ∈ Rd , sampled from a latent prior q(·), such as ψ ∼ q(·). The correlations between actions naturally arise because each action parameter θa is derived from the same latent parameter ψ. Specifically, the action parameters θa are conditionally independent given ψ and are sampled from a conditional prior pa as θa | ψ ∼ pa (·; fa (ψ)) for all a ∈ A. Here, pa is ′ parameterized by fa (ψ), where fa : Rd → Rd is a known prior function that encodes the hierarchical relationship between action parameters θa and the latent parameter ψ. This structure allows for sparsity, meaning that θa may depend only on a subset of ψ’s coordinates. Moreover, pa accounts for model uncertainty, allowing for cases where θa is not a deterministic function of ψ, i.e., θa ̸= fa (ψ). The reward distribution for action a in context x is given by p(· | x; θa ), which depends only on x and θa . To summarize, the structured prior is defined below, and its graphical representation is given in Figure 6.1. ψ ∼ q(·) , θa | ψ ∼ pa (·; fa (ψ)) , R | ψ, θ, X, A ∼ p(· | X; θA ) .
(6.3) ∀a ∈ A ,
To derive the posterior under this prior, we assume that: (i) (X, A) is independent of ψ, and given ψ, (X, A) is independent of θ; and (ii) given ψ, the parameters θa for all a ∈ A are independent.
: taken action
Figure 6.1: Graph representation of the structured prior. Now, we discuss how to perform off-policy learning under this general structured prior in Equation (6.3), before applying it to linear-Gaussian distributions in Section 6.3. 77
6.2.2
Off-Policy Learning
Off-policy learning relies on an estimate of the value function V (π; θ∗ ) obtained using the logged data Dn . In DMs, the estimator V̂dm in Equation (6.1) requires access to the learned reward r̂(x, a) ≈ r(x, a; θ∗ ). In our Bayesian setting, this requires access to the action posterior θa | Dn under the prior in Equation (6.3) since the reward is then estimated as r̂(x, a) = E [r(x, a; θ) | Dn ] for any (x, a) ∈ X × A, and this estimate is plugged into V̂dm in Equation (6.1) to estimate V (π; θ∗ ). Thus, we need to derive the posterior density of the action parameter θa , p(θa | Dn ), under the structured prior in Equation (6.3), which reads ∫︂ p(θa | Dn ) = p(θa | ψ, Dn )p(ψ | Dn ) dψ , (6.4) ψ
where ψ | Dn is the latent posterior and θa | ψ, Dn is the conditional action posterior. To compute p(θa | Dn ), we first compute p(θa | ψ, Dn ) and p(ψ | Dn ) and then integrate out ψ following Equation (6.4). First, p(θa | ψ, Dn ) ∝ La (θa )pa (θa ; fa (ψ)) ,
(6.5)
∏︁ with La (θa ) = (X,A,R)∈Sa p(R|X; θa ) is the likelihood of observations of action a (Sa = (Xi , Ai , Ri )i∈[n],Ai =a is the subset of Dn where Ai = a). Similarly, ∏︂ ∫︂ Lb (θb )pb (θb ; fb (ψ)) dθb q(ψ) , (6.6) p(ψ | Dn ) ∝ b∈A
θb
This allows us to further develop Equation (6.4) as ∫︂ ∏︂ ∫︂ Lb (θb )pb (θb ; fb (ψ)) dθb q(ψ) dψ . p(θa | Dn ) ∝ La (θa )pa (θa ; fa (ψ)) ψ
b∈A
(6.7)
θb
All the quantities inside the integrals in Equation (6.7) are given (the parameters of pa and q) or tractable (the terms in La ). Thus, if these integrals can be computed, then the posterior can be fully characterized in closed form, which we will do in Section 6.3 in the fully linear case. Otherwise, the posterior should be approximated. Finally, we act greedy with respect to our estimator V̂dm and define the learned policy as the one maximizing it: π̂g = argmaxπ∈Π V̂dm (π). If the set of policies Π contains deterministic policies, then π̂g (a | x) = 1{a = argmax r̂(x, b)} .
(6.8)
b∈A
In particular, we do not adopt the common pessimism approach (Jin et al., 2021). In pessimism, one constructs confidence intervals of the reward estimate r̂(x, a) of the form |r(x, a; θ) − r̂(x, a)| ≤ u(x, a), and then defines the learned policy as π̂p (a | x) = 1{a = argmaxb∈A r̂(x, b) − u(x, b)}. The advantage of one over another depends on the evaluation metric used. Our metric is the Bayesian suboptimality (BSO), defined in Section 6.4. It assesses the average performance of algorithms across multiple problems rather than the worst-case. The Greedy policy is more suitable for BSO optimization than pessimism (demonstrated theoretically and empirically in Sections C.3.3 and C.4.4). 78
6.3
Linear-Gaussian Case
In this section, we use linear functions fa combined with Gaussian distributions for the structured prior Equation (6.3). Precisely, we assume that the latent prior q(·) = ′ ′ ′ N (·; µ, Σ) is Gaussian with mean µ ∈ Rd and covariance Σ ∈ Rd ×d . Moreover, let ′ ′ Wa ∈ Rd×d be the mixing matrix for action a, we define fa (v) = Wa v for any v ∈ Rd . We define the conditional prior pa (·; fa (ψ)) = N (·; Wa ψ, Σa ) is Gaussian with mean fa (ψ) = Wa ψ ∈ Rd and covariance Σa ∈ Rd×d . The reward distribution p(· | x; θa ) is also linear-Gaussian as N (·; ϕ(x)⊤ θa , σ 2 ), where ϕ(·) outputs a d-dimensional representation of x and σ > 0 is the observation noise variance. The whole prior is (6.9)
ψ ∼ N (µ, Σ) , (︂ )︂ θa | ψ ∼ N Wa ψ, Σa ,
∀a ∈ A ,
R | ψ, θ, X, A ∼ N (ϕ(X)⊤ θA , σ 2 ) .
6.3.1
Applications
Mixed-effect modeling. Equation (6.9) allows modeling that action parameters depend on a linear mixture of effect parameters (Chapter 3). Precisely, let J be the number of effects and assume that d′ = dJ so that the latent parameter ψ is the concatenation of J, d-dimensional effect parameters, ψj ∈ Rd , such as ψ = (ψj )j∈[J] ∈ RdJ . Moreover, assume Rd×dJ where wa = (wa,j )j∈[J] ∈ RJ are the mixing that for any a ∈ A , Wa = wa⊤ ⊗ Id ∈∑︁ weights of action a. Then, Wa ψ = j∈[J] wa,j ψj for any a ∈ A. Sparsity, i.e., when an action a only depends on a subset of effects, is captured through the mixing weights wa : wa,j = 0 when action a is independent of the j-th effect parameter ψj and wa,j ̸= 0 otherwise. Also, the level of dependence between action a and effect j is quantified by the absolute value of wa,j . This mixed-effect model can be used in numerous applications (the reader can refer to the first paragraphs of Chapter 3 for examples). Low-rank modeling. Equation (6.9) can also model the case where the dimension of the latent parameter ψ is much smaller than that of the action parameters θa , i.e., when d′ ≪ d. Again, this is captured through the mixing matrices Wa , when Wa is low-rank.
6.3.2
Closed-Form Solutions for sDM
The conditional action posterior is known in closed-form as θa | ψ, Dn ∼ N (µ̃a , Σ̃a ), with −1 Σ̃−1 a µ̃a = Σa Wa ψ + Ba ,
−1 Σ̃−1 a = Σa + Ga ,
(6.10)
where Ga = σ −2
∑︂
1{Ai = a}ϕ(Xi )ϕ(Xi )⊤ ,
i∈[n]
Ba = σ −2
∑︂ i∈[n]
79
1{Ai = a}Ri ϕ(Xi ) .
This posterior has the standard form except that the prior mean Wa ψ now depends on the latent parameter ψ. Similarly, the effect posterior writes ψ | Dn ∼ N (µ̄, Σ̄), where ∑︂ −1 −1 Σ̄−1 = Σ−1 + Wa⊤ (Σ−1 a − Σa Σ̃a Σa )Wa , a∈A
Σ̄−1 µ̄ = Σ−1 µ +
∑︂
(6.11)
Wa⊤ Σ−1 a Σ̃a Ba .
a∈A
The latent posterior precision Σ̄−1 is the sum of the latent prior precision Σ−1 and the −1 −1 ⊤ learned action precisions Σ−1 a − Σa Σ̃a Σa , weighted by Wa Wa . The contribution of each action’s learned precision to the latent precision is proportional to Wa⊤ Wa . This intuition similarly applies to interpreting µ̄. Finally, from Equation (6.7), the action posterior is θa | Dn ∼ N (µ̂a , Σ̂a ), where (︁ )︁ ⊤ −1 Σ̂a = Σ̃a + Σ̃a Σ−1 µ̂a = Σ̃a Σ−1 (6.12) a Wa Σ̄Wa Σa Σ̃a , a Wa µ̄ + Ba . Finally, from Equation (6.9), the reward function is r(x, a; θ) = ϕ(x)⊤ θa . Thus, the estimated reward is r̂(x, a) = E[r(x, a; θ) | Dn ] = ϕ(x)⊤ µ̂a ,
∀(x, a) ∈ X × A .
This can then be plugged in Equation (6.8) for decision-making, leading to π̂g (a | x) = 1{a = argmax ϕ(x)⊤ µ̂b } . b∈A
To see why this is more beneficial than the standard prior in Equation (6.2), notice that the mean and covariance of the posterior of action a, µ̂a and Σ̂a , are now computed using the mean and covariance of the latent posterior, µ̄ and Σ̄. But µ̄ and Σ̄ are learned using the interactions with all the actions in Dn . Thus µ̂a and Σ̂a are also learned using the interactions with all the actions in Dn , in contrast with the standard prior in Equation (6.2) where they were learned using only the interaction with action a. The additional computational cost of considering the structured prior in Equation (6.9) is small. The computational and space complexities are O(K((d2 + d′ 2 )(d + d′ ))) and O(Kd2 ). For example, when d′ = O(d), these complexities become O(Kd3 ) and O(Kd2 ), respectively. This is exactly the cost of the standard prior in Equation (6.2). In contrast, this strictly improves the computational efficiency of jointly modeling the action parameters, where the complexities are O(K 3 d3 ) and O(K 2 d2 ) since the joint posterior of (θa )a∈A | Dn requires converting and storing a dK × dK covariance matrix. Remark 6. sDM with linear-Gaussian hierarchies can be used even with data generated from non-linear rewards, and we empirically investigate its robustness to misspecification. We found that this model performs well even if the true rewards are not generated from a linear-Gaussian distribution.
6.4
Analysis
6.4.1
Bayesian Metrics
The performance of a learned policy π̂ is evaluated using suboptimality (SO): so(π̂; θ∗ ) = V (π∗ ; θ∗ ) − V (π̂; θ∗ ) , 80
where π∗ = argmaxπ∈Π V (π; θ∗ ) is the optimal policy. This metric is well-suited when the environment is governed by a unique, fixed ground truth θ∗ . It applies to any policy π̂, whether learned through frequentist approaches (e.g., MLE) or Bayesian ones (e.g., ours). However, when the environment is modeled as a random variable θ∗ sampled from some prior distribution, SO becomes less appropriate. Thus, drawing on recent developments in Bayesian analysis for online bandits through Bayes regret (Russo and Van Roy, 2014), we introduce a new metric for offline settings, termed Bayes suboptimality, defined as: Bso(π̂) = E[V (π∗ ; θ∗ ) − V (π̂; θ∗ )] ,
(6.13)
where the expectation is taken over all random variables: the logged data Dn and θ∗ , which is treated as a random variable sampled from the prior. The BSO can be computed in two ways. One method involves taking the expectation under the prior θ∗ , followed by taking an expectation under data generated from a fixed environment θ∗ as Dn | θ∗ . The other method involves taking an expectation under the data Dn , followed by taking an expectation under the posterior θ∗ | Dn . The BSO is a reasonable metric for assessing the average performance of algorithms across multiple environments, due to the expectation over θ∗ . It is also known that Bayes regret captures the benefits of using informative priors (Chapter 3), and this is similarly achieved by the BSO.
6.4.2
Theoretical Results
Our theory relies on the important well-specified assumption: Assumption 1 (Well-specified priors). Action parameters θ∗,a and rewards are drawn from Equation (6.9). We also make simplifying assumptions for the sake of exposition. Assumption 2 (Diagonal covariances for simplicity). We assume Σa = σ02 Id , Σ = τ 2 Id′ , ∥ϕ(x)∥2 ≤ 1, and the matrices Wa are normalized such that λ1 (Wa Wa⊤ ) = λd (Wa Wa⊤ ) = 1. This yields our bound on the BSO of sDM. Theorem 2 (Covariance-Dependent Bound). Let π∗ (x) be the optimal action for context x. Then the BSO of sDM under the structured prior in Equation (6.9) satisfies √︃ [︂ ]︂ (2 log(2K) + 2)(σ02 + τ 2 ) , (6.14) Bso(π̂g ) ≤ αn E ∥ϕ(X)∥Σ̂π (X) + ∗ n √︂ √︁ where αn = d + 2 d log(Kn) + 2 log(Kn). Scaling of the bound in Theorem 2 aligns with existing frequentist results (Jin et al., 2021, Theorem 4.4). The main differences lie in the constants and the fact that this rate is achieved using greedy policies in Equation (6.8). This contrasts with the frequentist setting where pessimism is used (Jin et al., 2021) and known to be optimal (Jin et al., 2021, 81
Theorem 4.7). In fact, greedy policies are optimal when BSO is used as the performance metric. Specifically, Bso(π̂g ) ≤ Bso(π) for any policy π, including pessimistic ones. Therefore, in the Bayesian setting and when BSO is used as a performance metric, greedy policies should always be preferred to pessimistic ones. This fundamental difference is proven in Section C.3.3 and it is of independent interest beyond this work. Theorem 2 suggests that the BSO primarily depends on the posterior covariance of action π∗ (X) in the direction of the context ϕ(X). That is, when the uncertainty in the posterior distribution of the optimal action π∗ (X) is low on average across different contexts X and logged data Dn , then the BSO bound is correspondingly small. In particular, the tightness of the bound depends on the degree to which the logged data covers the optimal actions on average. Theorem 2 can highlight the advantages of using sDM over the non-structured prior in Equation (6.2). To see this, notice that the parameters of the non-structured prior in Equation (6.2), µa and Σa , are obtained by marginalizing out ψ in Equation (6.9). In this ⊤ ns case, µns a ← Wa µ and Σa ← Σa + Wa ΣWa . The corresponding posterior covariance is ⊤ −1 Σ̂ns + Ga )−1 , and is generally larger than the covariance of sDM, Σ̂a a = ((Σa + Wa ΣWa ) in Equation (6.12). This is more pronounced when the number of actions K is large and when the latent parameters are more uncertain than the action parameters. Thus, the BSO bound of sDM is smaller due to the reduced posterior uncertainty it exhibits. Also, note that even when π∗ (X) is unobserved in the logged data Dn , sDM’s posterior covariance Σ̂π∗ (X) can remain small since we use interactions with all actions to compute it. This contrasts with standard non-structured priors in Equation (6.2), where observing π∗ (X) is necessary; without such observations, the posterior covariance Σ̂π∗ (X) would simply be the prior covariance Σπ∗ (X) . √ Next, we provide another bound on the BSO that scales as O(1/ n). To simplify the exposition, we roughly present its scaling with n in Theorem 3 and defer the complete general statement to Section C.3.2. We make the following additional assumptions: Assumption 3. Let G = EX∼ν [XX ⊤ ] with g = λd (G). We assume that g > 0. Assumption 4 (Context-independent logging policy). A is independent of X, i.e., π0 (a | x) = π0 (a) = pa for all x and a. Equivalently, (Xi ) are i.i.d. ∼ ν and independent of (Ai ), with P(A = a) = pa . Theorem 3 (Scaling with n). For n large enough, the BSO of sDM under the structured prior in Equation (6.9) scales as (︄√︄ [︃ E Bso(π̂g ) = Õ
d nρX + 1
√︃
]︃ +
log K n
)︄ ,
where ρX = π0 (π∗ (X)). The above bound becomes smaller or larger depending on how well the logging policy π0 covers the optimal actions for each context x. 82
6.5
Experiments
We evaluate sDM using both synthetic and real datasets. We use the average reward of the learned policy relative to the optimal policy as the evaluation metric.
6.5.1
Synthetic Problems
Setting. We simulate synthetic data using the linear-Gaussian model in Equation (6.9) with σ = 1. The contexts X are sampled uniformly from [−1, 1]d , with d = 10. The ′ matrices Wa are sampled uniformly from [−1, 1]d×d , where we very d′ as d′ ∈ {5, 10., 20}. We set Σ = 3Id′ and Σa = Id , meaning the latent parameters are more uncertain than the ′ action parameters. The latent mean µ is randomly sampled from [−1, 1]d . The number of actions is varied as K ∈ {100, 1000}, and we use a uniform logging policy to collect data. Additional experiments with different logging policies are presented in Section C.4. Baselines. First, we use sDM under prior in Equation (6.9). Second, we examine DM (Bayes), which uses the standard non-structured prior in Equation (6.2), where parameters µa and Σa are obtained by marginalizing out the latent parameters ψ in Equation (6.9). Thus DM (Bayes) is a standard Bayesian DM that does not capture arm reward correlations. We also include DM (Freq), which estimates θ∗,a by the MLE. We include IPS (Horvitz and Thompson, 1952), self-normalized IPS (snIPS) (Swaminathan and Joachims, 2015b), and doubly robust (DR) (Dudik et al., 2014), which we optimize to learn the optimal policy. MIPS (Saito and Joachims, 2022) and PC (Sachdeva et al., 2024) are also included. Implementation details of baselines is provided in Section C.4.1.
Avg. relative reward
Results. In Figure 6.2, we plot the results and we observe that sDM consistently outperforms the baselines across all settings. This performance gap becomes even more significant when sample size n is small. These results highlight sDM’s enhanced efficiency in using available logged data, making it particularly beneficial in data-limited situations and scalable to large action spaces.
1.0 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1
Synthetic - OPL K=100, d'=5, d=10
0
500
1000 1500 20000
Number of samples n sDM (Ours)
DM (Bayes)
Synthetic - OPL K=1000, d'=5, d=10
500
Synthetic - OPL K=1000, d'=10, d=10
1000 1500 20000
Number of samples n DM (Freq)
500
1000 1500 20000
Number of samples n DR
Synthetic - OPL K=1000, d'=20, d=10
IPS
500
1000 1500 2000
Number of samples n snIPS
MIPS
PC
Figure 6.2: The average relative reward of the learned policy using one of the baselines on synthetic problems with varying n, K and d′ . Scaling to large action spaces. sDM achieves improved scalability compared to standard DM as it leverages data more efficiently. While it still learns a d-dim. parameter for each action a, it does so by considering interactions with all actions in the logged 83
Avg. relative reward
data Dn , instead of only using interactions with the specific action a. This is crucial, especially given that many actions may not even be observed in Dn . To show sDM’s improved scalability, we compare it to the most competitive baseline, DM (Bayes), for varying K ∈ [10, 100000] with n = 1000. The results in Figure 6.3 reveal that the performance gap between sDM and DM (Bayes) becomes more significant when the number of actions K increases. Hence, despite the necessity for sDM to learn distinct parameters for each action, accommodating practical scenarios like recommender systems where unique embeddings are learned for each product, it still enjoys good scalability. 1.0
Varying K, d'=5, d=10
0.9 0.8 0.7 sDM (Ours) DM (Bayes)
0.6 0.5 10
500
5000 100000
Number of actions K
Figure 6.3: sDM vs. DM (Bayes) for varying K.
6.5.2
MovieLens Problems
Setting. We use MovieLens 1M (Lam and Herlocker, 2016), which contains 1 million ratings representing the interactions between 6,040 users and 3,952 movies. To create a semi-synthetic environment, we first apply a low-rank factorization to the rating matrix, producing 5-dim. representations: xu ∈ R5 for user u ∈ [6040] and θa ∈ R5 for movie a ∈ [3952]. Movies are treated as actions, and contexts X are sampled randomly from the user vectors. The reward for movie a and user u is modeled as N (x⊤ u θa , 1), serving as proxy for ratings. A uniform logging policy is used to collect data. Baselines. We consider the same baselines as in synthetic data. A prior is not needed for DM (Freq), IPS, snIPS, and DR. However, for DM (Bayes), a standard prior in Equation (6.2) is inferred from data, where we set µa to be the mean of movie vectors across all dimensions, and Σa = diag(v), where v represents the variance of movie vectors across all dimensions. Unlike the synthetic experiments, the latent structure assumed by sDM is not inherently present in MovieLens. But we learn it by training a Gaussian Mixture Model (GMM) to cluster movies into J = 5 mixture components. This gives rise to the mixed-effect structure described in Section 6.3, which represents a specific instance of sDM with d′ = dJ = 25. MIPS also has access to movie clusters, while we use the knn smoothing implementation of PC (see (Sachdeva et al., 2024, Section 3)). Note that DM (Bayes), sDM, MIPS and PC use the same subset of data (of size 1000) to learn their priors/assumed structure and thus we compare them fairly. We conduct experiments with K ∈ {100, 1000} randomly selected movies. Results. Results are in Section 6.5.2. Even though the latent structure assumed by sDM is not inherently present in MovieLens, sDM still outperforms the baselines by learning it 84
Avg. relative reward
offline.
0.95 0.90 0.85 0.80 0.75 0.70 0.65 0.60
MovieLens - OPL K=100, d'=25, d=5
0
500 1000 1500 2000
Number of samples n sDM (Ours) DM (Bayes) DM (Freq)
MovieLens - OPL K=1000, d'=25, d=5
0
500 1000 1500 2000
Number of samples n DR IPS snIPS
MIPS PC
Figure 6.4: The average relative reward of the learned policy using one of the baselines on MovieLens problems with varying n, K and d′ .
6.6
Conclusion
We introduced sDM, a structured approach to off-policy learning that leverages latent structure among actions to enhance statistical efficiency while maintaining computational tractability, particularly in large action spaces with limited data coverage. Within a Bayesian framework, we proved that greedy policies outperform pessimistic ones under √ Bayesian suboptimality and established O(1/ n) convergence without requiring restrictive full-support assumptions. Our work has several limitations. First, our theoretical analysis assumes a well-specified prior; while we empirically observed robustness to misspecification, formal guarantees under prior mismatch remain an open question. Second, closed-form posterior updates are available only for linear-Gaussian hierarchies; extending to nonlinear reward models requires approximate inference, which may compromise computational efficiency or statistical accuracy. Third, the latent structure must be specified or learned pre-trained, adding an additional modeling task. Finally, extending sDM to handle nonlinear hierarchies, building on the diffusion-based approach of Chapter 4, is a promising avenue for future work.
85
Chapter 7
Optimization Matters More than Esimation
Contents 6.1
Setting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
76
6.2
Structured DM . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
76
6.2.1
Structured Priors . . . . . . . . . . . . . . . . . . . . . . . . .
76
6.2.2
Off-Policy Learning . . . . . . . . . . . . . . . . . . . . . . . .
78
Linear-Gaussian Case . . . . . . . . . . . . . . . . . . . . . . . . . . .
79
6.3.1
Applications . . . . . . . . . . . . . . . . . . . . . . . . . . . .
79
6.3.2
Closed-Form Solutions for sDM . . . . . . . . . . . . . . . . . .
79
Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
80
6.4.1
Bayesian Metrics . . . . . . . . . . . . . . . . . . . . . . . . .
80
6.4.2
Theoretical Results . . . . . . . . . . . . . . . . . . . . . . . .
81
Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
83
6.5.1
Synthetic Problems . . . . . . . . . . . . . . . . . . . . . . . .
83
6.5.2
MovieLens Problems . . . . . . . . . . . . . . . . . . . . . . .
84
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
85
6.3
6.4
6.5
6.6
This chapter challenges the dominant paradigm in off-policy learning (explored in Chapter 8), which frames the problem as finding a policy π̂ = argmaxπ V̂ (π) (or, with pessimism, π̂ = argmaxπ [V̂ (π) − pen(π)]), where V̂ is an IPS-based1 estimate of the true policy value V (π). The rationale behind these objectives is that maximizing a more accurate value estimate yields a better policy. However, this estimator-centric view neglects a crucial factor: the optimization landscape. 1
Recall that IPS is an importance-weighting estimator of the policy value. We use IPS-based to refer to any estimator derived from or inspired by importance weighting.
86
IPS-based objectives (Dudík et al., 2011; Dudík et al., 2012; Dudik et al., 2014; Wang et al., 2017; Farajtabar et al., 2018; Su et al., 2020; Metelli et al., 2021; Kuzborskij et al., 2021; Saito and Joachims, 2022) are highly non-concave under common policy parameterizations (Chen et al., 2019), prone to suboptimal local maxima and plateaus: issues that are exacerbated in large action spaces. Even sophisticated estimators designed to reduce variance fail to overcome this optimization barrier, as they often induce equally difficult landscapes. We make the following contributions. 1) We show that objective-aware policy parametrization can partially alleviate these difficulties by structuring the policy class to match the implicit biases of the estimator. Such parametrizations reduce the effective search space and can shorten optimization plateaus and local maxima. However, this strategy does not eliminate the fundamental non-concavity of IPS-based objectives, leaving optimization as the central bottleneck. 2) Motivated by this limitation, we advocate for an alternative approach based on policy-weighted log-likelihood (PWLL) objectives. Unlike traditional estimators, PWLL optimizes an objective Û (π) designed for ease of optimization rather than accuracy in estimating V (π). Although PWLL objectives perform poorly as value estimators, their favorable concave landscape makes them significantly more effective for policy learning. 3) Through theoretical and empirical analysis, we demonstrate that this optimization-centric approach consistently enables simpler PWLL objectives to outperform complex, state-of-the-art IPS-based methods, particularly in large action spaces. Setting and organization. This chapter considers the general setting of Section 5.1 with R ∈ [0, 1]. The remainder is organized as follows. Section 7.1 employs an asymptotic lens to analyze IPS-based objectives and derives objective-aware policy parametrizations that partially alleviate their optimization challenges. Section 7.2 introduces PWLL objectives and establishes their favorable optimization properties. Section 7.3 presents large-scale experiments. We conclude in Section 7.4.
7.1
Analysis of IPS-Based Objectives
IPS-based objectives optimize an estimator V̂ (π) of the policy value V (π). To understand the policies to which these estimators converge, we study their oracle policies π∗method = argmaxπ E[V̂ method (π)]. Taking the expectation removes sampling fluctuations and isolates the inductive bias of each objective: different estimators yield different oracle policies, even with infinite data. Crucially, oracle policies admit closed-form expressions, enabling precise characterization of each estimator’s implicit bias. This analysis motivates objective-aware parametrizations that align the policy class with the estimator’s bias to ease optimization: the first improvement we propose in this chapter.
7.1.1
Standard IPS-Based Objectives
The foundational IPS estimator (Horvitz and Thompson, 1952) re-weights observed rewards by the ratio between the target policy π and the logging policy π0 : n 1 ∑︂ π(Ai | Xi ) Ri . (7.1) V̂ips (π) = n i=1 π0 (Ai | Xi ) 87
In expectation, IPS selects the best-rewarding action among those in the support of π0 : [︃ ]︃ IPS ′ ′ π∗ (a | x) = 1 a = argmax r(x, a )1[π0 (a | x) > 0] . (7.2) a′ ∈A
Clipped IPS (cIPS). To mitigate the high variance of IPS, a widely used variant is cIPS (Bottou et al., 2013) that clips small propensity scores at a threshold τ ∈ (0, 1): n
π(Ai | Xi ) 1 ∑︂ Ri . V̂cips (π) = n i=1 max{π0 (Ai | Xi ), τ }
(7.3)
This clipping introduces a bias. The oracle policy down-weights the rewards of rare actions, causing it to favor actions that were frequent under π0 , even if they are suboptimal: [︂ π∗cIPS (a | x) = 1 a = argmax a′ ∈A
]︂ π0 (a′ | x) ′ r(x, a ) . max{π0 (a′ | x), τ }
(7.4)
Exponential smoothing (ES). Instead of hard clipping, ES (Aouali et al. (2023a), Chapter 8) smooths importance weights by raising propensities to a fractional power α ∈ (0, 1): n
1 ∑︂ π(Ai | Xi ) Ri . V̂es (π) = n i=1 π0 (Ai | Xi )α
(7.5)
Its oracle policy balances reward maximization with preference for frequent actions: [︃ ]︃ ES ′ ′ 1−α π∗ (a | x) = 1 a = argmax r(x, a )π0 (a | x) . (7.6) a′ ∈A
Another variant of ES regularizes the entire importance weight as ( ππ0 )β instead of only the denominator. In contrast to the deterministic policies derived from IPS, cIPS, and the ES formulation above, this approach yields a stochastic oracle policy: π∗ES (a | x) ∝ r(x, a)1/(1−β) π0 (a | x). Other regularizations include logarithmic smoothing (Sakhi et al., 2024), implicit exploration (Gabbianelli et al., 2024), harmonic correction (Metelli et al., 2021), shrinkage (Su et al., 2020). But we do not include as ES and cIPS are already representative of them. Doubly robust (DR). The DR estimator incorporates a reward model r̂(x, a) to reduce variance and enable generalization to actions outside π0 ’s support. A common clipped variant is: n
1 ∑︂ π(Ai | Xi ) (Ri − r̂(Xi , Ai )) + EA∼π(·|Xi ) [r̂(Xi , A)] . V̂dr (π) = n i=1 max{π0 (Ai | Xi ), τ }
(7.7)
Its oracle policy interpolates between the reward model prediction and an importance weighting correction for the reward model error: [︂ π∗DR (a | x) = 1 a = argmax r̂(x, a′ ) + a′ ∈A
]︂ π0 (a′ | x) ′ ′ (r(x, a ) − r̂(x, a )) . max{π0 (a′ | x), τ } 88
(7.8)
7.1.2
Large-Scale IPS-Based Objectives
In large action spaces, importance weights ππ(a|x) can become huge, leading to estimators 0 (a|x) with high variance. To mitigate this, modern methods compute marginalized importance weights over a lower-dimensional action representation, trading bias for reduced variance. Marginalized IPS (MIPS). MIPS (Saito and Joachims, 2022) tackles large action spaces by clustering actions. It maps each action a to a cluster c via a function h : A → C, where | C |≪| A |. Estimation is then performed at the cluster level: n
1 ∑︂ π(Ci | Xi ) V̂mips (π) = Ri , n i=1 π0 (Ci | Xi )
where Ci = h(Ai ) and π(c | x) =
∑︂
π(a | x).
(7.9)
a∈c
This cluster-level marginalization introduces bias: the oracle policy only selects the best cluster based on its average reward under π0 , and cannot differentiate between actions within that cluster: }︃]︃ [︃ {︃ ∑︁ a∈c′ π0 (a | x)r(x, a) MIPS ∑︁ . (7.10) π∗ (c | x) = I c = argmax c′ ∈C a∈c′ π0 (a | x) Hence, MIPS offers no specific guidance for selecting an action within the optimal cluster; any action is considered equally valid. Consequently, one possible induced action-level oracle under uniform tie-breaking is: {︂ ∑︁ [︂ }︂]︂ ′ π (a|x)r(x,a) a∈c ∑︁ 0 I h(a) = argmaxc′ ∈C a∈c′ π0 (a|x) π∗MIPS (a | x) = . |h(a)| where |h(a)| denotes the size of the cluster containing action a. Conjunct effect modeling (OffCEM). Building on MIPS, OffCEM (Saito et al., 2023) uses a reward model r̂ to correct for the cluster-level aggregation bias, in a doubly robust fashion: )︃ n (︃ 1 ∑︂ π(Ci | Xi ) V̂offcem (π) = (Ri − r̂(Xi , Ai )) + EA∼π(·|Xi ) [r̂(Xi , A)] . (7.11) n i=1 π0 (Ci | Xi ) The resulting oracle policy selects the action that maximizes the model-predicted reward r̂, plus a cluster-level correction term that accounts for model error: [︄ {︄ }︄]︄ ∑︁ π (ā | x)(r(x, ā) − r̂(x, ā)) ′ 0 ā∈h(a ) ∑︁ π∗OffCEM (a | x) = I a = argmax r̂(x, a′ ) + . ′ a ∈A ā∈h(a′ ) π0 (ā | x) (7.12) Two-stage decomposition (POTEC). In this chapter, we see POTEC (Saito et al., 2025) as an optimization strategy of OffCEM (rather than seeing it as a new estimator). It restricts the policy to a cluster-informed form, ∑︂ π(a | x) = π rm (a | x, c)π cl (c | x), c∈C
89
where π rm (a | x, c) = 1[a = argmaxa′ ∈c r̂(x, a′ )] is fixed, model-based policy that deterministically selects the best action within each cluster. Learning is then simplified to finding the optimal cluster-level policy π cl that maximizes the OffCEM objective in Equation (7.11): )︄ (︄ n cl ∑︂ ∑︂ π (C | X ) 1 i i (Ri − r̂(Xi , Ai )) + π cl (c | Xi )r̂c∗ (Xi ) , (7.13) V̂potec (π cl ) = n i=1 π0 (Ci | Xi ) c∈C where r̂c∗ (x) = maxa∈c r̂(x, a) is the estimated reward of the best action in cluster c. This practical decomposition has the same optimal oracle policy as OffCEM: π∗POTEC = π∗OffCEM . Policy convolution (PC). Moving beyond hard clustering, PC (Sachdeva et al., 2024) leverages the assumption that actions close in an embedding space yield similar rewards. For each action a, it aggregates over its neighborhood of nearest neighbors Nϵ (a) = {a′ : d(a, a′ ) < ϵ}, where d is a pre-defined distance metric (e.g., ℓ2 distance between action embeddings): n
V̂pc (π) =
1 ∑︂ π(Nϵ (Ai ) | Xi ) Ri , n i=1 π0 (Nϵ (Ai ) | Xi )
with π(Nϵ (a) | x) =
∑︂
π(a′ | x) .
(7.14)
a′ ∈Nϵ (a)
The induced oracle policy is deterministic: it selects the action a′ that maximizes an aggregated neighborhood score. Each logged neighbor ā ∈ Nϵ (a′ ) contributes its reward r(x, ā), weighted by the conditional probability of observing ā under the logging policy restricted to its neighborhood. ⎧ ⎫⎤ ⎡ ⎨ ∑︂ π (ā | x)r(x, ā) ⎬ 0 ⎦. π∗PC (a | x) = I ⎣a = argmax (7.15) ′ ⎩ π a ∈A 0 (Nϵ (ā) | x) ⎭ ′ ā∈Nϵ (a )
Other recent IPS variants for large action spaces (Peng et al., 2023; Cief et al., 2024; Taufiq et al., 2024) are often extensions of MIPS that relax its core assumptions. We focused on four methods (MIPS, OffCEM, POTEC, and PC), which we consider representative of this family. Since these variants largely share the same MIPS foundation and optimization procedure (with the notable exception of POTEC), we expect our findings to be generally applicable.
7.1.3
Optimization Challenges
The effectiveness of IPS-based estimators in off-policy learning is often limited by their challenging optimization landscape. These objectives become difficult to optimize when paired with standard, expressive policy classes such as the softmax. This section explores why this occurs and introduces objective-aware parametrization as a strategy to mitigate, though not entirely solve, the problem. 90
To analyze the optimization process, we consider policies parametrized by a softmax function over an effective action space 2 Aeff ⊆ A, which is the set of actions that can be assigned non-zero probability. By default, Aeff = A, but we explain below why restricting it to match the structure of the estimator’s oracle policy can be beneficial. Specifically, the policy takes the form: πθ (a | x) = ∑︁
exp(sθ (x, a)) 1a∈Aeff , ′ a′ ∈Aeff exp(sθ (x, a ))
∀a ∈ A ,
(7.16)
where sθ (x, a) is a learnable score function. Common choices are linear softmax scores: lightweight: sθ (x, a) = ϕ(x, a)⊤ θ ,
heavyweight: sθ (x, a) = ϕ(x)⊤ θa , (7.17)
which we call lightweight parametrization (a single shared parameter vector θ, corresponding to a joint reward model) and heavyweight parametrization (separate parameters θa for each action, corresponding to a disjoint reward model). The size of the effective action space, Keff = |Aeff |, is the critical factor governing optimization difficulty. The following propositions (proofs in Section D.2, adapted from Chen et al. (2019); Mei et al. (2020a)) reveal the severity of the problem. First, gradient-based methods can become trapped in suboptimal regions for extended periods. Proposition 3 (Optimization plateaus). For any IPS-based estimator V̂ that is linear in π, even with a linear softmax policy, there exist problem instances where gradient ascent remains trapped in a suboptimal region for O(Keff ) iterations. Second, the optimization landscape has numerous poor local maxima. Proposition 4 (Local maxima). Under similar conditions, the optimization landscape for IPS-based objectives can contain a number of local maxima that is exponential in Keff . These results highlight that Keff plays a central role in optimization difficulty. The standard choice of Aeff = A, which sets Keff = K, leads to optimization failure in large action spaces where K can reach millions: learning must navigate a landscape with potentially O(K)-length plateaus and exponentially many local maxima. Surprisingly, even sophisticated methods designed specifically for large action spaces often fall into this trap. At first glance, methods such as MIPS, OffCEM, and PC appear to operate in a smaller space because their objectives involve marginalized probabilities: π(Ci | Xi ) in MIPS and OffCEM, or π(Nϵ (Ai ) | Xi ) in PC. However, these marginalized terms are defined as sums over an underlying action-level policy: ∑︂ ∑︂ π(Ci | Xi ) = π(a | Xi ) , and π(Nϵ (a) | x) = π(a′ | x) . a′ ∈Nϵ (a)
a∈Ci 2
The effective action space can also depend on context x, i.e., Aeff (x) ⊆ A. We omit this dependence for notational simplicity.
91
Then, if π(a | x) is a softmax over A, then Keff = K and Propositions 3 and 4 apply with Keff = K which is large. The only exception is POTEC, which fixes the intra-cluster policy π rm and only optimizes a cluster-level policy π cl . This reduces the effective action space to Aeff = C with Keff = |C| ≪ K, directly mitigating the optimization pathologies. Design implications: objective-aware parametrization The choice of Keff introduces a fundamental trade-off. A smaller effective action space simplifies the optimization landscape, but risks excluding the optimal action and reduces policy expressiveness. If Aeff is chosen arbitrarily, it may degrade performance. The challenge is to find the sweet spot: a parametrization constrained enough to be optimizable, yet expressive enough to contain the objective’s maximizer. This is precisely where our asymptotic analysis helps. The oracle policy π∗method reveals the minimal sufficient set of actions required to maximize each objective. By aligning the policy parametrization with this structure, we can reduce Keff without sacrificing performance: the core principle of our proposed objective-aware parametrization. For instance, the oracle policies for IPS, cIPS, and ES are confined to the support of the logging policy, S0 (x). This implies that Aeff = S0 is sufficient, reducing Keff from K to |S0 | ≪ K. Similarly, for OffCEM and MIPS, the cluster-level structure of their oracle policies suggests a two-stage decomposition similar to that of POTEC, reducing Keff to |C|. We summarize these observations as claims, validated empirically in Section 7.3: Claim 1. For IPS, cIPS, and ES, restricting the policy support to S0 reduces Keff and yields superior learned policies. Claim 2. For OffCEM and MIPS, a two-stage POTEC-style decomposition that optimizes at the cluster level outperforms action-level parametrization. While objective-aware parametrization mitigates the optimization pathologies of Propositions 3 and 4 by reducing Keff , it only treats the symptoms without curing the underlying non-concavity. In the next section, we propose a more fundamental shift: abandoning value estimation in favor of inherently tractable objectives.
7.2
Analysis of PWLL objectives
To overcome the optimization challenges of IPS-based objectives, we consider policyweighted log-likelihood (PWLL) objectives. These methods trade accurate value estimation for a well-behaved, concave optimization landscape, leading to more robust and effective policy learning. General form. Given a positive weighting function g(r, p0 ), the PWLL objective is: n
1 ∑︂ g(Ri , π0 (Ai | Xi )) log π(Ai | Xi ). Ûg (π) = n i=1 92
(7.18)
The key motivation behind PWLL is to replace the linear dependence on the policy in IPS-based estimators, responsible for plateaus and local maxima in Section 7.1, with a concave transformation. Softmax policies are parametrized through scores sθ (x, a), and the map s ↦→ log softmax(s) is concave. Consequently, for common linear parametrizations in Equation (7.17), the composition log πθ (a | x) is concave in θ. This removes the optimization pathologies inherent to IPS-based objectives. Proposition 5 (proof in Section D.2.) formalizes this advantage. Proposition 5. For linear softmax policies πθ , the PWLL objective Ûg (πθ ) is concave in θ. With ℓ2 regularization, it is strongly concave. Proposition 5 makes PWLL appealing for stochastic optimization. In Section D.3, we show that under standard assumptions of bounded feature norms ∥ϕ(x, a)∥ and weights g(Ri , π0 (Ai , Xi )), these objectives satisfy the regularity conditions necessary to invoke established convergence theorems (Garrigos and Gower, 2023). This allows us to derive problem-dependent convergence guarantees: stochastic gradient ascent attains a global √ O(1/ T ) rate in the general (concave) case (Proposition 12), accelerating to a geometric rate under ℓ2 -regularization (Proposition 13). Beyond optimization properties, PWLL also admits a simple statistical interpretation. Ûg (π) in Equation (7.18) is a weighted log-likelihood: the term log π(Ai | Xi ) performs standard behavior cloning, while the weight g(Ri , π0 (Ai | Xi )) determines how desirable 3 each logged sample is. This turns off-policy learning into a form of logging-aware and reward-weighted maximum-likelihood estimation. Different choices of g encode different notions of desirability. For example, the weighting g(r, π0 (a | x)) =
r max{π0 (a | x), τ }
emphasizes samples with high reward while reducing the influence of actions that the logging policy selected very frequently. At the same time, the clipping at τ prevents extremely rare actions from receiving disproportionately large weights, ensuring that their contribution is attenuated once π0 (a | x) falls below the threshold. In this view, desirable samples are those that provide strong reward evidence without allowing very small propensities to dominate the updates. Many other PWLL variants arise from different choices of g (see below), each specifying a distinct prioritization scheme for the logged data, while all benefit from the concavity induced by the logarithmic term. To illustrate the qualitative difference between PWLL and IPS-based objectives, we construct a simple offline bandit problem with K = 3 actions and visualize the resulting optimization landscapes in a two-parameter policy space. Concretely, we consider a noncontextual setting with deterministic mean rewards r = (0.9, 0.7, 0.2) and a logging policy π0 whose support places almost all mass on action 3 (π0 (1) = 0.002, π0 (2) = 0.003, π0 (3) = 0.995). We generate a fixed dataset of n = 60 logged samples (Ai , Ri ) by drawing actions Ai ∼ π0 and binary rewards from the corresponding Bernoulli distributions, Ri ∼ Bern(r(Ai )). To obtain a two-dimensional visualization, we parameterize the target 3
By how desirable an action is, we mean how strongly this action should influence the learned policy.
93
PWLL Landscape (3D View)
OPE Landscape (3D View) PWLL Landscape (2D Projection)
200
Global Min (4.81, 4.41)
235
10
210
150
185
5 2
50
135
0
110 85
5
60 35
10
15
10
5
5
0
10
1
(a) PWLL (3D view)
15
15 15
10
5
0
5
1
10
15
OPE Landscape (2D Projection) Global Min (-1.01, -0.60)
0.16
10
0.36 0.56
5
0.76 0.96
0
1.16 5
1.36 1.56
10
1.76
2
10
2
0 15 10 5 0 5 10 15
160
Loss
100
0.2 0.4 0.6 0.8 1.0 1.2 1.4 1.6 1.8 15 10 5 0 5 10 15
Loss
15
2
250
15
(b) PWLL (2D projection)
15
10
5
5
0
10
1
(c) IPS-based (3D view)
15
15 15
10
5
0
1
5
10
15
(d) IPS-based (2D projection)
Figure 7.1: Optimization landscapes on a toy example. PWLL (cLPI) vs IPS-based (cIPS). ∑︁ policy using a softmax over three logits: πθ (a) = eθa / b∈[3] eθb , fixing the logit associated with action 3 as θ3 = 1, and letting the remaining two logits be free parameters (θ1 , θ2 ). In Figure 7.1, the PWLL landscape is concave with well-scaled gradients, and optimization trajectories converge reliably from roughly any initialization. In contrast, the IPS-based landscape consists of flat regions, separated by a narrow band of extremely steep curvature. This creates both vanishing and exploding gradients, severe ill-conditioning, and high sensitivity to initialization and learning rate. This aligns with the optimization pathologies in Propositions 3 and 4. Remark 7 (Beyond linear-softmax policies). The concavity guarantee of Proposition 5 assumes linear-softmax policies. In many large-scale recommendation systems, a deep encoder is pre-trained and kept fixed, and only a final linear head is optimized for the downstream task; in this case, the policy is still linear-softmax in the trainable parameters, and PWLL objectives retain their concavity. When the full network is trained end-toend, concavity no longer holds. Yet, PWLL’s gradients g(Ri , π0 (Ai |xi ))∇θ log πθ (Ai | Xi ) match the structure of cross-entropy gradients, which are known to produce stable and well-scaled updates in deep architectures. Thus, even without formal guarantees, PWLL maintains substantially more benign optimization dynamics than IPS-based objectives. Local policy improvement (LPI). Liang and Vlassis (2022) set g(r, p0 ) = r, which optimizes the log-likelihood of actions weighted by their observed rewards: n
Ûlpi (π) =
1 ∑︂ Ri log π(Ai | Xi ) . n i=1
(7.19)
The oracle policy balances reward-seeking with imitation of the logging policy: π∗LPI (a | x) ∝ r(x, a)π0 (a | x) .
(7.20)
r Clipped LPI (cLPI). uses importance-weight clipping, setting g(r, p0 ) = max(p : 0 ,τ ) n
1 ∑︂ Ri Ûclpi (π) = log π(Ai | Xi ) . n i=1 max{π0 (Ai | Xi ), τ } 94
(7.21)
In a similar spirit to cIPS, its oracle policy corrects for action frequency under π0 , downweighting the influence of rare actions due to the clipping: π∗cLPI (a | x) ∝ r(x, a)
π0 (a | x) . max{π0 (a | x), τ }
(7.22)
KL regularization (RegKL). To further amplify the reward signal relative to the logging policy prior, RegKL uses an exponential weighting function g(r, p0 ) = exp(r/β): n
Ûregkl (π) =
1 ∑︂ exp(Ri /β) log π(Ai | Xi ) . n i=1
(7.23)
The oracle policy is proportional to the logging policy, weighted by the exponentiated reward: [︁ ]︁ (7.24) π∗RegKL (a | x) ∝ Er∼p(·|x,a) exp(r/β) π0 (a | x). The temperature parameter β smoothly interpolates between behavior cloning (β → ∞) and greedy reward maximization (β → 0). Note that BPR (Rendle et al., 2012) can be seen as an approximate PWLL objective, and we included it in our experiments. In fact, this general form of PWLL lends itself to numerous variations by modifying the weighting function g. For instance, one could introduce variants inspired by regularized IPS like exponential smooting (Chapter 8). While many such variants can be proposed for specific use cases, the central message of our work is that the well-behaved optimization landscape of the PWLL family is of greater practical importance than the estimation accuracy of IPS-based objectives. Thus, an exploration of these PWLL variants is beyond our scope. We contend that the foundational methods analyzed above, LPI, cLPI, and RegKL, along with the widely used BPR are sufficient to demonstrate the inherent advantages of PWLL objectives. Finally, PWLL resembles reward- or advantage-weighted behavioral cloning objectives in RL (Nair et al., 2020; Wang et al., 2020; Peng et al., 2019; Peters, 2006). While those methods address multi-step MDPs and often focus on mitigating distributional shift and bootstrapping errors, we focus on offline contextual bandits with large action spaces: identifying objectives and parametrizations that remain optimizable as K grows, rather than accurately estimating V (π). PWLL is critic-free and uses logged rewards and propensities through a weighting function g(Ri , π0 (Ai | Xi )) that induces concave optimization landscapes for common policy classes. This yields substantial gains in largeK bandits without the overhead of value-function estimation. PWLL’s optimizationcentric perspective complements the usual KL-regularized or trust-region interpretations of these RL methods.
7.3
Empirical Analysis
We conduct our empirical evaluation on three large-scale recommendation datasets: MovieLens (K = 60k) (Lam and Herlocker, 2016), Twitch (K = 200k) (Rappaz et al., 2021), and 95
GoodReads (K = 1M) (Wan et al., 2019). These benchmarks feature action spaces with up to one million items, representing some of the largest settings studied in the offline policy learning literature. For all experiments, we employ the common softmax inner-product policies. We compare methods from both objective families. For IPS-based objectives, we include IPS, ES, DR, MIPS, OffCEM, POTEC, and PC in Section 7.1. For PWLL objectives, we evaluate LPI, cLPI, RegKL, and BPR in Section 7.2. All implementation details are provided in Section D.4.
7.3.1
Optimization is the Main Bottleneck
To test our central hypothesis that optimization challenges are a more significant barrier than estimation accuracy, we evaluate how objectives perform under various optimization configurations. If an algorithm’s success is highly dependent on specific hyperparameters like batch size or learning rate, it suggests a difficult, non-robust optimization landscape. This experiment directly probes the practical trainability of each method, a key aspect our paper argues is often overlooked. The results strongly support our claim. As shown in Figure 7.2, IPS-based objectives are highly sensitive to batch size and learning rate schedule: minor changes can cause performance collapse, making them difficult to tune and train reliably. In contrast, PWLL objectives remain robust, achieving consistently high reward across all configurations. This stability translates directly into better learned policies: PWLL objectives outperform IPSbased objectives on all datasets. Even POTEC, a state-of-the-art method designed for large action spaces, is surpassed by the much simpler and easier-to-optimize cLPI. One might assume that an objective designed for estimation fidelity, such as a low-MSE IPS-based estimator, would naturally yield a better policy. Our findings show this is not the case. The superiority of PWLL objectives, which are poor value estimators by design, provides compelling evidence against this estimator-centric view. This reinforces our main takeaway: in large action space settings, a tractable optimization landscape is a more critical feature for a learning objective than its statistical accuracy. For completeness, an experiment tracking the MSE of methods is given in Section D.4. The figure also supports Claim 2. Indeed, there is a consistent performance gap between POTEC and OffCEM. Both methods are designed to maximize the same asymptotic objective as we show in Section 7.1; their statistical goals are identical. The divergence in performance, therefore, can be attributed entirely to their differing optimization strategies. POTEC’s use of a two-stage, cluster-level optimization proves far more effective than OffCEM’s naive, action-level parametrization.
7.3.2
Objective-Aware Parametrization
To empirically validate Claim 1, we compare a naive, whole-action-space parametrization against our proposed objective-aware approach, which restricts the policy’s effective action space to the logging policy support, S0 . As shown for the IPS objective in Figure 7.3, the naive approach is highly unstable, with performance collapsing under simple learning configurations. In contrast, the objective-aware version is very robust, achieving high reward consistently across all batch sizes and schedules. This benefit extends even to 96
Effect of Optimization Hyperparameters on Performance Reward on MovieLens (K=60K)
LR Schedule: Warmup Cosine
Reward on Twitch (K=200K)
LR Schedule: None
0.30
0.30
0.25
0.25
0.25
0.20
0.20
0.20
0.15
0.15
0.15
0.10
0.10
0.10
0.05
0.05
0.05
0.00
0.00
24
25
26
27
24
28
0.00 25
27
26
24
28
0.30
0.30
0.30
0.25
0.25
0.25
0.20
0.20
0.20
0.15
0.15
0.15
0.10
0.10
0.10
0.05
0.05
0.05
0.00
0.00
24
Reward on GoodReads (K=1M)
LR Schedule: One Cycle
0.30
25
26
27
24
28
27
26
24
28
0.35
0.35
0.30
0.30
0.30
0.25
0.25
0.25
0.20
0.20
0.20
0.15
0.15
0.15
0.10
0.10
0.10
0.05
0.05
0.05
0.00
0.00 25
26
27
28
24
26
27
28
25
26
27
28
26
27
28
0.00 25
0.35
24
25
0.00 25
Batch Size
27
26
28
24
25
Batch Size IPS ES
DR MIPS
OffCEM PC
POTEC LPI
Batch Size cLPI RegKL
BPR
Figure 7.2: Effect of batch size and learning rate schedule on final validation reward using three large-scale datasets. IPS-based objectives are highly sensitive, while PWLL objectives are robust.
97
Effect of Objective-Aware Parametrization on Performance Reward - MovieLens (K=60K)
LR Schedule: None
LR Schedule: Warmup Cosine
LR Schedule: One Cycle
0.30
0.30
0.30
0.25
0.25
0.25
0.20
0.20
0.20
0.15
0.15
0.15
0.10 24
0.10 25
26
Batch Size IPS (Objective-Aware Defined on Support)
27
28
24
0.10 25
27
26
Batch Size IPS (Whole Action Space)
28
24
25
26
27
28
Batch Size
cLPI (Objective-Aware Defined on Support)
cLPI (Whole Action Space)
Figure 7.3: The effect of objective-aware parametrization for IPS and cLPI on MovieLens. inherently stable PWLL objectives like cLPI, which achieve even better performance with the restricted support. This provides strong evidence for Claim 1: aligning the policy structure with the objective’s inductive bias simplifies the optimization landscape, leading to greater stability and superior learned policies. This finding holds across all datasets, with full results available in Section D.4.
7.4
Conclusion
The dominant approach to off-policy learning focuses on developing sophisticated IPSbased estimators while neglecting a crucial factor: the optimization landscape. We demonstrated, both theoretically and empirically, that this landscape becomes prohibitively difficult to optimize in large action spaces, undermining the practical effectiveness of even state-of-the-art estimators. Our analysis motivates two strategies. First, objective-aware policy parametrizations align the policy class with the estimator’s inductive bias, reducing the effective search space. Second, PWLL objectives abandon value estimation entirely in favor of inherently concave optimization landscapes. Experiments confirm that this focus on optimization tractability yields more robust learning, reduced sensitivity to hyperparameters, and superior policies. Our work has several limitations. First, PWLL objectives are not value estimators: they cannot be used for off-policy evaluation or policy selection (choosing the best policy from a finite candidate set) when accurate value estimates and their comparison are required. Second, the concavity guarantee of Proposition 5 holds only for linear-softmax policies; when training deep networks end-to-end, PWLL retains favorable gradient structure but loses formal concavity guarantees, although IPS-based objectives face even more severe optimization challenges in this setting. Third, PWLL’s oracle policies inherently depend on the logging policy (e.g., π∗LPI ∝ r(x, a)π0 (a | x)), which may be suboptimal when π0 has poor coverage of high-reward actions; however, this limitation is shared by IPS-based methods, whose oracle policies similarly depend on π0 ’s support.
98
Chapter 8
Principled Pessimism for Exponential Smoothing and Beyond
Contents Analysis of IPS-Based Objectives . . . . . . . . . . . . . . . . . . . . .
87
7.1.1
Standard IPS-Based Objectives . . . . . . . . . . . . . . . . .
87
7.1.2
Large-Scale IPS-Based Objectives . . . . . . . . . . . . . . . .
89
7.1.3
Optimization Challenges . . . . . . . . . . . . . . . . . . . . .
90
7.2
Analysis of PWLL objectives . . . . . . . . . . . . . . . . . . . . . . .
92
7.3
Empirical Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
95
7.3.1
Optimization is the Main Bottleneck . . . . . . . . . . . . . .
96
7.3.2
Objective-Aware Parametrization . . . . . . . . . . . . . . . .
96
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
98
7.1
7.4
Having explored structured direct methods in Chapter 6 and optimization-focused objectives in Chapter 7, we now turn to inverse propensity scoring (IPS). Despite its current practical limitations in large action spaces, many practitioners remain committed to IPSbased methods for their unbiasedness and theoretical guarantees, which enable principled safe off-policy learning. In this chapter, we improve IPS through exponential smoothing, a differentiable importance-weight regularization technique that enables a controlled biasvariance trade-off. Then, we adopt the pessimistic framework introduced in Section 5.1, deriving principled uncertainty penalties for our regularized estimators for safe policy learning. Prior work on pessimistic off-policy learning has derived objectives from generalization bounds (Swaminathan and Joachims, 2015a; London and Sandler, 2019), but these approaches suffer from critical limitations: (i) they provide only one-sided bounds that fail to control estimation error in absolute value, limiting their ability to certify estimator quality, (ii) the resulting bounds are intractable and incompatible with stochastic optimization, and (iii) the pessimistic objectives require careful hyperparameter tuning. 99
We address these limitations by deriving tractable two-sided PAC-Bayes generalization bounds that can be optimized directly via stochastic gradient ascent. Unlike prior work (Sakhi et al., 2022), our analysis applies to standard IPS without assuming bounded importance weights, requiring only bounded second moments. Our bounds reveal that the optimal importance-weight smoothing parameter α depends on the quality of the logging policy. Furthermore, our framework generalizes to a broad class of importanceweight regularization techniques, yielding unified pessimistic objectives that enable fair comparison across different importance-weight regularization techniques. We present this extension in the final two sections of this chapter. The chapter is organized as follows. Section 8.1 presents background on regularized IPS and pessimism. Section 8.2 identifies the shortcomings of hard clipping and introduces our exponential smoothing estimators. Section 8.3 leverages PAC-Bayes theory to derive two-sided generalization bounds within the pessimistic framework. Section 8.4 discusses implications of our results. Section 8.5 demonstrates favorable performance across diverse benchmarks. Finally, Sections 8.6 and 8.7 extend the framework to other importanceweight regularizations and compare them under a unified pessimistic objective.
8.1
Background
We consider the off-policy setting in Section 5.1, where we have access to logged data Dn = {(Xi , Ai , Ri )}ni=1 collected by a known logging policy π0 . As additional notation, we let µπ be the joint distribution of (X, A, R); µπ (x, a, r) = ν(x)π(a|x)p(r|x, a) ,
so that (Xi , Ai , Ri ) ∼ µπ0 .
Our goal remains to find a policy π̂ ∈ Π that maximizes the value V (π) = EX∼ν,A∼π(·|X) [r(X, A)] .
8.1.1
Regularized IPS
This chapter focuses on the IPS estimator (Horvitz and Thompson, 1952; Dudík et al., 2012), which estimates the value V (π) by re-weighting the samples as n
1 ∑︂ Ri w(Ai |Xi ), V̂ips (π) = n i=1
(8.1)
where w(a|x) = π(a|x)/π0 (a|x) are the importance weights. While IPS provides an unbiased estimate of V (π) when the common support condition holds (i.e., π0 (a|x) = 0 implies π(a|x) = 0), its variance grows with these importance weights (Swaminathan et al., 2017), which can be arbitrarily large when the target policy π and logging policy π0 differ significantly. To mitigate this variance issue, it is common to transform the importance weights using regularization functions that introduce controlled bias to reduce variance. A regularized IPS estimator takes the form: n
1 ∑︂ V̂ (π) = Ri ŵ(Ai |Xi ), n i=1 100
(8.2)
where ŵ(a|x) ≤ w(a|x) are the regularized importance weights. A common importanceweight regularization approach is clipping where ŵ(a|x) = min( ππ(a|x) , M ), M > 0. 0 (a|x)
8.1.2
Pessimistic Objectives
Within the pessimistic framework introduced in Section 5.1, we seek to maximize: π̂ = argmaxπ∈Π [V̂ (π)−pen(π)] where pen(·) is a penalty term. The construction of this penalty has been approached in various ways, but generally relies on lower confidence bounds on the policy value: Evaluation bounds (Metelli et al., 2021) provide confidence intervals for a fixed target policy π, showing that with probability at least 1 − δ: |V (π) − V̂ (π)| ≤ f (δ, π, π0 , n).
(8.3)
Essentially, Equation (8.3) indicates that for a fixed policy π ∈ Π, the event |V (π) − V̂ (π)| ≤ f (δ, π, π0 , n) holds with high probability. However, this event depends on the target policy π. Thus Equation (8.3) is useful for evaluating a single target policy when having access to multiple logged data sets Dn . This poses a problem for off-policy learning, where we optimize over a potentially infinite space of policies using a single logged data set Dn . This is the fundamental theoretical limitation of using evaluation bounds similar to Equation (8.3) in off-policy learning. While one can transform Equation (8.3) into a generalization bound that holds uniformly over all π ∈ Π via a union bound, this typically introduces intractable complexity terms, making the resulting pessimistic objectives, which maximize the lower confidence bound, equally intractable. One-sided generalization bounds (Swaminathan and Joachims, 2015a; London and Sandler, 2019; Sakhi et al., 2022) address this limitation by providing bounds that hold simultaneously for all policies ∈∈ Π. For δ ∈ (0, 1), with probability at least 1 − δ: V (π) ≥ V̂ (π) − g(δ, Π, π, π0 , n),
∀π ∈ Π,
(8.4)
where the function g now depends on the policy space Π. This leads to the pessimistic objective: π̂ = argmax V̂ (π) − g(δ, Π, π, π0 , n).
(8.5)
π∈Π
However, one-sided bounds fail to attest to the quality of the estimator. To illustrate, consider a degenerate estimator V̂ poor (π) = 0 for all π ∈ Π. Since V (π) ∈ [0, 1], we trivially have a one-sided bound V (π) ≥ V̂ poor (π) with probability 1, yet this estimator is entirely uninformative about the true rewards. Two-sided generalization bounds resolve this issue by controlling both the upper and lower deviations leading to |V (π) − V̂ (π)| ≤ g(δ, Π, π, π0 , n),
∀π ∈ Π.
(8.6)
These bounds ensure the estimator quality and enable oracle inequalities of the form V (π̂) ≥ V (π∗ ) − 2g(δ, Π, π∗ , π0 , n), where π̂ is learned using Equation (8.5) and π∗ = 101
argmaxπ∈Π V (π) is the optimal policy. This shows the appeal of pessimism: the suboptimality gap depends on the bound evaluated at the optimal policy π∗ , meaning the estimator only needs to be precise for near-optimal policies rather than uniformly across the policy class Π. Alternative approaches include heuristics that simplify theoretical bounds for tractability (Swaminathan and Joachims, 2015a; London and Sandler, 2019; Wang et al., 2023), often penalizing by empirical variance or policy divergence while discarding complexity terms. Recent work on implicit pessimism (Gabbianelli et al., 2024; Sakhi et al., 2024) (published after the work in this chapter) shows that careful analysis of specific importance-weight regularizations can yield bounds where the penalty is policyindependent. In this case, maximizing the lower bound reduces to maximizing the estimator directly: the pessimism becomes implicit in the regularization itself. In this work, we derive a tractable two-sided PAC-Bayesian generalization bound for our exponential smoothing estimator and then generalize it to other importance-weight regularizations.
8.2
Exponential Smoothing
Importance-weight clipping (Swaminathan and Joachims, 2015a) yields the following commonly used estimators n (︁ )︁ 1 ∑︂ Ri min w(Ai |Xi ), M , IPS-min Ṽ (π) = n i=1 m
n
π(Ai |Xi ) 1 ∑︂ . IPS-max V̂ (π) = Ri n i=1 max(π0 (Ai |Xi ), τ ) τ
(8.7)
Here IPS-min clips the weights while IPS-max only clips π0 in the denominator since π is always smaller than 1. For instance, M ∈ R+ in Ṽ m (π) trades the bias and variance of the estimator. When M is large, the bias of Ṽ m (π) is small but its variance may be large. On the other hand, the variance goes to 0 when M ≈ 0 since in that case Ṽ m (π) ≈ 0 for any π ∈ Π. Similarly, τ ∈ [0, 1] trades the bias and variance of V̂ τ (π) and can be seen as τ ≈ M1 . This hard clipping has some limitations. First, min(·, M ) leads to non-differentiable objectives that may require additional care in optimization (Papini et al., 2019). Also, min(·, M ) is constant on [M, ∞) leading to objectives with zero gradients for any policy π that satisfies w(Ai |Xi ) > M for any i ∈ [n]. More importantly, hard clipping is sensitive to the choice of the clipping threshold M . In practice, tuning M is challenging and may cause the learned policy to match the logging policy, leading to minimal improvements. To see this, consider the following illustrative example. For simplicity, suppose that the problem is non-contextual, in which case the reward function r only depends on the actions a ∈ A. It follows that policies do not depend on x ∈ X ; they are now probability distributions π(·) over A. Also, assume that A = [100] and that the reward received after taking action a ∈ [100] is binary. That is, R ∼ 102
Bern(r(a)) where r(a) = 0.1 − 10−3 (a − 1) is the expected reward of action a, and for any p ∈ [0, 1], Bern(p) is the Bernoulli distribution with parameter p. This means that the best action is 1 and the worst is 100. Finally, the logging policy π0 (·) is ϵ-greedy centered at action 50. That is π0 (50) = 1 − ϵ, and for any a ̸= 50, π0 (a) = 99ϵ , with ϵ = 0.05. Now consider 100 deterministic policies πa (·) for a ∈ [100] such that πa (·) is the Dirac distribution centered at a. In Figure 8.1, we plot the estimated reward of the policies πa using either IPS in Equation (8.1) √ or IPS-min in Equation (8.7). We generate n = 50k samples and set M = 100 = O( n) as suggested by Ionides (2008). With this choice of M , IPS-min underestimates the reward of all policies πa for a ̸= 100 since their weights πa /π0 are either 0 or 99/ϵ > M . The estimated reward of IPS-min is maximized in π50 ≈ π0 only. Thus, if we optimize Ṽ m (·) over Dirac policies, we will converge to the logging policy despite its bad performance. Although the other variant of hard clipping, IPS-max in Equation (8.7), is differentiable, it is still sensitive to τ and may induce high bias similar to Figure 8.1. This is due to some loss of information related to the preferences of the logging policy. Indeed, for two actions a and a′ such that π0 (a|Xi ) ≪ π0 (a′ |Xi ) < τ for an observed context Xi , the propensity scores π0 (a|Xi ) and π0 (a′ |Xi ) will be clipped to the same value τ . Thus the information that, for context Xi , action a′ is preferred by the logging policy than action a will be lost.
0.25 true ips ips-min, M=100
0.20 0.15 0.10 0.05 0.00
0
20
40
60
80
100
Figure 8.1: Effect of hard clipping on the estimation quality. The x-axis corresponds to actions a ∈ [100]. The y-axis is the estimated reward of each of the 100 policies πa using either IPS or IPS-min. The cyan line is the true reward for each policy πa . To mitigate this, we propose the following exponential smoothing correction for IPS. Our estimators are defined as n
IPS-α : V̂ α (π) =
1 ∑︂ Ri ŵα (Ai |Xi ) , α ∈ [0, 1] , n i=1 n
1 ∑︂ IPS-β : Ṽ (π) = Ri w̃πβ (Ai |Xi ) , β ∈ [0, 1] , n i=1 β
103
(8.8)
β
π(a|x) β where ŵα (a|x) = ππ(a|x) α and w̃π (a|x) = π (a|x)β . Here standard IPS is recovered for α = 1 0 (a|x) 0 and β = 1. These estimators yield smooth, everywhere-differentiable objectives and avoid the flat regions induced by hard clipping; this improves optimization in practice. Also, in contrast with IPS-max in Equation (8.7), V̂ α (π) preserves the preferences of the logging policy. Precisely, for two actions a and a′ such that π0 (a|Xi ) < π0 (a′ |Xi ) for an observed context Xi , we still have π0 (a|Xi )α < π0 (a′ |Xi )α and the information that action a′ is preferred by the logging policy than action a is preserved.
While a similar correction to IPS-β was proposed in Korba and Portier (2022), its use in off-policy learning is novel. Also, Su et al. (2020); Metelli et al. (2021) regularized w 1w the importance weights w as λ1λ+w 2 , λ1 > 0 and 1−λ +λ w , λ2 ∈ [0, 1], respectively. Thus, 2 2 the expression of both corrections is very different from ours. More importantly, these corrections entail different properties than ours. Roughly speaking, our correction allows us to simultaneously (1) control a tuning parameter α ∈ [0, 1] that is in a bounded domain [0, 1], (2) without constraining the resulting importance weights to be bounded, (3) and to obtain tractable PAC-Bayes generalization bounds as the correction ππα is linear in 0 π; a technical requirement of PAC-Bayes analysis. In contrast, Metelli et al. (2021); Su et al. (2020) do not provide generalization guarantees; they focus on estimation accuracy (e.g., through mean squared error) and only propose heuristics for off-policy learning. Those heuristics are not based on theory, in contrast with ours which is directly derived from our generalization bound. Also, our approach has favorable empirical performance (Section E.3.6). Although Korba and Portier (2022, Lemma 1) show that smoothing the importance weights similarly to IPS-β in Equation (8.8) reduces the variance, it might still be unclear how α and β trade the bias and variance of our estimators in off-policy learning. To see this, let α ∈ [0, 1], then we have [︁ ]︁ |B(V̂ α (π))| ≤ EX∼ν,A∼π(·|X) 1 − π0 (A|X)1−α , (8.9) [︂ ]︂ 1 [︁ π(A|X) ]︁ , V V̂ α (π) ≤ EX∼ν,A∼π(·|X) n π0 (A|X)2α−1 with B(V̂ α (π)) = E[V̂ α (π)]−V (π) and V[V̂ α (π)] = E[(V̂ α (π)−E[V̂ α (π)])2 ] are respectively the bias and the variance of V̂ α (π). The bound of the bias in Equation (8.9) is minimized in α = 1 (standard IPS); in which case it is equal to 0 (standard IPS is unbiased). In contrast, the bound of the variance is minimized in α = 0. Thus if the variance is small or n is large enough such that E[π(A|X)/π0 (A|X)2α−1 ]/n → 0, then we set α → 1. Otherwise, we set α → 0. This shows that α trades the bias and variance of V̂ α . More details and a similar discussion for Ṽ β (π) are deferred to Section E.1.
8.3
PAC-Bayes Analysis for Off-Policy Learning
We now derive generalization bounds for our estimator. We opt for the PAC-Bayes framework for the following reasons. First, it is known to provide some of the tightest generalization bounds in challenging scenarios (Farid and Majumdar, 2021), for aggregated and randomized predictors (Alquier, 2021). Second, the bounds have a Kullback–Leibler (KL) divergence (Van Erven and Harremos, 2014) term DKL (Q∥P) that depends on a 104
fixed prior P and a learning posterior Q (see Section 8.3.1 for a brief introduction). This quantity can be seen as a complexity measure, similarly to the covering number (Maurer and Pontil, 2009). The difference is that complexity measures are uniform on the space of policies while the KL term in PAC-Bayes depends on the prior P and the posterior Q. This allows getting sharper bounds when the former is well chosen. Third, the PAC-Bayes perspective fits very well with off-policy learning. In fact, a policy π can be written as an aggregation of predictors under some distribution Q. Thus the prior P can be associated with the logging policy π0 that we want to improve upon while the posterior Q is related to the learning policy π. Fourth, London and Sandler (2019) showed that PAC-Bayes can lead to tractable and scalable objectives, an important consideration for this thesis.
8.3.1
Elements of PAC-Bayes
Let Z = X × Y be an instance space: e.g., X and Y are the input and output space in supervised learning. Let H = {h : X → Y} denote a hypothesis space of mappings from X to Y (predictors). Also, let L : H × Z → R be a loss function and assume access to data Dn = (Zi )i∈[n] drawn from an unknown distribution D. Let Risk(h) = EZ∼D [L(h, Z)] be the risk of h ∈ H while
n 1 ∑︂ ˆ︃ L(h, Zi ) Riskn (h) = n i=1
is its empirical counterpart. Then the main focus in PAC-Bayes is to study the generalization capabilities of random hypotheses Q on H by controlling the gap between the expected risk [︂ ]︂ under Q, Eh∼Q [Risk(h)], and the expected empirical risk under Q, ˆ︃ Eh∼Q Riskn (h) . For example, assume that L(h, Z) ∈ [0, 1] for any (h, Z) ∈ H × Z, let P be a fixed prior distribution on H and let δ ∈ (0, 1). Then with probability at least 1 − δ over Dn ∼ Dn , the following inequality holds simultaneously for any posterior distribution Q on H: √︄ √ 2 n ]︂ [︂ D (Q∥P) + log KL δ ˆ︃ n (h) + Eh∼Q [Risk(h)] ≤ Eh∼Q Risk . 2n This was originally proposed by McAllester (1998), and the reader may refer to Alquier (2021); Guedj (2019) for more elaborate introductions of PAC-Bayes theory. Connection to value functions. The loss L is often chosen as the negative reward, L(h, Z) = −r(h, Z). In this case, minimizing the Risk(h) is equivalent to maximizing the value function. Thus, PAC-Bayes bounds on the risk directly translate into guarantees on the discrepancy between empirical and true value, providing a principled way to reason about generalization in off-policy learning.
8.3.2
PAC-Bayes for Off-Policy Learning
Let H = {h : X → A} be a hypothesis space of mappings from X (contexts) to A (actions). Given a policy π and a context x ∈ X , the action distribution π(·|x) is induced 105
by a distribution Q over H (London and Sandler, 2019) such as [︁ ]︁ π(a|x) = πQ (a|x) = Eh∼Q 1{h(x)=a} .
(8.10)
This is not an assumption since any policy π has this form when H is rich enough (Sakhi et al., 2022, Theorem 2). From Equation (8.10), we observe that policies can be seen as an aggregation Eh∼Q [·] (under some distribution Q on the pre-defined hypothesis space H) of deterministic decision rules 1{h(x)=a} . This allows formulating off-policy learning as a PAC-Bayes problem. Before showing how this is achieved, we start by providing two practical policies of such form. Example 1 (softmax and}︁ mixed-logit policies). We define the hypothesis space {︁ H = hθ,γ ; θ ∈ RdK , γ ∈ RK of mappings hθ,γ (x) = argmaxa∈A ϕ(x)⊤ θa + γa . Here ϕ(x) outputs a d-dimensional representation of x, and γa is a standard Gumbel perturbation, γa ∼ G(0, 1) for any a ∈ A. Then exp(ϕ(x)⊤ θa ) , ⊤ a′ ∈A exp(ϕ(x) θa′ ) [︁ ]︁ (i) = Eγ∼G(0,1)K 1{hθ,γ (x)=a} ,
πθsof (a|x) = ∑︁
(8.11)
where (i) follows from the Gumbel-Max trick (GMT) (Luce, 2012; Maddison et al., 2014). Thus a softmax policy πθsof can be written as in Equation (8.10). Now we also consider random parameters θ ∼ N (µ, σ 2 IdK ) with µ ∈ RdK and σ > 0. Then, let mixL Q = N (µ, σ 2 IdK ) × G(0, 1)K , it follows that πQ = πµ,σ is a mixed-logit policy and it reads ]︃ [︃ exp(ϕ(x)⊤ θa ) mixL , πµ,σ (a|x) = Eθ∼N (µ,σ2 Id ) ∑︁ ⊤ a′ ∈A exp(ϕ(x) θa′ ) [︁ ]︁ = Eθ∼N (µ,σ2 Id ) ,γ∼G(0,1)K 1{hθ,γ (x)=a} . (8.12) Example 2 (Gaussian policies): Sakhi et al. (2022) removed the Gumbel {︁ noise γ in }︁ Equation (8.12) and consequently defined the hypothesis space as H = hθ ; θ ∈ RdK of mappings hθ (x) = argmaxa∈A ϕ(x)⊤ θa for any x ∈ X . Then, let Q = N (µ, σ 2 IdK ), it gaus follows that πQ = πµ,σ reads [︁ ]︁ gaus πµ,σ (a|x) = Eθ∼N (µ,σ2 Id ) 1{hθ (x)=a} . (8.13) To see why removing the Gumbel noise can be beneficial, the reader may refer to Section E.3.2. After motivating the definition of policies in Equation (8.10), we are in a position to relate our estimators to the general PAC-Bayes framework in Section 8.3.1. One technical requirement of our proof is that the estimator should be linear in π. Thus we focus on V̂ α (·) since Ṽ β (π) is non-linear in π. Let h ∈ H, x ∈ X , a ∈ A and r ∈ [0, 1], we define the objective Uα as Uα (h, x, a, r) = 106
1{h(x)=a} r. π0 (a|x)α
(8.14)
Using the definition in Equation (8.10) and the linearity of the expectation, we have that V̂ α (·) in Equation (8.8) can be written as ]︄ [︄ n 1 ∑︂ α Uα (h, Xi , Ai , Ri ) . V̂ (πQ ) = Eh∼Q n i=1 Moreover, the expectation of V̂ (πQ ) reads V α (πQ ) = Eh∼Q E(X,A,R)∼µπ0 [Uα (h, X, A, R)] . Finally, the main quantity of interest, the value V (πQ ) , can be expressed in terms of the objective with α = 1 , U1 , as V (πQ ) = Eh∼Q E(X,A,R)∼µπ0 [U1 (h, X, A, R)] . Since V̂ α (πQ ) is an unbiased estimator of V α (πQ ), PAC-Bayes can be used to bound V α (πQ ) − V̂ α (πQ ). This will allow bounding our quantity of interest V (πQ ) − V̂ α (πQ ).
8.3.3
Main Result
To ease the exposition, we assume that the rewards are deterministic. Then, in logged data Dn , Ri = r(Xi , Ai ) for any i ∈ [n]. Note that the same result holds for stochastic rewards. We discuss our result and sketch its proof in Section 8.4. The complete proof can be found in Section E.2.1. Theorem 4. Let λ > 0, n ≥ 1, δ ∈ (0, 1), α ∈ [0, 1], and let P be a fixed prior on H, then with probability at least 1 − δ over draws Dn ∼ µnπ0 , the following holds simultaneously for any posterior Q on H √︃ kl1 (πQ ) kl2 (πQ ) λ + Bnα (πQ ) + + Varαn (πQ ) . |V (πQ ) − V̂ α (πQ )| ≤ 2n nλ 2 √
where kl1 (πQ ) = DKL (Q∥P) + ln 4 δ n , and kl2 (πQ ) = DKL (Q∥P) + ln 1 Varαn (πQ ) =
n ∑︂
n i=1
4 , δ [︃
EA∼π0 (·|Xi )
Bnα (πQ ) = 1 −
n [︁ ]︁ 1 ∑︂ EA∼πQ (·|Xi ) π01−α (A|Xi ) , n i=1
]︃ πQ (A|Xi ) πQ (Ai |Xi )Ri2 + . π0 (A|Xi )2α π0 (Ai |Xi )2α
We start by clarifying that the prior P can be any fixed distribution on H. If we have access to P0 on H such that π0 = πP0 , then it is natural to set P = P0 . But this is just a choice and one may use priors that do not depend on π0 . Now we explain the main terms in our bound. First, the terms kl1 (πQ ) and kl2 (πQ ) contain the divergence DKL (Q∥P) which penalizes posteriors Q that differ a lot from the prior P. Moreover, Bnα (πQ ) is the bias conditioned on the contexts (Xi )i∈[n] ; Bnα (πQ ) = 0 when α = 1 and Bnα (πQ ) > 0 otherwise. Also, the first term in Varαn (πQ ) resembles the theoretical second moment of 107
the regularized importance weights ππα (without the reward) when they are seen as random 0 variables. Similarly, the second term in Varαn (πQ ) resembles the empirical second moment √ of ππα R (with the reward). Finally, if Varαn (πQ ) is bounded, then we can set λ = 1/ n, in 0 √ which case our bound scales as O(1/ n + Bnα (π√ Q )). In practice, we set α ≈ 1 leading to Bnα (πQ ) ≈ 0 and the bound would scale as O(1/ n). One of the main strengths of our result is that it holds for standard IPS with α = 1 under the assumption that Var1n (πQ ) is bounded. This assumption is less restrictive than assuming that the importance weight as a random variable, πQ (A|X)/π0 (A|X), is bounded, a required assumption for traditional concentration bounds. In contrast, Varαn (πQ ) only involves the expectations of the random variables πQ (A|Xi )/π0 (A|Xi )2α , and ratios of π0 evaluated at observed contexts and actions and (Xi , Ai )i∈[n] , that have non-zero probabilities under π0 by definition. Our result holds for fixed λ > 0 and α ∈ [0, 1]. In Section E.2.2, we extend this to any potentially data-dependent λ ∈ (0, 1) and α ∈ (0, 1]. The assumption that R ∈ [0, 1] can be relaxed to R ∈ [0, B] up to additional factors B 2 and B in Varαn (πQ ) and kl1 (πQ ), respectively. Finally, our bound is suitable for stochastic gradient ascent (Robbins and Monro, 1951) since data-dependent quantities are not inside a square root. This is important for scalability. Limitations. Our bound in Theorem 4 has two main limitations. (i) Using it to directly derive a data-independent suboptimality gap bound is not straightforward. This difficulty arises because our bound involves empirical quantities such as Bnα (πQ ) and Varαn (πQ ), whose dependence on the logged data prevents expressing the gap purely as a function of n. However, obtaining data-independent suboptimality guarantees was not the goal of this chapter. Instead, our focus was on deriving tractable and theoretically grounded bounds for exponential smoothing, that also perform well in practice when used for pessimistic objectives. (ii) Our result provides symmetric deviation bounds that simultaneously control the upper and lower deviations of V̂ α (πQ ) from V (πQ ). Yet, recent work (Gabbianelli et al., 2024), published after the paper corresponding to this chapter, indicates that the tails of regularized IPS estimators are inherently asymmetric. Consequently, tighter bounds may arise from developing asymmetric two-sided bounds that treat each deviation separately. We explored this direction in our follow-up work (Sakhi et al., 2024), where we derived some of the tightest bounds in the literature, with strong empirical performance.
8.3.4
Adaptive and Data-Driven Tuning of α
Theorem 4 assumes that α is fixed (although we extend it for data-dependent α in Section E.2.2). However, providing a procedure to tune α in an adaptive and data-dependent fashion is important in practice. Thus we propose to set √︃ 2kl2 (πQ ) Varαn (πQ ) , (8.15) α∗ = argmin Bnα (πQ ) + n α∈[0,1] where all the terms are defined in Theorem 4. Roughly speaking, α∗ establishes a biasvariance trade-off; it minimizes the sum of the√︂ bias term Bnα (πQ ) and the square root of the second moment term Varαn (πQ ), weighted by
108
2kl2 (πQ ) . Here Equation (8.15) is obtained n
by minimizing the bound in Theorem 4 with respect to both α and λ as follows. √︂ First, we 2kl (π ) minimize the bound in Theorem 4 with respect to λ; the minimizer is λ∗ = n Var2α (πQQ ) . n Then, the bound in Theorem 4 evaluated at λ = λ∗ becomes √︃ √︃ kl1 (πQ ) 2kl2 (πQ ) Varαn (πQ ) α + Bn (πQ ) + . (8.16) 2n n Finally, α∗ is defined as the minimizer of Equation (8.16) with respect to α ∈ [0, 1], and √︂ kl1 (πQ ) does not appear in Equation (8.15) as it does not depend on α. Note that α∗ 2n depends on both logged data Dn and the learning policy πQ . Thus it is adaptive; its value changes in each iteration during optimization.
8.4
Discussion
We start by interpreting and comparing our results to related work. Then, we present the technical challenges in Section 8.4.2. After that, we sketch our proof in Section 8.4.3.
8.4.1
Interpretation and Comparison to Related Work
Theorem 4 gives insight into the number of samples needed so that the performance of π̂ is close to that of the optimal policy π∗ . To simplify the problem, we consider the Gaussian policies in Equation (8.13) and assume that there exists Q∗ = N (µ∗ , IdK ) with µ∗ ∈ RdK such that the optimal policy is π∗ = πQ∗ . Also, we let the prior P = N (µ0 , IdK ) and assume that π0 is uniform. This is possible since as we said before, the prior P does not have to depend on the logging policy π0 . Then we have that DKL (Q∗ ∥P) = ∥µ∗ − µ0 ∥2 /2, Bnα (πQ∗ ) = 1−1/K 1−α and Varαn (πQ∗ ) ≤ 2K 2α . The last inequality is not tight but it allows getting an easy-to-interpret term that does not depend on n. Now let ϵ > 2(1 − K α−1 ) for α ∈ [1 − log 2/ log K, 1]. This condition on α ensures that ϵ ∈ [0, 1] and it is mild as α is often close to 1. Then, it holds with high probability that (︂ ∥µ − µ ∥2 + K 2α )︂2 ∗ 0 ˜︁ =⇒ V (π̂) ≥ V (πQ∗ ) − ϵ , n> ϵ − 2(1 − K α−1 ) ˜︁ This gives an intuition on the sample where we omit constant and logarithmic terms in >. complexity for our procedure. In particular, fewer samples are needed in four cases. The first is when ϵ is large, which means that we afford to learn a policy whose performance is far from the optimal one. The second is when the prior P is close to Q∗ , that is when ∥µ∗ − µ0 ∥ is small. This highlights that the choice of the prior P is important. The third is when the second-moment term K 2α is small. The fourth is when the bias Bnα (πQ∗ ) is small. In particular, when α = 1, the bias is 0. In contrast, the second-moment term is minimized in α = 0. This is where the choice of α matters. The proofs of these claims and more detail can be found in Section E.2.4. Our chapter derives a tractable generalization bound for an estimator other than clipped IPS in Equation (8.7), which also holds for the standard IPS in Equation (8.1). The bounds in Swaminathan and Joachims (2015a); London and Sandler (2019); Sakhi et al. 109
(2022) have a multiplicative dependency on the clipping threshold (M or 1/τ in Equation (8.7)). Standard IPS is recovered when M → ∞ (or τ = 0) in which case their bounds are infinite. We successfully avoid any similar dependency on α. Moreover, Swaminathan and Joachims (2015a); London and Sandler (2019) only used their generalization bounds to inspire pessimistic objectives. Although we directly optimize our theoretical bound (Theorem 4) in our experiments, our analysis also inspires a pessimistic objective where we simultaneously penalize the L2 distance, the variance and the bias. That is, we find µ ∈ RdK that maximizes V̂ α (πµ ) − λ1 ∥µ − µ0 ∥2 − λ2 Varαn (πµ ) − λ3 Bnα (πµ ) .
(8.17)
Here λ1 , λ2 and λ3 are tunable hyper-parameters, πµ can be the Gaussian policy in Equagaus tion (8.13), πµ = πµ,1 , with a fixed σ = 1, and µ0 is the mean of the prior P = N (µ0 , IdK ). Existing works either penalize the L2 distance or the variance. For completeness, we also show that this pessimistic objective should be preferred over existing ones in Section E.3.5.
8.4.2
Technical Challenges
London and Sandler (2019); Sakhi et al. (2022) derived PAC-Bayes generalization bounds for the estimator IPS-max in Equation (8.7). Extending their analyses to our case is not straightforward. First, their estimator IPS-max is upper bounded by 1/τ , and thus they relied on traditional techniques for [0, 1]-objectives (Alquier, 2021). In contrast, our objective in Equation (8.14) is not upper-bounded, and controlling it without assuming that the importance weights are bounded is challenging. Moreover, their bounds have a multiplicative dependency on 1/τ , hence they explode as τ → 0. This makes them vacuous for small values of τ and inapplicable to the standard IPS estimator in Equation (8.1) recovered for τ = 0. In contrast, our bound does not have a similar dependency on α and it is also valid for standard IPS recovered for α = 1. Moreover, we derive two-sided inequalities rather than one-sided ones for the important reasons that we priorly discussed. This requires carefully controlling in closed-form the absolute value of the bias. Prior works only used that the bias is negative which was enough to obtain one-sided inequalities. Explaining other challenges requires stating a result that inspired our analysis: Kuzborskij and Szepesvári (2019) derived PAC-Bayes generalization bounds for unbounded losses by only controlling their second moments. Recently, Haddouche and Guedj (2022) proposed a similar result using Ville’s inequality (Bercu and Touati, 2008). Adapting their theorem to our problem is given Proposition 6. We slightly adapt their proof to get a two-sided inequality for a negative loss. The proof is deferred to Section E.2.3. Proposition 6. Let λ > 0, n ≥ 1, δ ∈ (0, 1), α ∈ [0, 1] and let P be a fixed prior on H, then with probability at least 1−δ over draws Dn ∼ µnπ0 , the following holds simultaneously for all posteriors, Q, on H |V α (πQ ) − V̂ α (πQ )| ≤
n DKL (Q∥P) + log 2δ λ ∑︂ πQ (Ai |Xi ) 2 + R λn 2n i=1 π02α (Ai |Xi ) i [︃ ]︃ λ πQ (A|X) 2 + E(X,A,R)∼µπ0 R , 2 π02α (A|X)
110
(8.18)
[︁ πQ (A|X) 2 ]︁ There are two main issues with Proposition 6. First, the term E(X,A,R)∼µπ0 π2α R in 0 (A|X) 2 Equation (8.18) is intractable. One could bound R by 1, but the resulting term will still be intractable due to the expectation over the unknown distribution of contexts ν. Second, we need an upper bound of |V (πQ ) − V̂ α (πQ )| while Proposition 6 only provides one for |V α (πQ )− V̂ α (πQ )|. Thus it remains to quantify the approximation error |V (πQ )−V α (πQ )|. This will also require computing an expectation over X ∼ ν, which is intractable.
8.4.3
Sketch of Proof for Theorem 4
We conclude by showing how the technical challenges above were solved. First, We decompose V (πQ ) − V̂ α (πQ ) as V (πQ ) − V̂ α (πQ ) = I1 + I2 + I3 ,
where
n
1 ∑︂ V (πQ |Xi ) , I1 = V (πQ ) − n i=1 I2 =
n n 1 ∑︂ 1 ∑︂ α V (πQ |Xi ) − V (πQ |Xi ) , n i=1 n i=1 n
1 ∑︂ α V (πQ |Xi ) − V̂ α (πQ ) , I3 = n i=1 where V (πQ |Xi ) = EA∼πQ (·|Xi ) [r(Xi , A)] ,
α
V (πQ |Xi ) = EA∼π0 (·|Xi )
[︂ π (A|X ) Q
i
π0 (A|Xi )
]︂ r(Xi , A) . α
I1 is the estimation error of the empirical mean of the value using n i.i.d. contexts (Xi )i∈[n] . This term is introduced to avoid the intractable expectation over X ∼ ν. Moreover, I2 is the bias term conditioned on the contexts (Xi )i∈[n] and we bound it in closedform. Finally, I3 is the estimation error of the value conditioned on the contexts (Xi )i∈[n] . Again, this conditioning allows us to avoid the intractable expectation over X ∼ ν and to consequently bound |I3 | by tractable terms. First, Alquier (2021, Theorem 3.3) yields that with probability at least 1 − 2δ , it holds for any Q on H that √︄ √ DKL (Q∥P) + log 4 δ n |I1 | ≤ . 2n Also, |I2 | is bounded similarly to Equation (8.9) as n [︁ ]︁ 1 ∑︂ EA∼πQ (·|Xi ) 1 − π01−α (A|Xi ) . |I2 | ≤ n i=1
Bounding |I3 | is achieved by expressing it using martingale difference sequences (fi (Ai , h))i∈[n] that we construct as follows. Let (Fi )i∈{0}∪[n] be a filtration adapted to (Si )i∈[n] where Si = (Aℓ )ℓ∈[i] for any i ∈ [n], we define [︃ ]︃ 1{h(Xi )=A} r(Xi , A) 1{h(Xi )=Ai } Ri fi (Ai , h) = EA∼π0 (·|Xi ) − . α π0 (A|Xi ) π0 (Ai |Xi )α 111
Then we show that for any h ∈ H, (fi (Ai , h))i∈[n] is a martingale difference sequence. After that, we apply Haddouche and Guedj (2022, Theorem 5) and obtain that with probability at least 1 − δ/2, it holds for any Q on H that DKL (Q∥P) + log 4δ λ + Eh∼Q [Varn (h)] , λ 2 [︁ ]︁ ∑︁ ∑︁ where Mn (h) = ni=1 fi (Ai , h) and Varn (h) = ni=1 fi (Ai , h)2 +E fi (Ai , h)2 |Fi−1 . Then notice that Eh∼Q [Mn (h)] can be expressed in terms of I3 as |Eh∼Q [Mn (h)]| ≤
Eh∼Q [Mn (h)] =
n ∑︂
V α (πQ |Xi ) − nV̂ α (πQ ) = nI3 ,
i=1
Moreover, Eh∼Q [Varn (h)] is bounded by [︃ ]︃ n ∑︂ πQ (A|Xi ) πQ (Ai |Xi ) 2 EA∼π0 (·|Xi ) + R . 2α 2α i π (A|X ) π (A |X ) 0 i 0 i i i=1 Thus with probability at least 1 − 2δ , it holds for any Q that [︃ ]︃ n n DKL (Q∥P) + log 4δ λ ∑︂ πQ (Ai |Xi ) 2 λ ∑︂ πQ (A|Xi ) |I3 | ≤ + R + EA∼π0 (·|Xi ) . nλ 2n i=1 π0 (Ai |Xi )2α i 2n i=1 π0 (A|Xi )2α Our result is obtained by bounding |I1 | + |I2 | + |I3 |. One shortcoming of our analysis is that Varαn (πQ ) is not exactly and only resembles the sum of the theoretical and empirical second moments of our estimator. Precisely, the terms πQ /π02α should be πQ2 /π02α . This problem arises due to our definition of the martingale difference sequences (fi (Ai , h))i∈[n] in Equation (8.14). Precisely, in our proof, we compute the square fi (Ai , h)2 . However, the square of an indicator function is the indicator function itself. Thus applying the expectation afterwards, Eh∼Q [fi (Ai , h)2 ], leads to πQ appearing instead of πQ2 . This issue is inherent in the PAC-Bayes formulation and seminal works (London and Sandler, 2019; Sakhi et al., 2022) would suffer the same issue. Solving this would be beneficial and we leave it to future work.
8.5
Experiments for Exponential Smoothing
We briefly present our experiments. More details and discussions can be found in Section E.3. We consider the standard supervised-to-bandit conversion (Agarwal et al., 2014) where we transform a supervised training set Sntr to a logged bandit data Dn as described in Algorithm 3 in Section E.3.1. Here the action space A is the label set and the context space X is the input space. Then, Dn is used to train our policies. After that, we evaluate the value of the learned policies on the supervised test set Sntsts as described in Algorithm 4 in Section E.3.1. Roughly speaking, the resulting value quantifies the ability of the learned policy to predict the true labels of the inputs in the test set. This is our performance metric; the higher the better. We use 4 image classification datasets MNIST (LeCun et al., 1998), FashionMNIST (Xiao et al., 2017), EMNIST (Cohen et al., 2017) and CIFAR100 (Krizhevsky et al., 2009). 112
reward of the learned policy
MNIST, K=10, d=784
0.9
FashionMNIST, K=10, d=784
0.8
0.8
0.7
0.7
0.6
0.6
0.5
0.5
0.4
0.4 0.3
0.3
0.2
0.2
0.1 0.0
0.2
0.4
0.6
0.8
1.0
inverse-temperature parameter ´0 Ours, Gaussian Ours, Mixed-Logit
EMNIST, K=47, d=784
0.6 0.5
0.25
0.4
0.20
0.3
0.15
0.2
0.10
0.1
0.1 0.0
0.2
0.4
0.6
0.8
1.0
London et al., Gaussian London et al., Mixed-Logit
0.05
0.0 0.0
inverse-temperature parameter ´0
CIFAR, K=100, d=2048
0.30
0.2
0.4
0.6
0.8
1.0
inverse-temperature parameter ´0
Sakhi et al. 1, Gaussian Sakhi et al. 1, Mixed-Logit
0.00 0.0
0.2
0.4
0.6
0.8
1.0
inverse-temperature parameter ´0
Sakhi et al. 2, Gaussian Sakhi et al. 2, Mixed-Logit
Logging
Figure 8.2: The reward of the learned policy using one of the baselines with varying quality of the logging policy η0 ∈ [0, 1]. The logging policy is defined as π0 = πηsof in Equation (8.11), where µ0 = (µ0,a )a∈A ∈ RdK 0 ·µ0 and η0 ∈ [0, 1] is the inverse-temperature parameter. The higher η0 , the better the performance of π0 . When η0 = 0, π0 is uniform. The parameters µ0 are learned using 5% of the training set Sntr . In our experiments, we consider both, Gaussian and mixedlogit policies, in Equation (8.12) and Equation (8.13), for which we set the prior as P = N (η0 µ0 , IdK ) and P = N (η0 µ0 , IdK ) × G(0, 1)K , respectively. Given that µ0 are learnt on 5% of Sntr , we train our policies on the remaining 95% portion of Sntr to match our theory that requires the prior to not depend on training data. The policies are trained using Adam (Kingma and Ba, 2014) with a learning rate of 0.1 for 20 epochs. Main results. We compare our bound to those in London and Sandler (2019); Sakhi et al. (2022); discarding the intractable bound in Swaminathan and Joachims (2015a) as it requires computing a covering number. Here we do not include the pessimistic objectives in Swaminathan and Joachims (2015a); London and Sandler (2019) since we directly optimize our bounds. But we make such a comparison in Section E.3.5 for completeness, showing the favorable performance of our bound and the newly proposed pessimistic objective in Equation (8.17). Also, we do not compare to Su et al. (2020); Metelli et al. (2021) since they do not provide generalization guarantees; they focus on estimation accuracy and only propose a heuristic for off-policy learning. However, we still show the favorable performance of our approach in off-policy learning compared to Su et al. (2020); Metelli et al. (2021) in Section E.3.6 for completeness. Prior methods are not named. Thus we refer to them as (Author, Policy) where Author ∈ {Ours, London et al., Sakhi et al. 1, Sakhi et al. 2} and Policy ∈ {Gaussian, Mixed-Logit}. Here Ours, London et al., Sakhi et al. 1 and Sakhi et al. 2 correspond to Theorem 4, London and Sandler (2019, Theorem 1), Sakhi et al. (2022, Proposition 1), and Sakhi et al. (2022, Proposition 3), respectively. Since we have two classes of policies, each bound leads to two baselines. For example, London and Sandler (2019, Theorem 1) leads to (London et al., Gaussian) and (London et al., MixedLogit). More details are provided in Section E.3.3. √ In Figure 8.2, we √ report the value of the learned policies. Here we fix τ = 1/ 4 n ≈ 0.06 and α = 1 − 1/ 4 n ≈ 0.94 so that when n is large enough, both V̂ τ (π) and V̂ α (π) approach V̂ ips (π) (Ionides, 2008). This is because standard IPS should be preferred when n → ∞. To have a fair comparison, we fixed α instead of tuning it in an adaptive fashion as described in Section 8.3.4. However, we also provide the results with an adaptive α 113
in Figure 8.3. Let us start with interpreting Figure 8.2 (with fixed α and τ ). Overall, our method outperforms all the baselines. We also observe that Gaussian policies behave better than mixed-logit policies. However, this is less significant for our method where the performances of both Gaussian and mixed-logit policies are comparable. Moreover, our method reaches the maximum value even when the logging policy has an average performance. In contrast, the baselines only reach their best value when the logging policy is well-performing (η0 ≈ 1), in which case minor to no improvements are made. Finally, the baselines induce a better value when the logging policy is uniform (η0 = 0). But our method has a better value when η0 > 0, which is more common in practice. Larger action spaces. The experiments above did not consider very large values of K. However, Chapter 7 evaluated IPS-based methods, including exponential smoothing and clippped IPS, on datasets with up to one million actions. In those experiments, exponential smoothing outperformed clipped IPS, though the improvements were modest compared to the gains observed here.
MNIST, K=10, d=784
0.9
reward of the learned policy
reward of the learned policy
Choice of hyperparameters. Our choice of τ and α does not affect the above conclusions. In Figure 8.3 (left-hand side), we compare our method with the best baseline, (Sakhi et al. 2) with Gaussian policies, for 20 evenly spaced values of τ ∈ (0, 1) and α ∈ (0, 1). We also include the results using the adaptive tuning procedure of α described in Section 8.3.4 (green curve). This procedure is reliable since the performance with an adaptive α (green curve) is comparable with the best possible choice of α. Also, our method consistently outperforms the best baseline (Sakhi et al. 2) with the best value of τ when the logging policy is not uniform (η0 > 0). Also, there is no very bad choice of α, in contrast with τ = 10−5 (dark blue plot) which led to minimal improvement upon all logging policies. This might be due to the 1/τ dependency in existing bounds.
0.8 0.7 0.6 0.5 Logging Ours Sakhi et al. 2 Ours, Adaptive ®
0.4 0.3 0.2 0.1 0.0
0.2
0.4
0.6
0.8
1.0
inverse-temperature parameter ´0
0.90
MNIST, K=10, d=784
0.85 0.80 0.75 0.70 Modest Logging Good Logging
0.65 0.60 0.0
0.2
0.4
0.6
0.8
1.0
smoothing parameter ®
Figure 8.3: On the left-hand side is the reward of the learned policy with varying τ ∈ (0, 1), α ∈ (0, 1) and η0 ∈ [0, 1], and for an adaptive α using the procedure in Section 8.3.4 (green curve). The blue-to-cyan and red-to-yellow colors correspond to varying values of τ and α, respectively. The lighter the color, the higher the value of τ or α. The green curve corresponds to the reward of the learned policy with an adaptive and data-dependent α (Section 8.3.4). On the right-hand side is the average reward of the learned policies using our method across the modest and good logging groups, η0 ∈ [0, 0.5] (red) and η0 ∈ [0.5, 1] (green), respectively. 114
To see the effect of α, we consider the following experiment. We split the logging policies into two groups. The first is called modest logging which corresponds to logging policies π0 whose η0 is between 0 and 0.5. This group includes the uniform policy and other average-performing policies. The second is called good logging and it includes the logging policies whose η0 is between 0.5 and 1. Then, for each α, we compute the average value of the learned policy, with that value of α, across these two groups. This leads to the two red and green curves in Figure 8.3 (right-hand side). Overall, we observe that α ≈ 0.7 leads to the best performance across the modest logging group. Thus when the performance of the logging policy is bad or average, which is common in practice, importance-weight regularization can be critical. In contrast, when the performance of the logging policy is already good and n is large enough, importance-weight regularization might not be needed and α ≈ 1 would also lead to good performance. This is one of the main strengths of our bound; it holds for the standard IPS recovered with α = 1. This result goes against the belief that clipped IPS should always be preferred to standard IPS. Here, our bound applied to standard IPS outperformed clipping by a large margin when the logging policy is relatively well-performing. Similar results for the other datasets are deferred to Section E.3.4.
8.6
Extension to Other Regularizations
The experiments above demonstrated that exponential smoothing substantially outperforms clipping. However, we compared exponential smoothing with our pessimistic objective against clipping with pessimistic objectives specifically designed for it. This makes it difficult to isolate whether the gains stem from exponential smoothing as a regularization technique or from our pessimistic objective. Moreover, exponential smoothing and clipping are only two instances within a broader class of importance-weight regularizations. While numerous methods have been proposed to stabilize IPS through importance-weight transformations (Bottou et al., 2013; Swaminathan and Joachims, 2015a; Su et al., 2020; Metelli et al., 2021), most focus on estimation accuracy rather than learning performance. As highlighted in Chapter 7, improved estimation does not necessarily yield improved policies, motivating a reassessment of importance-weight regularization specifically within the learning paradigm. Moreover, existing approaches followed a case-by-case basis: each regularization technique comes with its own theoretical analysis and corresponding pessimistic objective. This inconsistency makes it impossible to determine whether empirical improvements arise from the regularizer itself or from its specific objective formulation. This reveals a critical gap: the absence of a unified framework providing principled pessimistic objectives across diverse importance-weight regularizations. We address this by developing a generic PAC-Bayesian generalization bound that applies uniformly to a broad family of regularizations, enabling fair comparison within a single theoretical framework. Recall that the regularized IPS estimator has the form: n
1 ∑︂ Ri ŵ(Xi , Ai ), V̂ (π) = n i=1 115
(8.19)
where ŵ(X, A) are the regularized importance weights. We further assume that ŵ(X, A) = g(π(A|X), π0 (A|X)) for some function g : [0, 1] × [0, 1] → R+ . Examples of ŵ include clipping (Clip) (London and Sandler, 2019), exponential smoothing (ES) (Aouali et al., 2023a), implicit exploration (IX) (Gabbianelli et al., 2024), and harmonic (Har) (Metelli et al., 2021), defined as Clip : ES : IX : Har :
8.6.1
π(a | x) , τ ∈ [0, 1] , max(π0 (a | x), τ ) π(a | x) , α ∈ [0, 1] , ŵ(x, a) = π0 (a | x)α π(a | x) ŵ(x, a) = , γ ∈ [0, 1] , π0 (a | x) + γ w(x, a) , λ ∈ [0, 1] . ŵ(x, a) = (1 − λ)w(x, a) + λ ŵ(x, a) =
(8.20)
Generalization Bounds
⃓ ⃓ ⃓ ⃓ PAC-Bayes theory (Section 8.3) allows bounding ⃓Eθ∼Q [V (πθ ) − V̂ (πθ )]⃓, with n
1 ∑︂ V̂ (πθ ) = ŵθ (Xi , Ai )Ri , n i=1
V (πθ ) = EX∼ν,A∼πθ (·|X) [r(X, A)] ,
where we make the dependence of ŵθ on θ explicit to avoid confusion when taking the expectation Eθ∼Q . Below is our first general result that extends Theorem 4 to any regularization function g, instead of just exponential smoothing. Its proof follows exactly the same techniques we employed for Theorem 4. Theorem 5. Let λ > 0, n ≥ 1, δ ∈ (0, 1), and let P be a fixed prior on Θ. The following inequality holds with probability at least 1 − δ for any distribution Q on Θ: ⃓ ⃓ √︃ kl (π ) kl (π ) λ ⃓ ⃓ 1 Q 2 Q + + Bn (Q) + Varn (Q) , (8.21) ⃓Eθ∼Q [V (πθ ) − V̂ (πθ )]⃓ ≤ 2n nλ 2 √
where kl1 (πQ ) = DKL (Q∥P) + log 4 δ n , kl2 (πQ ) = DKL (Q∥P) + log 4δ , and n ]︁ [︁ 1 ∑︂ Varn (Q) = Eθ∼Q EA∼π0 (·|Xi ) [ŵθ (Xi , A)2 ] + ŵθ (Xi , Ai )2 Ri2 , n i=1 n [︁ ]︁ 1 ∑︂ ∑︂ Bn (Q) = Eθ∼Q |πθ (A|Xi ) − π0 (A|Xi )ŵθ (Xi , A)| . n i=1 A∈A
The terms in the above bound have similar interpretations to those in Theorem 4. Linear vs. non-linear regularization. If ŵ(X, A) is linear in πθ (X, A) (i.e., g linear in its first variable), then V̂ is also linear in πθ , yielding ⃓ ⃓ ⃓ ⃓ ⃓ ⃓ ⃓ ⃓ ⃓Eθ∼Q [V (πθ ) − V̂ (πθ )]⃓ = ⃓V (πQ ) − V̂ (πQ )⃓ , 116
where we define (similar to Section 8.3.2) πQ = Eθ∼Q [πθ ] .
(8.22)
As seen in allows translating the bound in Theorem 5, which ⃓ ⃓ Section 8.3.2, this technique ⃓ ⃓ controls ⃓Eθ∼Q [V (πθ ) − V̂ (πθ )]⃓, into a bound that controls |V (πQ ) − V̂ (πQ )|, the quantity of interest in off-policy learning. The main requirement is to find linear importance-weight regularizations and policies that satisfy Equation (8.22). Fortunately, many importanceweight regularizations, such as Clip, IX, and ES in Equation (8.20), are linear in π, and several practical policies adhere to the formulation in Equation (8.22); see Section 8.3.2 for an in-depth explanation of such policies, including softmax, and Gaussian policies. In Corollary 1, we specialize Theorem 5 to linear importance-weight regularizations of πθ (a|x) the form ŵθ (x, a) = h(π , where h(π0 (a|x)) ≥ π0 (a|x) for all (x, a) ∈ X × A. We 0 (a|x)) additionally assume that the base policies πθ are deterministic, i.e., πθ (a | x) ∈ {0, 1}, which implies πθ (a|x)2 = πθ (a|x). This assumption is only needed here and it is mild: the PAC-Bayes policies πQ defined in Equation (8.22) are mixtures of deterministic policies under Q, and common policy classes such as softmax, mixed-logit, and Gaussian policies admit such representations (Section 8.3.2). Under these assumptions, Theorem 5 yields the following result. Corollary 1. Assume the regularized importance weights can be written as ŵθ (x, a) = πθ (a|x) with h : [0, 1] → R+ verifies h(p) ≥ p for any p ∈ [0, 1]. Moreover, for any h(π0 (a|x)) distribution Q in the parameter space Θ, we define πQ = Eθ∼Q [πθ ] where πθ is binary. Then, let λ > 0, n ≥ 1, δ ∈ (0, 1), and let P be a fixed prior on Θ, The following inequality holds with probability at least 1 − δ for any distribution Q on Θ ⃓ ⃓ √︃ kl (Q) kl2 (Q) λ ⃓ ⃓ 1 + Bn (πQ ) + + Varn (πQ ) , ⃓V (πQ ) − V̂ (πQ )⃓ ≤ 2n nλ 2
(8.23)
where kl1 (Q) and kl2 (Q) are defined in Theorem 5, and ]︃ [︃ n 1 ∑︂ πQ (Ai |Xi ) πQ (A|Xi ) Varn (πQ ) = Ri2 , EA∼π0 (·|Xi ) + 2 2 n i=1 h(π0 (A|Xi )) h(π0 (Ai |Xi )) n
Bn (πQ ) = 1 −
1 ∑︂ ∑︂ πQ (A|Xi ) π0 (A|Xi ) . n i=1 A∈A h(π0 (A|Xi ))
The main benefit of Corollary 1 compared to Theorem 5 is that it eliminates the need for the expectation Eθ∼Q [·], which is now embedded in the definition of policies in Equation (8.22). For example, Corollary 1 allows us to recover the main result of ES above Aouali et al. (2023a) when h(p) = pα , α ∈ [0, 1]. Similarly, we can apply it to IX (Gabbianelli et al., 2024) by setting h(p) = p + γ, γ ≥ 0, and to Clip (London and Sandler, 2019) by setting h(p) = max(p, τ ), τ ∈ [0, 1]. However, if ŵθ (x, a) is not linear in πθ (a|x), then this technique cannot be used, and the original expectation Eθ∼Q [·] in Theorem 5 must be retained. 117
8.6.2
Pessimistic Objectives
Theorem 5 yields two pessimistic objectives. Bound optimization. The first approach directly maximizes the lower bound from Theorem 5: √︃ [︂ ]︂ kl1 (Q) kl2 (Q) λ argmax Eθ∼Q V̂ (πθ ) − − Bn (Q) − − Varn (Q) , (8.24) 2n nλ 2 Q The main challenge is that the objective involves expectations under Q. We address this using the local reparameterization trick (Kingma et al., 2015), which expresses gradients of expectations as expectations of gradients, estimated via Monte Carlo sampling. Specifically, we consider softmax policies πθsof (a|x) from Equation (8.11) and set Q = N (µ, σ 2 IdK ) with learnable parameters µ ∈ RdK and σ > 0. All terms in Equation (8.24) take the form Eθ∼N (µ,σ2 IdK ) [f (πθsof (a|x))], which can be rewritten as: Eθ∼N (µ,σ2 IdK ) [f (πθsof (a|x))] [︃ (︃ )︃]︃ exp(ϕ(x)⊤ µa + σϵa ) = Eϵ∼N (0,∥ϕ(x)∥22 IK ) f ∑︁ . ⊤ a′ ∈A exp(ϕ(x) µa′ + σϵa′ ) This expectation is approximated by sampling ϵi ∼ N (0, ∥ϕ(x)∥22 IK ) and computing the empirical mean; gradients are estimated similarly. However, this approach can exhibit high variance when K is large. For linear importance-weight regularizations, this can be mitigated by optimizing the bound in Corollary 1. For the general case, we propose a practical alternative. Heuristic optimization. The second approach avoids the challenges of direct bound optimization at the cost of additional hyperparameters. Inspired by Theorem 5, we maximize the estimated value penalized by bias, variance, and proximity to the logging policy: ˜ n (πθ ) − λ3 B̃n (πθ ) , V̂ (πθ ) − λ1 ∥θ − θ0 ∥2 − λ2 Var
(8.25)
˜ n (πθ ) and B̃n (πθ ) are the terms inside the expectations in Varn (Q) and Bn (Q), where Var respectively, θ0 parameterizes the logging policy π0 , and λ1 , λ2 , λ3 are hyperparameters. Both objectives in Equations (8.24) and (8.25) are amenable to stochastic gradient optimization and are generic across importance-weight regularizations, enabling fair comparison. We empirically compare these objectives and evaluate different regularization techniques in Section 8.7.
8.7
Experiments for Other Regularizations
We adopt the experimental setting of Section 8.5 and conduct two main experiments. In Section 8.7.1, we fix the importance-weight regularization to Clip (Equation (8.20)) and compare our pessimistic objective against PAC-Bayesian objectives from the literature specifically designed for clipping. The goal is to demonstrate that our objective not 118
only applies more broadly but also outperforms existing alternatives. In Section 8.7.2, having validated our pessimistic objective, we fix it and compare across importance-weight regularizations. The goal is to determine whether any particular regularization technique yields superior off-policy learning performance.
8.7.1
Varying Pessimistic Objectives, Fixed Regularization
We examine the impact of different pessimistic objectives on learned policy performance, π(a|x) in Equafixing the importance-weight regularization to Clip: ŵ(x, a) = max(π 0 (a|x),τ ) √ tion (8.20), with τ = 1/ 4 n following Ionides (2008). For fair comparison, we consider PAC-Bayesian objectives from prior work where the theoretical bound is optimized directly. Specifically, we include London et al. (London and Sandler, 2019, Theorem 1), and two bounds from Sakhi et al. (2022): Sakhi et al. 1 (Sakhi et al., 2022, Proposition 1), based on Catoni (2007), and Sakhi et al. 2 (Sakhi et al., 2022, Proposition 3), a Bernstein-type bound. Since these baselines use linear importance-weight regularization (Section 8.6.1), we compare against our bound in Corollary 1. Following Sakhi et al. (2022); Aouali et al. (2023a), we optimize over Gaussian policies (Equation (8.13)), which perform better in this setting. We also include the logging policy as a baseline.
reward of the learned policy
Figure 8.4 plots the reward of learned policies as a function of logging policy quality η0 ∈ [0, 1]. Our objective outperforms all baselines across a wide range of logging policies. Thus, in addition to being generic across importance-weight regularizations, our approach proves more effective than objectives tailored specifically for Clip. This advantage holds when η0 is not too close to zero: a realistic scenario where logging policies typically outperform uniform random selection. Note that all methods (including ours) improve upon the logging policy (dashed black lines). 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1 0.0
MNIST, K=10, d=784
0.2
0.4
0.6
0.8
1.0
FashionMNIST, K=10, d=784 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1 0.0 0.2 0.4 0.6 0.8 1.0
inverse-temperature parameter ´0 inverse-temperature parameter ´0 Ours Sakhi et al. 1 Logging London et al. Sakhi et al. 2
Figure 8.4: Performance of the learned policy with different PAC-Bayes pessimistic objectives (our Corollary 1 and those in London and Sandler (2019); Sakhi et al. (2022)) using the Clip IPS estimator in Equation (8.20) .
8.7.2
Varying Regularization, Fixed Pessimistic Objective
Having demonstrated the favorable performance of our pessimistic objective, we now compare different importance-weight regularization techniques: Clip, Har, IX, and ES 119
MNIST, K=10, d=784
0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1 0.0
Logging Clip
MNIST, K=10, d=784
0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1 0.0
Logging Har
MNIST, K=10, d=784
0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1 0.0
Logging IX
MNIST, K=10, d=784 Logging ES
inverse-temperature parameter ´0
inverse-temperature parameter ´0
inverse-temperature parameter ´0
inverse-temperature parameter ´0
FashionMNIST, K=10, d=784
FashionMNIST, K=10, d=784
FashionMNIST, K=10, d=784
FashionMNIST, K=10, d=784
0.8 Logging 0.7 Clip 0.6 0.5 0.4 0.3 0.2 0.1 0.0 −0.6 −0.4 −0.2 0.0
0.7 0.6 0.5
0.2
0.4
0.6
inverse-temperature parameter ´0
0.7
Logging Har
0.6 0.5
0.7
Logging IX
0.6 0.5
Logging ES
0.4
0.4
0.4
0.3
0.3
0.3
0.2
0.2
0.2
0.1
0.1
0.1
0.0 −0.6 −0.4 −0.2 0.0
0.0 −0.6 −0.4 −0.2 0.0
0.0 −0.6 −0.4 −0.2 0.0
0.2
0.4
0.6
inverse-temperature parameter ´0
0.2
0.4
0.6
inverse-temperature parameter ´0
0.2
0.4
0.6
inverse-temperature parameter ´0
Av. reward w.r.t. hyperparametersAv. reward w.r.t. hyperparameters
reward reward
0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1 0.0
0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1 0.0
MNIST, K=10, d=784
inverse-temperature parameter ´0
0.7
FashionMNIST, K=10, d=784
0.6 0.5 0.4 0.3 0.2 0.1 0.0 −0.6 −0.4 −0.2 0.0
0.2
0.4
0.6
inverse-temperature parameter ´0
Figure 8.5: Performance of the policy learned by Bound optimization (i.e, Equation (8.24)) for different importance-weight regularizations. The x-axis reflects the quality of the logging policy η0 ∈ [−0.5, 0.5]. In the first four columns, we plot the reward of the learned policy using a fixed importance-weight regularization technique (Clip, Har, IX, or ES as defined in Equation (8.20)) for various values of its hyperparameter within [0, 1]. In the last column, we report the mean reward across these hyperparameter values. (Equation (8.20)). We evaluate both pessimistic objectives from Section 8.6.2, optimizing over softmax policies. For bound optimization, we use Theorem 5 rather than Corollary 1, since Har is non-linear in π. We set λ to its optimal value λ∗ minimizing the bound. While our theory requires λ to be fixed a priori (since λ∗ is data-dependent), we found this yields good empirical performance. For heuristic optimization (Equation (8.25)), we set λ1 = λ2 = λ3 = 10−5 . Figures 8.5 and 8.6 present learned policy rewards as a function of logging policy quality η0 , for bound optimization and heuristic optimization respectively. We vary η0 ∈ [−0.5, 0.5], including logging policies worse than uniform (η0 < 0) to highlight settings requiring stronger regularization, though such scenarios are rarely encountered in practice. Rows correspond to MNIST and FashionMNIST. The first four columns show results for each regularization technique across hyperparameter values in [0, 1]; the last column reports mean reward across hyperparameters for each regularization technique to assess sensitivity to hyperparameters. Bound optimization (Figure 8.5). All regularizations improve over the logging policy (all curves above the dashed baseline), with Har showing less improvement. Clip, IX, and ES achieve comparable performance despite regularizing importance weights differently. These results align with the generality of our bound and suggest that the choice of regularization has limited impact when optimizing the theoretical bound directly. Heuristic optimization (Figure 8.6). Heuristic optimization achieves better performance than bound optimization, likely due to practical limitations of Monte Carlo estimation in high dimensions (Section 8.6.2). The far-right column reveals comparable average performance across regularizations, with two exceptions: ES outperforms the others while Har underperforms. This clarifies our results from Section 8.5: the superior performance of exponential smoothing is more related to our pessimistic objective than the smooth regularization itself. Here, the smooth regularization adds some improvements compared 120
0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1 0.0
Logging Clip
MNIST, K=10, d=784
0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1 0.0
Logging Har
MNIST, K=10, d=784
0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1 0.0
Logging IX
MNIST, K=10, d=784 Logging ES
inverse-temperature parameter ´0
inverse-temperature parameter ´0
inverse-temperature parameter ´0
inverse-temperature parameter ´0
FashionMNIST, K=10, d=784
FashionMNIST, K=10, d=784
FashionMNIST, K=10, d=784
FashionMNIST, K=10, d=784
0.8 Logging 0.7 Clip 0.6 0.5 0.4 0.3 0.2 0.1 0.0 −0.6 −0.4 −0.2 0.0
0.2
0.4
0.6
inverse-temperature parameter ´0
0.8 Logging 0.7 Har 0.6 0.5 0.4 0.3 0.2 0.1 0.0 −0.6 −0.4 −0.2 0.0
0.2
0.4
0.6
inverse-temperature parameter ´0
0.8 Logging 0.7 IX 0.6 0.5 0.4 0.3 0.2 0.1 0.0 −0.6 −0.4 −0.2 0.0
0.2
0.4
0.6
inverse-temperature parameter ´0
0.8 Logging 0.7 ES 0.6 0.5 0.4 0.3 0.2 0.1 0.0 −0.6 −0.4 −0.2 0.0
0.2
0.4
0.6
inverse-temperature parameter ´0
Av. reward w.r.t. hyperparametersAv. reward w.r.t. hyperparameters
reward reward
MNIST, K=10, d=784
0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1 0.0
0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1 0.0
MNIST, K=10, d=784
inverse-temperature parameter ´0 FashionMNIST, K=10, d=784 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1 0.0 −0.6 −0.4 −0.2 0.0 0.2 0.4 0.6 inverse-temperature parameter ´0
Figure 8.6: Performance of the policy learned by Heuristic optimization in Equation (8.25) for different importance-weight regularizations. The x-axis reflects the quality of the logging policy η0 ∈ [−0.5, 0.5]. In the first four columns, we plot the reward of the learned policy using a fixed importance-weight regularization technique (Clip, Har, IX, or ES as defined in Equation (8.20)) for various values of its hyperparameter within [0, 1]. In the last column, we report the mean reward across these hyperparameter values. to others, but the improvments are not significant compared to the improvments we get by simply changing the pessimistic learnin principle even with the standard clipping regularization(Figure 8.4) Larger action spaces. The experiments in this section did not consider very large values of K. However, Chapter 7 evaluated numerous IPS-based methods on datasets with up to one million actions. In those experiments, exponential smoothing outperformed other IPS-based methods, though the improvements were modest. Combined with the results above, this reinforces our conclusion: the choice of the objective has a larger impact on learning performance than the choice of importance-weight regularization, which primarily affects estimation accuracy.
8.8
Conclusion
In this chapter, we investigated importance-weight regularization techniques within the pessimistic paradigm, with particular focus on exponential smoothing as a principled alternative to hard clipping. Our key contributions include: (i) tractable two-sided PAC-Bayes generalization bounds that, unlike prior work, apply to both regularized and standard IPS estimators and are amenable to stochastic gradient optimization; and (ii) the first unified framework for comparing diverse importance-weight regularizations under a common pessimistic objective. This work addresses fundamental theoretical limitations in existing approaches, including the reliance on one-sided inequalities and the misapplication of evaluation bounds in off-policy learning. Rather than using theoretical bounds merely as inspiration for heuristics, we directly optimize them, representing a first step toward making IPS-based pessimism more practical. Our work has two primary limitations. First, the inclusion of empirical bias and variance terms in our bounds makes deriving data-independent suboptimality gaps challenging. 121
Second, two-sided bounds for regularized IPS can be loose as they treat both tails symmetrically, whereas recent work indicates significant asymmetry between lower and upper tails. We address both limitations in Sakhi et al. (2024), which investigates tail-specific bounds to achieve significantly tighter guarantees and sharp suboptimality results. This chapter serves practitioners committed to IPS-based methods, whose appeal is wellfounded: unbiasedness, theoretical guarantees, and the ability to derive principled pessimistic objectives for safe policy learning in high-stakes scenarios. However, from a purely empirical perspective, particularly in very large action spaces, we would favor the PWLL objectives introduced in Chapter 7, which consistently outperform IPS-based methods. The reader can find a direct comparison of these approaches on datasets with up to one million actions in that chapter.
122
Chapter 9
Conclusions and Future Work
This thesis addressed, from both a practical and theoretical perspective, the core obstacle to deploying contextual bandits in modern applications: scalability to large action spaces while maintaining computational tractability. On-Policy Learning (Part I). We introduced structured Bayesian models that enable principled information sharing across actions and derived scalable exploration algorithms. meTS in Chapter 3 couples action parameters through shared latent effects, yielding regret and complexity that scale with an effective number of actions rather than K. dTS in Chapter 4 further develops this idea by using a pre-trained diffusion model to encode richer structure. These algorithms perform well in practice in their theoretical form, without additional tweaks or hyperparameter tuning. Off-Policy Learning (Part II). We tackled both pillars: DM and IPS approaches. sDM in Chapter 6 extends the structured modeling of Part I to the offline regime. We then showed in Chapter 7 that, in large action spaces, optimization matters more than estimation: estimator-based objectives induce highly non-concave landscapes, whereas policy-weighted log-likelihoods produce concave objectives for common policy classes and win decisively at scale. Finally, we developed a principled pessimistic framework for regularized IPS in Chapter 8: smooth importance-weight regularization (exponential smoothing) paired with two-sided PAC-Bayes bounds, and a unified analysis that clarifies when regularization matters and how to compare it across methods. Our additional work on logarithmic smoothing (Sakhi et al., 2024) sharpens the concentration analysis further and yields tighter learning guarantees. Thesis message and practical implications. Scaling to large action spaces causes methods that perform well in small settings (e.g., standard IPS) to fail at scale. This thesis advances three design principles to address this challenge: (i) encode structure to shrink the effective action space; (ii) prioritize objectives with favorable optimization properties over faithful but intractable estimators; and (iii) when relying on IPS with pessimism, couple differentiable importance-weight corrections with theoretically grounded, data-driven bounds amenable to stochastic gradient descent. Together, these principles yield algorithms that are statistically efficient, computationally tractable, and numerically stable. 123
Future work. This thesis opens several promising directions for future research. A key theoretical challenge is to establish robust guarantees under model misspecification, extending the Bayesian analysis of sDM, meTS, and dTS beyond the well-specified setting. For on-policy learning, developing a comprehensive nonlinear diffusion theory for Thompson sampling remains an open problem. In the off-policy setting, future work could investigate the extensions and applications of our methods to LLM and diffusion model fine-tuning, where the objective closely mirrors offline contextual bandit objectives. Moreover, integrating these approaches into large-scale recommender pipelines requires efficient action retrieval, slate constraints, and systems-level optimization. Some of these aspects, such as coupling decision-making with approximate maximum inner product search, were explored in our applied studies (see Additional Contributions in Section 1.3) but omitted from this manuscript. These practical directions have a tangible impact on the online advertising industry and beyond, and are worth pursuing.
124
Chapter A
Supplementary Materials for Chapter 3
A.1
Preliminaries
In this section, we recall some basic properties of matrix operations. (a) The mixed-product property. We have that (A ⊗ B)(C ⊗ D) = AC ⊗ BD for any matrices A, B, C, D such that the products AC and BD exist. (b) Transpose. We have that (A ⊗ B)⊤ = A⊤ ⊗ B⊤ for any matrices A, B. (c) Vectorization. Let A ∈ Rn×m , B ∈ Rm×p , then Vec(AB) = (Ip ⊗ A) Vec(B) = (B⊤ ⊗ In ) Vec(A). (d) For any matrix A, we have that I1 ⊗ A = A. (e) For any positive semi-definite matrices A and B, we have that λ1 (A⊗B) = λ1 (A)λ1 (B). (f) For any matrix A and any positive semi-definite matrix B such that the product A⊤ BA exists, the following inequality holds λ1 (A⊤ BA) ≤ λ1 (B)λ1 (A⊤ A).
A.2
Posterior Derivations
Here we provide the derivations of the effect posterior and action posteriors for the setting presented in Section 3.1.1. Precisely, we present the proof for Proposition 1 in Section A.2.1 and the proof of Proposition 2 in Section A.2.2.
A.2.1
Effect Posterior Derivation
Proof of Proposition 1 (derivation of qt ). First, from basic properties of matrix operations, we observe that the using Kronecker ∑︁ mean of the action parameter can be rewritten Ld products. Specifically, ℓ∈[L] ba,ℓ ψℓ = Γa Ψ, where Ψ = (ψℓ )ℓ∈[L] ∈ R is the concatenated d×Ld effect vector and Γa = b⊤ . Thus, the model in Equation (3.2) (up to round a ⊗ Id ∈ R 125
t ∈ [T ]) can be written as Ψ ∼ N (µΨ , ΣΨ ) , θa | Ψ ∼ N (Γa Ψ, Σ0,a ) ,
∀a ∈ A ,
Ri | Xi , Ai , θ, Ψ ∼ N (Xi⊤ θAi , σ 2 ) ,
∀i ∈ [t − 1] .
(A.1)
Under this model, conditional on (θa )a∈A and (Xi , Ai )i<t , the rewards (Ri )i<t are independent and each Ri depends on Ψ only through θAi . Hence ∫︂ p((Ri )i<t | (Xi , Ai )i<t , Ψ) = p((Ri )i<t | (Xi , Ai )i<t , θ) p(θ | Ψ) dθ. θ∈RdK
Moreover, since p(θ | Ψ) =
∏︁
a∈A p0,a (θa | Ψ) and
p((Ri )i<t | (Xi , Ai )i<t , θ) =
∏︂ ∏︂
N (Ri ; Xi⊤ θa , σ 2 ) =
∏︂
Lt,a (θa ),
a∈A
a∈A i∈St,a
the integral factorizes across arms: p((Ri )i<t | (Xi , Ai )i<t , Ψ) =
∏︂ ∫︂
Lt,a (θa ) p0,a (θa | Ψ) dθa .
a∈A
It follows that the joint effect posterior in round t reads qt (Ψ) ∝ p((Ri )i<t | (Xi , Ai )i<t , Ψ) q0 (Ψ) , ∏︂ ∫︂ Lt,a (θa )p0,a (θa | Ψ) dθa q0 (Ψ) = a∈A
=
θa
∏︂ ∫︂ a∈A ⏞ θa
where Lt,a (θa ) =
∏︁
(A.2)
Lt,a (θa )N (θa ; Γa Ψ, Σ0,a ) dθa N (Ψ; µΨ , ΣΨ ) , ⏟⏟ ⏞
(A.3)
Ia (Ψ)
⊤ 2 i∈St,a N (Ri ; Xi θa , σ ).
We compute the integral term Ia (Ψ) using Lemma 1. Specifically, we obtain that Ia (Ψ) is proportional to a Gaussian density on Ψ, denoted N (Ψ; µ̄t,a , Σ̄t,a ), where
)︁ (︁ −1 ⊤ −1 −1 −1 −1 Σ̄−1 t,a = Γa Σ0,a − Σ0,a (Gt,a + Σ0,a ) Σ0,a Γa , )︁ (︁ −1 −1 −1 µ̄t,a = Σ̄t,a Γ⊤ a Σ0,a (Gt,a + Σ0,a ) Bt,a , and Gt,a and Bt,a are defined in Section 3.2.2. Consequently, the effect posterior qt (Ψ) is proportional to the product of K + 1 multivariate Gaussian distributions: the prior N (µΨ , ΣΨ ) and the likelihood contributions N (µ̄t,a , Σ̄t,a ) for each a ∈ A. Since the product of Gaussians is Gaussian, qt = N (µ̄t , Σ̄t ), where the precision matrix is the sum of the individual precisions: ∑︂ ∑︂ (︁ −1 )︁ −1 −1 −1 ⊤ −1 −1 −1 −1 Σ̄−1 = Σ + Σ̄ = Σ + Γ Σ − Σ (G + Σ ) Σ t,a t t,a a 0,a 0,a 0,a 0,a Γa . Ψ Ψ a∈A
a∈A
126
Using that Γa = b⊤ a ⊗ Id , we rewrite the term inside the sum as: )︁ (︁ −1 )︁ ⊤ (︁ −1 −1 −1 −1 −1 −1 −1 −1 −1 Γ⊤ Σ − Σ (G + Σ ) Σ Γ = (b ⊗ I ) Σ − Σ (G + Σ ) Σ t,a t,a a a d a 0,a 0,a 0,a 0,a 0,a 0,a 0,a 0,a (ba ⊗ Id ) (︁ −1 )︁ ⊤ −1 −1 −1 −1 = (ba ba ) ⊗ Σ0,a − Σ0,a (Gt,a + Σ0,a ) Σ0,a . Similarly, for the mean µ̄t , we have: (︄ µ̄t = Σ̄t Σ−1 Ψ µΨ +
)︄ ∑︂
Σ̄−1 t,a µ̄t,a
a∈A
)︄
(︄ = Σ̄t Σ−1 Ψ µΨ +
∑︂
−1 −1 −1 Γ⊤ a Σ0,a (Gt,a + Σ0,a ) Bt,a
.
a∈A
Using the mixed (Kronecker) product property, (︁ −1 )︁ −1 −1 −1 −1 −1 Γ⊤ a Σ0,a (Gt,a + Σ0,a ) Bt,a = ba ⊗ Σ0,a (Gt,a + Σ0,a ) Bt,a . This recovers the expressions in Proposition 1. To reduce clutter in the following lemma, we fix an action a ∈ A and a round t. We drop the sub-indices a and t, so that we have the following correspondences: Γ ← Γa ,
Σ0 ← Σ0,a ,
N ← Nt,a ,
θ ← θa ,
(Xi , Ri )i∈[N ] ← (Xi , Ri )i∈St,a ,
Lemma 1 (Gaussian posterior update). Let Γ ∈ Rd×Ld , Σ0 ∈ Rd×d , and σ > 0. Consider a dataset of N observations (Xi , Ri )N i=1 . Then, (︄ )︄ ∫︂ ∏︂ N N (Ri ; Xi⊤ θ, σ 2 ) N (θ; ΓΨ, Σ0 ) dθ ∝ N (Ψ; µN , ΣN ) , θ
i=1
where (︂ )︂ (︁ )︁ ⊤ −1 −1 −1 −1 −1 Σ−1 = Γ Σ − Σ G + Σ Σ Γ, N 0 0 0 0 N (︂ )︂ (︁ )︁−1 ⊤ −1 Σ−1 GN + Σ−1 BN . 0 N µN = Γ Σ 0 with GN = σ −2
⊤ −2 i=1 Xi Xi and BN = σ
∑︁N
i=1 Ri Xi .
∑︁N
Proof. Let v = σ −2 and Λ0 = Σ−1 0 . We denote the integral in the lemma by f (Ψ). Completing the square for θ, we have: ]︄ [︄ ∫︂ N 1 v ∑︂ (Ri − Xi⊤ θ)2 − (θ − ΓΨ)⊤ Λ0 (θ − ΓΨ) dθ f (Ψ) ∝ exp − 2 2 θ i=1 (︄ )︄ (︄ N )︄ ∫︂ N [︂ 1 (︂ ∑︂ ∑︂ ∝ exp − θ⊤ v Xi Xi⊤ + Λ0 θ − 2θ⊤ v Ri Xi + Λ0 ΓΨ 2 θ i=1 ⏞ i=1 ⏟⏟ ⏞ VN−1
)︂]︂ + (ΓΨ) Λ0 (ΓΨ) dθ . ⊤
127
To reduce clutter, let GN = v
N ∑︂
Xi Xi⊤ ,
VN = (GN + Λ0 )−1 ,
UN = VN−1 ,
i=1
BN = v
N ∑︂
Ri Xi
and
βN = VN (BN + Λ0 ΓΨ) .
i=1
We have that UN VN = VN UN = Id , and thus ]︃ [︃ ∫︂ )︁ 1 (︁ ⊤ ⊤ ⊤ f (Ψ) ∝ exp − θ UN θ − 2θ UN VN (BN + Λ0 ΓΨ) + (ΓΨ) Λ0 (ΓΨ) dθ , 2 [︃ ]︃ ∫︂θ )︁ 1 (︁ ⊤ ⊤ ⊤ = exp − θ UN θ − 2θ UN βN + (ΓΨ) Λ0 (ΓΨ) dθ , 2 ]︃ [︃ ∫︂θ )︁ 1 (︁ ⊤ ⊤ ⊤ = exp − (θ − βN ) UN (θ − βN ) − βN UN βN + (ΓΨ) Λ0 (ΓΨ) dθ , 2 θ [︃ ]︃ (︁ )︁ 1 ⊤ ⊤ ∝ exp − −βN UN βN + (ΓΨ) Λ0 (ΓΨ) , 2 [︃ )︂]︃ 1 (︂ ⊤ ⊤ = exp − − (BN + Λ0 ΓΨ) VN (BN + Λ0 ΓΨ) + (ΓΨ) Λ0 (ΓΨ) , 2 [︃ ]︃ (︁ ⊤ )︁)︁ 1 (︁ ⊤ ⊤ ⊤ ∝ exp − Ψ Γ (Λ0 − Λ0 VN Λ0 ) ΓΨ − 2Ψ Γ Λ0 VN BN , 2 ]︃ [︃ 1 ⊤ −1 ⊤ −1 = exp − Ψ ΣN Ψ + Ψ ΣN µN , 2 where ⊤ Σ−1 N = Γ (Λ0 − Λ0 VN Λ0 ) Γ , (︁ ⊤ )︁ Σ−1 N µN = Γ Λ0 VN BN .
(A.4)
Plugging the expression of VN concludes the proof.
A.2.2
Action Posterior Derivation
Proof of Proposition 2 (Derivation of pt,a ). This proposition is a direct application of Lemma 2; in which case we get that the posterior pt,a is a multivariate Gaussian distribution N (µ̃t,a , Σ̃t,a ), where −1 Σ̃−1 t,a = Gt,a + Σ0,a , (︄
µ̃t,a = Σ̃t,a Bt,a + Σ−1 0,a
L ∑︂
)︄ ba,ℓ ψt,ℓ
.
ℓ=1
To reduce clutter, we consider a fixed action a ∈ [K] and round t ∈ [T ], and drop subindexing by t and a in Lemma 2. In summary, fix a ∈ [K] and t ∈ [T ] such that we have the following correspondences: bℓ ← ba,ℓ ,
Σ0 ← Σ0,a ,
N ← Nt,a ,
θ ← θa , 128
(Xi , Ri )i∈[N ] ← (Xi , Ri )i∈St,a .
Lemma 2. Consider the following model (︄ L )︄ ∑︂ θ|Ψ∼N bℓ ψℓ , Σ0 , )︁ (︁ ℓ=1 Ri | Xi , θ ∼ N Xi⊤ θ, σ 2 ,
∀i ∈ [N ] . (︂ )︂ Let H = {X1 , R1 , . . . , XN , RN } then we have that p(θ | Ψ, H) = N θ; µ̃N , Σ̃N , where −2 Σ̃−1 N = σ
N ∑︂
Xi Xi⊤ + Σ−1 0 ,
i=1
(︄ µ̃N = Σ̃N
σ −2
N ∑︂
Xi Ri + Σ−1 0
i=1
Proof. Let v = σ −2 ,
L ∑︂
)︄ bℓ ψℓ
.
ℓ=1
Λ0 = Σ−1 0 . Then the action posterior decomposes as
p(θ | Ψ, H) ∝ p((Ri )i∈[N ] | Ψ, θ, (Xi )i∈[N ] )p(θ | Ψ) , = p((Ri )i∈[N ] | θ, (Xi )i∈[N ] )p(θ | Ψ) , =
N ∏︂
N (Ri ; Xi⊤ θ, σ 2 )N (θ;
i=1
L ∑︂
bℓ ψℓ , Σ0 ) ,
ℓ=1
N L ∑︂ 1 (︂ ∑︂ 2 ⊤ ⊤ 2 ⊤ ⊤ = exp − bℓ ψℓ v (Ri − 2Ri Xi θ + (Xi θ) ) + θ Λ0 θ − 2θ Λ0 2 i=1 ℓ=1 )︄ )︄⊤ (︄ L (︄ L )︂]︂ ∑︂ ∑︂ , bℓ ψℓ Λ0 + bℓ ψℓ
[︂
[︄
(︄
(︄
N ∑︂
ℓ=1 N ∑︂
ℓ=1 L ∑︂
1 ∝ exp − θ⊤ (v Xi Xi⊤ + Λ0 )θ − 2θ⊤ v Xi Ri + Λ0 bℓ ψℓ 2 i=1 i=1 ℓ=1 (︃ (︂ )︂−1 )︃ ∝ N θ; µ̃N , Λ̃N , where Λ̃N = v
A.3
⊤ i=1 Xi Xi + Λ0 , and Λ̃N µ̃N = v
∑︁N
∑︁N
i=1 Xi Ri + Λ0
)︄)︄]︄ ,
∑︁L
ℓ=1 bℓ ψℓ .
Regret Proofs
In this section, we establish a more general version of Theorem 1. As explained in Section 3.1.1, we analyze meTS in the linear setting under the assumption of a fully wellspecified model. That is, the true action parameters and rewards are generated according to the same hierarchical structure assumed by meTS: Ψ∗ ∼ N (µΨ , ΣΨ ) , L (︂ ∑︂ )︂ θ∗,a | Ψ∗ ∼ N ba,ℓ ψ∗,ℓ , Σ0,a , ℓ=1 Rt | Xt , At , θ∗ , Ψ∗ ∼ N (Xt⊤ θ∗,At , σ 2 ) ,
129
(A.5) ∀a ∈ A , ∀t ∈ [T ] ,
where the subscript ∗ denotes the true action and latent parameters. To derive the regret bound, we proceed as follows: First, we provide a compact problem formulation in Section A.3.1. Next, we employ total covariance decomposition to derive the posterior covariance of θ∗,a | Ht in Section A.3.2. Finally, we present preliminary eigenvalue results in Section A.3.3 before completing the proof in Section A.3.4.
A.3.1
Problem Reformulation for Regret Analysis
Here, we aim at rewriting Equation (A.5) in a compact form to simplify regret analysis. We first introduce K independent multivariate Gaussian variables Za ∼ N (0, Σ0,a ) for a ∈ [K], and the following matrix Ψ∗,mat = [ψ∗,1 , . . . , ψ∗,L ] ∈ Rd×L . First, we have that Vec(Ψ∗,mat ) = Ψ∗ where Ψ∗ is defined in Equation (A.5). Moreover ∑︁ notice that Lℓ=1 ba,ℓ ψ∗,ℓ = Ψ∗,mat ba , where ba = (ba,ℓ )ℓ∈[L] and thus given matrix Ψ∗,mat we have that ∀a ∈ [K] .
θ∗,a = Ψ∗,mat ba + Za ,
(A.6)
We vectorize Equation (A.6) to obtain θ∗,a = Vec(θ∗,a ) = Vec(Ψ∗,mat ba + Za ) = Vec(Ψ∗,mat ba ) + Za ,
(A.7)
where we used that if X ∈ Rd (a column vector), then X = Vec(X) and that Vec(·) is a linear transformation. Also, we know from (c) in Section A.1 that Vec(AB) = (B⊤ ⊗ In ) Vec(A) for any A ∈ Rn×m , B ∈ Rm×p . Therefore, (A.8)
θ∗,a = Γa Ψ∗ + Za , where Γa = b⊤ a ⊗ Id and we used that Vec(Ψ∗,mat ) = Ψ∗ . It follows that
(A.9)
θ∗,a | Ψ∗ ∼ N (Γa Ψ∗ , Σ0,a ) , This allows us to rewrite our model as a single-parent hierarchical model
A.3.2
(A.10)
Ψ∗ ∼ N (µΨ , ΣΨ ) , θ∗,a | Ψ∗ ∼ N (Γa Ψ∗ , Σ0,a ) ,
∀a ∈ [K] ,
Rt | Xt , At , θ∗ , Ψ∗ ∼ N (Xt⊤ θ∗,At , σ 2 ) ,
∀t ∈ [T ] .
Derivation of cov [θ∗,a | Ht ]
Let Gt,a = σ −2
∑︂
Xi Xi⊤ ,
i∈St,a
Bt,a = σ −2
∑︂ i∈St,a
130
Ri Xi .
Lemma 3 (Expression of cov [θ∗,a | Ht ]). Consider the model in Equation (A.10), then we have ⊤ −1 Σ̂t,a = cov [θ∗,a | Ht ] = Σ̃t,a + Σ̃t,a Σ−1 0,a Γa Σ̄t Γa Σ0,a Σ̃t,a ,
∀a ∈ [K] .
where Σ̄t =
(︂
Σ−1 Ψ +
K ∑︂
ba b⊤ a ⊗
(︁
−1 −1 −1 −1 Σ−1 0,a − Σ0,a (Gt,a + Σ0,a ) Σ0,a
)︁ )︂−1
a=1
Σ̃t,a = Gt,a + Σ−1 0,a (︁
)︁−1
.
Proof. Before proceeding with the proof, we emphasize that cov [Ψ∗ | Ht ] = cov [Ψ | Ht ] = Σ̄t ,
E [Ψ∗ | Ht ] = E [Ψ | Ht ] = µ̄t ,
and cov [θ∗,a | Ψ∗ , Ht ] = cov [θa | Ψ, Ht ] = Σ̃t,a ,
E [θ∗,a | Ψ∗ , Ht ] = E [θa | Ψ, Ht ] = µ̃t,a ,
where the explicit expressions of these covariances and expectations are provided in Proposition 1 and Proposition 2, respectively. These equalities hold because the true action parameters and rewards are assumed to follow the exact generative process defined by the meTS model. ∑︁ Now let Λ0,a = Σ−1 0,a . Proposition 2 and the fact that ℓ∈[L] ba,ℓ ψ∗,ℓ = Γa Ψ∗ where Γa = ⊤ ba ⊗ Id (Section A.3.1) yield cov [θ∗,a | Ψ∗ , Ht ] = (Gt,a + Λ0,a )−1 E [θ∗,a | Ψ∗ , Ht ] = cov [θ∗,a | Ψ∗ , Ht ] (Bt,a + Λ0,a Γa Ψ∗ ) First, given Ht , cov [θ∗,a | Ψ∗ , Ht ] = (Gt,a + Λ0,a )−1 is constant (does not depend on Ψ∗ ). Thus E [cov [θ∗,a | Ψ∗ , Ht ] | Ht ] = cov [θ∗,a | Ψ∗ , Ht ] = (Gt,a + Λ0,a )−1 . In addition, given Ht , both (Gt,a + Λ0,a )−1 and Bt,a are constant. Thus cov [E [θ∗,a | Ψ∗ , Ht ] | Ht ] = cov [cov [θ∗,a | Ψ∗ , Ht ] Λ0,a Γa Ψ∗ | Ht ] −1 = (Gt,a + Λ0,a )−1 Λ0,a Γa cov [Ψ∗ | Ht ] Γ⊤ a Λ0,a (Gt,a + Λ0,a ) −1 = (Gt,a + Λ0,a )−1 Λ0,a Γa Σ̄t Γ⊤ . a Λ0,a (Gt,a + Λ0,a )
Finally, total covariance decomposition (Weiss, 2005) concludes the proof.
A.3.3
Preliminary Eigenvalues Results
Next we present some preliminary upper bounds on the maximum eigenvalues of our covariance matrices. 131
• Definitions: Let λ1,0 = maxa∈[K] λ1 (Σ0,a ) , λd,0 = mina∈[K] λd (Σ0,a ) , λ1,Ψ = λ1 (ΣΨ ) , and κb = maxa∈[K] ∥ba ∥22 . • upper bound of λ1 (Γa Γ⊤ a ): λ1 (Γa Γ⊤ a ) ≤ κb ,
∀a ∈ [K] .
(A.11)
∀a ∈ [K] .
(A.12)
Similarly, we have that λ1 (Γ⊤ a Γa ) ≤ κb , • upper bound of λ1 (Σ̂t,a ): λ1 (Σ̂t,a ) ≤ λ1,0 + 1
λ21,0 λ1,Ψ κb , λ2d,0
∀a ∈ [K] .
(A.13)
1
2 • upper bound of λ1 (ΣΨ2 Σ̄−1 T +1 ΣΨ ): 1
1
2 λ1 (ΣΨ2 Σ̄−1 T +1 ΣΨ ) ≤ 1 + Kλ1,Ψ κb
(︂ 1 λd,0
−
)︂ 1 (︁ )︁ . 1 λ21,0 κσx2T + λd,0
(A.14)
Proof. We start with Equation (A.11). First, recall that Γa = b⊤ a ⊗ Id for any a ∈ [K]. 2 2 ⊤ ⊤ Thus Γa Γa = (ba ⊗ Id )(ba ⊗ Id ) = ∥ba ∥2 Id for any a ∈ [K]. Then λ1 (Γa Γ⊤ a ) = ∥ba ∥2 ≤ κb . ⊤ ⊤ The second result follows from the fact that λ1 (Γa Γa ) = λ1 (Γa Γa ). Now we prove the result in Equation (A.13). This follows from the expression of Σ̂t,a in Lemma 3. Precisely, we have that ⊤ −1 Σ̂t,a = Σ̃t,a + Σ̃t,a Σ−1 0,a Γa Σ̄t Γa Σ0,a Σ̃t,a ,
∀a ∈ [K] .
(︁ )︁−1 where Σ̃t,a = Gt,a + Σ−1 . Thus Weyl’s inequality combined with the properties in 0,a Section A.1 yields that ⊤ −1 λ1 (Σ̂t,a ) ≤ λ1 (Σ̃t,a ) + λ1 (Σ̃t,a )λ1 (Σ−1 0,a )λ1 (Γa Σ̄t Γa )λ1 (Σ0,a )λ1 (Σ̃t,a ) ≤ λ1,0 +
λ21,0 λ1,Ψ κb λ2d,0
⊤ In the last inequality, we used that λ1 (Γa Σ̄t Γ⊤ a ) ≤ λ1 (Σ̄t )λ1 (Γa Γa ), ((f) in Section A.1), 1 λ1 (Σ−1 0,a ) ≤ λd,0 , and λ1 (Σ̃t,a ) ≤ λ1,0 .
Finally, we prove the result in Equation (A.14). First, we rewrite the precision matrix of the effect posterior Σ̄−1 t using the compact notation introduced in Section A.3.1. Precisely, it follows from Equation (A.4) that (i) −1 Σ̄−1 t = ΣΨ +
(ii)
K ∑︂
= Σ−1 Ψ +
(︁ −1 )︁ −1 −1 −1 −1 − Σ (G + Σ ) Σ Γ⊤ Σ t,a 0,a Γa , 0,a 0,a 0,a a
a=1 K ∑︂
(︂ )︂ −1 −1 −1 Γ⊤ Σ − Σ Σ̃ Σ t,a a 0,a 0,a 0,a Γa .
a=1
132
−1 Here, (i) and (ii) are the same; (ii) follows from plugging Σ̃t,a = (Gt,a + Σ−1 in (i). 0,a ) Then we have that 1
1
2 λ1 (ΣΨ2 Σ̄−1 T +1 ΣΨ ) K )︂ (︂ (︂ ∑︂ 1 )︂ 1 −1 −1 −1 2 Γ Σ − Σ Σ̃ Σ Σ = λ1 ILd + ΣΨ2 Γ⊤ a Ψ 0,a T +1,a 0,a 0,a a
a=1
≤ 1 + λ1,Ψ ≤ 1 + λ1,Ψ ≤ 1 + λ1,Ψ ≤ 1 + λ1,Ψ
K ∑︂ a=1 K ∑︂ a=1 K ∑︂ a=1 K ∑︂
(︂ (︂ )︂ )︂ −1 −1 −1 λ1 Γ⊤ Σ − Σ Σ̃ Σ a 0,a 0,a T +1,a 0,a Γa , (︂ )︂ −1 −1 −1 λ1 (Γ⊤ Γ )λ Σ − Σ Σ̃ Σ a a 1 0,a 0,a T +1,a 0,a κb
(︂
(︂ )︂)︂ (︁ −1 )︁ −1 −1 λ1 Σ0,a + λ1 −Σ0,a Σ̃T +1,a Σ0,a ,
(︃ κb
a=1 K ∑︂
)︂ (︂ 1 −1 −1 − λd Σ0,a Σ̃T +1,a Σ0,a λd,0
)︃
(︃
)︃ )︂ (︁ (︁ −1 )︁ (︂ )︁ 1 −1 κb ≤ 1 + λ1,Ψ − λd Σ0,a λd Σ̃T +1,a λd Σ0,a , λd,0 a=1 (︃ K (︂ )︂)︃ ∑︂ 1 1 − λd Σ̃T +1,a ≤ 1 + λ1,Ψ κb λd,0 λ21,0 a=1 (︄ )︄ K ∑︂ 1 1 (︁ )︁ , − 2 = 1 + λ1,Ψ κb −1 λ λ λ G + Σ d,0 1 T +1,a 1,0 0,a a=1 ⎛ ⎞ K ∑︂ 1 1 (︂ )︂ ⎠ κb ⎝ ≤ 1 + λ1,Ψ − κx T 1 λ 2 d,0 λ1,0 σ2 + λd,0 a=1 ⎛ ⎞ 1 1 (︂ )︂ ⎠ . − = 1 + Kλ1,Ψ κb ⎝ λd,0 λ2 κx T + 1 1,0
A.3.4
σ2
λd,0
Regret Proof
Here we prove a more general version of Theorem 1 where we do not assume that the covariance matrices Σ0,a and ΣΨ are diagonal. We still assume that there exists κx > 0 such that ∥Xt ∥22 ≤ κx for any t ∈ [T ]. Theorem 6 (General version of Theorem 1). For any δ ∈ (0, 1), the Bayes regret of meTS in the mixed-effect model in Section 3.1.1 is bounded as BR(T ) ≤
√︁
2T (Ra (T ) + Re (T )) log(1/δ) + cT δ , 133
(A.15)
with c =
√︃
2 κ π x
(︁
λ1,0 +
λ21,0 λ1,Ψ κb )︁ K, λ2d,0
κb = maxa∈[K] ∥ba ∥22 , λ1,0 = maxa∈[K] λ1 (Σ0,a ) , λd,0 =
mina∈[K] λd (Σ0,a ) , λ1,Ψ = λ1 (ΣΨ ) and (︁ κx λ1,0 T κx λ1,0 )︁ , ca = , Ra (T ) = dKca log 1 + 2 σ d log(1 + κxσλ21,0 ) (︁ κx λ1,0 )︁ 2 )︂)︂ (︂ (︂ 1 κ κ λ λ 1 + 1 x b 1,Ψ 1,0 σ2 )︁ , ce = Re (T ) = dLce log 1 + Kκb λ1,Ψ − 2 (︁ κx T 2 λ (︁ )︁ . 1 κ κ λ x b 1,0 1,Ψ λd,0 λ1,0 σ2 + λ 2 log 1 + λ 2 d,0 2 d,0
σ λd,0
2 In particular, the result in Theorem 1 is retrieved when λ1,0 = λd,0 = σ02 , and λ1,Ψ = σΨ .
Proof. Consider our model rewritten in Equation (A.10). Then, the posterior distribution of the action parameter θ∗,a | Ht is a multivariate Gaussian distribution N (µ̂t,a , Σ̂t,a ) for some µ̂t,a ∈ Rd and Σ̂t,a ∈ Rd×d (Lemma 3). Now we let θt,∗ = (Xt⊤ θ∗,a )a∈[K] ∈ RK be the concatenation of the expected rewards of actions in round t. Notice that the context Xt is known in round t, and thus we include it in the history Ht . This is important, with slight abuse of notation, Ht now denotes Ht ← Ht ∪ {Xt }. Then, the joint posterior of the expected rewards, θt,∗ | Ht , is also a multivariate Gaussian N (θ̌t , Σ̌t ) for θ̌t = (Xt⊤ µ̂t,a )a∈[K] ∈ RK and some covariance Σ̌t ∈ RK×K . This follows from the properties of Gaussian distributions (Koller and Friedman, 2009) and the fact that Xt is now included in Ht . Let At ∈ {0, 1}K and At,∗ ∈ {0, 1}K be indicator vectors of the taken action At and optimal action At,∗ , respectively (The vector representations are in bold letters while the integer representations are in regular letters). Then the Bayes regret can be rewritten and consequently decomposed following standard analysis (Russo and Van Roy, 2014) as BR(T ) = E
[︄ T ∑︂
]︄ Xt⊤ θ∗,At,∗ − Xt⊤ θ∗,At
,
(A.16)
t=1 T [︂ ∑︂ ]︂ ⊤ =E A⊤ θ − A θ t,∗ t,∗ t t,∗ ,
(A.17)
t=1
=
T ∑︂
⃓ ]︁]︁ ⃓ ]︁]︁ [︁ [︁ [︁ [︁ ⃓ ⃓ E E A⊤ + E E A⊤ . t,∗ (θt,∗ − θ̌t ) Ht t (θ̌t − θt,∗ ) Ht
t=1
This follows from the fact that θ̌t = (Xt⊤ µ̂t,i )i∈[K] is deterministic given Ht (since Ht now includes Xt ), and that At,∗ and At are i.i.d. given Ht . Moreover, given Ht , θ̌t −θt,∗ is a zeromean multivariate random variable independent of At and thus E[A⊤ t (θ̌t − θt,∗ ) | Ht ] = 0. Therefore, we only need to bound the first term in (A.16). With slight abuse of notation, let A be the set of all possible indicator vectors of actions a ∈ [K]. Precisely, an action a ∈ [K] is also represented by an indicator vector a ∈ A ⊂ {0, 1}K (in bold letter). Then we define the following events }︂ {︂ √︁ ∀δ ∈ (0, 1) , ∀a ∈ A . Et,a (δ) = |a⊤ (θt,∗ − θ̌t )| ≤ 2 log(1/δ)∥a∥Σ̌t , Fix history Ht , we split the expectation over the two complementary events Et,At,∗ (δ) and 134
Ēt,At,∗ (δ), and use the Cauchy-Schwarz inequality to obtain ⃓ ]︁ ⃓ ]︁ √︁ [︁ [︁ ⃓ 2 log(1/δ) E ∥At,∗ ∥Σ̌t ⃓ Ht (A.18) E A⊤ t,∗ (θt,∗ − θ̌t ) Ht ≤ [︁ ⊤ {︁ }︁ ⃓ ]︁ + E At,∗ (θt,∗ − θ̌t )1 Ēt,At,∗ (δ) ⃓ Ht . Now the second term in Equation (A.18) can be bounded as follows. For any a ∈ A , let Za = a⊤ (θt,∗ − θ̌t ). Then we have that {︁ [︁ }︁ ⃓ ]︁ ⃓ Ht (θ − θ̌ )1 E A⊤ Ē (δ) t,∗ t t,A t,∗ t,∗ [︂ {︂ }︂ ⃓ ]︂ √︁ (i) ⃓ = E ZAt,∗ 1 |ZAt,∗ | > 2 log(1/δ)∥At,∗ ∥Σ̌t ⃓ Ht , [︂ {︂ }︂ ⃓ ]︂ (ii) √︁ ⃓ ≤ E |ZAt,∗ |1 |ZAt,∗ | > 2 log(1/δ)∥At,∗ ∥Σ̌t ⃓ Ht , }︂ ⃓ ]︂ [︂ {︂ (iii) ∑︂ √︁ ⃓ ≤ E |Za |1 |Za | > 2 log(1/δ)∥a∥Σ̌t ⃓ Ht , a∈A
]︄ u2 √ ≤ u exp − du , √ 2∥a∥2Σ̌t ∥a∥Σ̌t 2π u= 2 log(1/δ)∥a∥Σ̌t a∈A √︃ [︃ 2 ]︃ ∫︂ ∞ (v) ∑︂ (vi) 2 u 2 ≤ ∥a∥Σ̌t √ u exp − du ≤ λmax,t Kδ . √ 2 π 2π u= 2 log(1/δ) a∈A
(iv) ∑︂
2
[︄
∫︂ ∞
(A.19)
In (i), we simply rewrite the terms using the random variable ZAt,∗ . In (ii), we use bound the expectation of the ranthe fact that ZAt,∗ ≤{︂ |ZAt,∗ |. In (iii), we upper }︂ √︁ dom variable |ZAt,∗ |1 |ZAt,∗ | > 2 log(1/δ)∥At,∗ ∥Σ̌t with the sum of the expectations {︂ }︂ √︁ of |Za |1 |Za | > 2 log(1/δ)∥a∥Σ̌t for a ∈ A since all these random variables are nonnegative. Moreover, (iv) follows from the facts that given Ht , Za ∼ N (0, ∥a∥2Σ̌t ), and that if Z ∼ N (0, σ 2 ), then for any ϵ ≥ 0 , P(|Z| > ϵ) ≤ 2P(Z > ϵ). In (v), we use the change of variables u ← u/∥a∥Σ̌t . Finally, in (vi), we compute the integral and set λmax,t = maxa∈A ∥a∥Σ̌t . We combine Equation (A.18) and Equation (A.19) with the fact that At and At,∗ are i.i.d. given Ht to obtain that √︃ ⃓ ]︁ ⃓ ]︁ √︁ [︁ [︁ ⊤ 2 λmax,t Kδ . (A.20) E At,∗ (θt,∗ − θ̌t ) ⃓ Ht ≤ 2 log(1/δ) E ∥At ∥Σ̌t ⃓ Ht + π The bound in Equation (A.20) holds for any history Ht and thus we take an additional expectation and get that [︄ T ]︄ [︄ T ]︄ √︃ T ∑︂ ∑︂ ∑︂ √︁ 2 ⊤ ⊤ At,∗ θt,∗ − At θt,∗ ≤ 2 log(1/δ) E ∥At ∥Σ̌t + Kδ λmax,t , BR(T ) = E π t=1 t=1 t=1 ⎡⌜ ⎤ √︃ ⃓ T T ⃓ ∑︂ ∑︂ (i) √︁ 2 ≤ 2T log(1/δ) E ⎣⎷ ∥At ∥2Σ̌t ⎦ + Kδ λmax,t , π t=1 t=1 ⌜ [︄ ]︄ ⃓ √︃ T T ⃓ ∑︂ ∑︂ (ii) √︁ 2 ⎷ 2 ≤ 2T log(1/δ) E ∥At ∥Σ̌t + Kδ λmax,t , π t=1 t=1 135
where we use the Cauchy-Schwarz inequality in (i), and (ii) follows from the concavity of the square root. Now note that any a ∈ A is an indicator vector and that Σ̌t is the covariance of the joint posterior of the expected rewards (Xt⊤ θ∗,a )a∈[K] | Ht . Therefore, for any a ∈ A, ∥a∥2Σ̌t = σ̌a2 is the variance of Xt⊤ θ∗,a | Ht . But we know that θ∗,a | Ht is a multivariate Gaussian and its covariance is Σ̂t,a (Lemma 3). Thus the variance of X ⊤ θ∗,a | Ht is σ̌a2 = Xt⊤ Σ̂t,a Xt . It follows that for any a ∈ A , ∥a∥2Σ̌t = Xt⊤ Σ̂t,a Xt = ∥Xt ∥2Σ̂ . In part,a
ticular, ∥At ∥2Σ̌t = Xt⊤ Σ̂t,At Xt . Combining this with Equation (A.13) yields that λmax,t = √︃(︂ √︂ )︂ λ2 λ1,Ψ κb κx . Then maxa∈A ∥a∥Σ̌t = maxa∈A ∥Xt ∥Σ̂t,a ≤ maxa∈A λ1 (Σ̂t,a )κx ≤ λ1,0 + 1,0λ2 d,0 √︃ (︂ )︂ λ21,0 λ1,Ψ κb 2 we let c = π λ1,0 + λ2 κx K which allows us to write d,0
⌜ [︄ ⃓ T ⃓ ∑︂ √︁ ⎷ BR(T ) ≤ 2T log(1/δ) E ∥Xt ∥2Σ̂
t,At
∥Xt ∥2Σ̂t,A = σ 2 t
+ cT δ .
t,At
t=1
√︃ [︂ ∑︁T 2 Now we focus on the the term E t=1 ∥Xt ∥Σ̂
]︄
]︂
(A.21)
that we decompose and bound as
)︂ Xt⊤ Σ̂t,At Xt (i) 2 (︂ −2 ⊤ −1 −1 −2 ⊤ ⊤ = σ σ X Σ̃ X + σ X Σ̃ Σ Γ Σ̄ Γ Σ Σ̃ X t,At t t,At 0,At At t At 0,At t,At t , t t σ2
(ii)
−1 ⊤ ≤ ca log(1 + σ −2 Xt⊤ Σ̃t,At Xt ) + c1 log(1 + σ −2 Xt⊤ Σ̃t,At Σ−1 0,At ΓAt Σ̄t ΓAt Σ0,At Σ̃t,At Xt ) ,
(A.22)
−1 ⊤ where (i) follows from Σ̂t,At = Σ̃t,At + Σ̃t,At Σ−1 0,At ΓAt Σ̄t ΓAt Σ0,At Σ̃t,At , and we use the following inequality in (ii) (︃ )︃ x x u log(1 + x) ≤ max log(1 + x) , log(1 + x) = x= x∈[0,u] log(1 + x) log(1 + x) log(1 + u)
which holds for any x ∈ [0, u], where constants ca and c1 are derived as κx λ1,0 ca = , log(1 + σ −2 κx λ1,0 )
cΨ c1 = , log(1 + σ −2 cΨ )
κx κb λ21,0 λ1,Ψ cΨ = , λ2d,0
The derivation of ca uses that −1 −1 −1 Xt⊤ Σ̃t,At Xt ≤ λ1 (Σ̃t,At )∥Xt ∥2 ≤ λ−1 d (Σ0,At + Gt,At )κx ≤ λd (Σ0,At )κx = λ1 (Σ0,At )κx ≤ λ1,0 κx .
The derivation of c1 follows from −1 −1 ⊤ 2 2 ⊤ Xt⊤ Σ̃t,At Σ−1 0,At ΓAt Σ̄t ΓAt Σ0,At Σ̃t,At Xt ≤ λ1 (Σ̃t,At )λ1 (Σ0,At )λ1 (ΓAt Σ̄t ΓAt )κx ,
λ21 (Σ0,At )λ1,Ψ λ1 (ΓAt Γ⊤ At )κx , 2 λd (Σ0,At ) λ21,0 λ1,Ψ κb κx ≤ . λ2d,0 ≤
136
The first inequality follows from Weyl’s inequality and the fact that λ1 (Σ̄t ) ≤ λ1 (ΣΨ ) = λ1,Ψ and λ1 (Σ̃t,At ) ≤ λ1 (Σ0,At ). Now we focus on bounding the logarithmic terms in Equation (A.22). First Term in Equation (A.22) We first rewrite this term as 1
(i)
1
2 2 log(1 + σ −2 Xt⊤ Σ̃t,At Xt ) = log det(Id + σ −2 Σ̃t,A Xt Xt⊤ Σ̃t,A ), t t
−1 −2 ⊤ = log det(Σ̃−1 t,At + σ Xt Xt ) − log det(Σ̃t,At ) , −1 = log det(Σ̃−1 t+1,At ) − log det(Σ̃t,At ) ,
where (i) follows from the Weinstein–Aronszajn identity. Now note that for any a ̸= At , −1 the arm-a precision does not update at round t, hence Σ̃−1 t+1,a = Σ̃t,a and the increment is zero; therefore we may sum over all a ∈ [K] without changing the value. Then, we sum over all rounds t ∈ [T ], and get a telescoping that leads to T ∑︂
log det(Id + σ
−2
1
1 2
2 )= Σ̃t,At Xt Xt⊤ Σ̃t,A t
T ∑︂
−1 log det(Σ̃−1 t+1,At ) − log det(Σ̃t,At ) ,
t=1
t=1
=
=
T ∑︂ K ∑︂
−1 log det(Σ̃−1 t+1,a ) − log det(Σ̃t,a ) =
t=1 a=1 K ∑︂
K ∑︂ T ∑︂
−1 log det(Σ̃−1 t+1,a ) − log det(Σ̃t,a ) ,
a=1 t=1
−1 log det(Σ̃−1 T +1,a ) − log det(Σ̃1,a ) ,
a=1
(︃ )︃ K K (ii) ∑︂ 1 1 1 1 1 (i) ∑︂ −1 −1 2 2 2 2 d log Tr(Σ0,a Σ̃T +1,a Σ0,a ) = log det(Σ0,a Σ̃T +1,a Σ0,a ) ≤ d a=1 a=1 (︃ )︃ (︃ )︃ K ∑︂ κx λ1 (Σ0,a )T κx λ1,0 T d log 1 + ≤ ≤ Kd log 1 + . 2d 2d σ σ a=1 where (i) follows from the fact that Σ̃1,a = Σ0,a and we use the inequality of arithmetic and geometric means in (ii). Second Term in Equation (A.22) First, we rewrite the covariance matrix of the effect posterior Σ̄t using the compact notation introduced in Section A.3.1. Precisely, it follows from Equation (A.4) that (i) −1 Σ̄−1 t = ΣΨ +
(ii)
K ∑︂
= Σ−1 Ψ +
)︁ (︁ −1 −1 −1 −1 −1 Γ⊤ a Σ0,a − Σ0,a (Gt,a + Σ0,a ) Σ0,a Γa ,
a=1 K ∑︂
(︂ )︂ −1 −1 −1 Γ⊤ Σ − Σ Σ̃ Σ 0,a 0,a t,a 0,a Γa . a
(A.23)
a=1 −1 Recall that (i) and (ii) are the same; (ii) follows from plugging Σ̃t,a = (Gt,a + Σ−1 in 0,a )
137
1
2 (i). Now let u = σ −1 Σ̃t,A Xt . Then it follows from (ii) in Equation (A.23) that t (︂ )︂ −1 −1 −1 −1 −1 −1 −1 ⊤ −2 ⊤ −1 −1 Σ̄−1 − Σ̄ = Γ Σ − Σ ( Σ̃ + σ X X ) Σ − (Σ − Σ Σ̃ Σ ) ΓAt , t t,A t t+1 t At t 0,At 0,At t,At 0,At 0,At 0,At 0,At (︂ )︂ −1 −1 −1 −2 ⊤ −1 = Γ⊤ At Σ0,At (Σ̃t,At − (Σ̃t,At + σ Xt Xt ) )Σ0,At ΓAt , )︂ (︂ 1 1 1 1 −1 −1 −2 2 ⊤ 2 −1 ⊤ 2 2 = ΓAt Σ0,At Σ̃t,At (Id − (Id + σ Σ̃t,At Xt Xt Σ̃t,At ) )Σ̃t,At Σ0,At ΓAt , (︂ )︂ 1 1 −1 −1 ⊤ −1 2 2 Σ Σ̃ = Γ⊤ (I − (I + uu ) ) Σ̃ Σ d At 0,At t,At d t,At 0,At ΓAt , )︃ (︃ 1 1 uu⊤ (i) ⊤ −1 2 2 Σ̃ = ΓAt Σ−1 Σ̃ 0,At t,At t,At Σ0,At ΓAt , ⊤ 1+u u (︃ )︃ Xt Xt⊤ −1 −1 −2 ⊤ Σ̃t,At Σ0,At ΓAt . (A.24) = σ ΓAt Σ0,At Σ̃t,At 1 + u⊤ u
In (i) we use the Sherman-Morrison formula. Moreover, we have that ∥Xt ∥2 ≤ κx . Therefore, 1 + u⊤ u = 1 + σ −2 Xt⊤ Σ̃t,At Xt ≤ 1 + σ −2 κx λ1 (Σ0,At ) ≤ 1 + σ −2 κx λ1,0 = c2 . This allows us to bound the second logarithmic term in Equation (A.22) as −1 ⊤ log(1 + σ −2 Xt⊤ Σ̃t,At Σ−1 0,At ΓAt Σ̄t ΓAt Σ0,At Σ̃t,At Xt ) , (i)
−1 −1 −2 ⊤ ⊤ ≤ c2 log(1 + c−1 2 σ Xt Σ̃t,At Σ0,At ΓAt Σ̄t ΓAt Σ0,At Σ̃t,At Xt ) ,
(ii)
1
1
−1 −1 ⊤ −2 2 ⊤ 2 = c2 log det(ILd + c−1 2 σ Σ̄t ΓAt Σ0,At Σ̃t,At Xt Xt Σ̃t,At Σ0,At ΓAt Σ̄t ) , [︂ ]︂ (iii) −1 −1 −1 −2 ⊤ ⊤ −1 Σ Σ̃ X X = c2 log det(Σ̄−1 + c σ Γ Σ̃ Σ Γ ) − log det( Σ̄ ) , t,A t t,A A t t t t 2 At 0,At t t 0,At ]︃ [︃ (iv) Xt Xt⊤ −1 −1 −2 ⊤ Σ̃t,At Σ−1 ≤ c2 log det(Σ̄−1 + σ Γ Σ Σ̃ t At 0,At t,At 0,At ΓAt ) − log det(Σ̄t ) , 1 + u⊤ u [︁ ]︁ (v) −1 = c2 log det(Σ̄−1 ) − log det( Σ̄ ) . t+1 t
Here (i) follows from the fact that log(1 + x) ≤ c2 log(1 + c−1 2 x) for any x ≥ 0 and c2 ≥ 1. In (ii), we use the Weinstein–Aronszajn identity. In (iii), we use the log product formula ⊤ and the fact that the det is a multiplicative map. In (iv), we use that c−1 2 ≤ 1/(1 + u u). Finally, (v) follows from Equation (A.24). Now we sum over all rounds and get telescoping T ∑︂
−1 ⊤ log(1 + σ −2 Xt⊤ Σ̃t,At Σ−1 0,At ΓAt Σ̄t ΓAt Σ0,At Σ̃t,At Xt ) ,
t=1 1 1 ]︁ [︁ −1 −1 2 2 ≤ c2 log det(Σ̄−1 T +1 ) − log det(Σ̄1 ) = c2 log det(ΣΨ Σ̄T +1 ΣΨ ) , (︃ )︃ (i) 1 1 1 −1 2 2 ≤ c2 Ld log Tr(ΣΨ Σ̄T +1 ΣΨ ) , Ld (︂ (ii) (iii) 1 1 (︁ 1 )︁)︂ 1 2 (︁ )︁ Σ )) 1 + Kκ λ − , ≤ c2 Ld log(λ1 (ΣΨ2 Σ̄−1 ≤ c Ld log 2 b 1,Ψ T +1 Ψ λd,0 λ21,0 κσx2T + λ 1 d,0
138
In (i) we use the inequality of arithmetic and geometric means. In (ii) we bound all eigenvalues in the trace by the maximum eigenvalue. In (iii) we use the result in Equation (A.14). We combine the upper bounds for both logarithmic terms and get
E
[︄ T ∑︂ t=1
]︄ ∥Xt ∥2Σ̂t,A
t
(︁ κx λ1,0 T )︁ ≤ Kdca log 1 + σ2d (︂ (︁ 1 )︁)︂ 1 (︁ )︁ + Ldc1 c2 log 1 + Kκb λ1,Ψ − . λd,0 λ21,0 κσx2T + λ 1 d,0
Finally, we set ce = c1 c2 , which concludes the proof for the general case. To retrieve the 2 result in Theorem 1, we only need to set λ1,0 = λd,0 = σ02 and λ1,Ψ = σΨ since we assumed 2 2 that ΣΨ = σΨ ILd and that Σ0,a = σ0 Id for any a ∈ [K]. In that case, the second term simplifies as (︂ (︂ )︁)︂ (︁ 1 )︁)︂ (︁ 1 1 1 2 )︁ (︁ )︁ − 2 (︁ κx T = log 1 + Kκ σ − , log 1 + Kκb λ1,Ψ b Ψ λd,0 λ1,0 σ2 + λ 1 σ02 σ04 κσx2T + σ12 d,0
0
(︁ 2 = log 1 + Kκb σΨ
A.4
)︁ T κx . 2 T κx σ0 + σ 2
Additional Experiments
We provide additional experiments where we evaluate meTS using synthetic and real-world problems, and compare it to baselines that either ignore or partially use effect parameters. In each plot, we report the averages and standard errors of the quantities. Both settings are described in Section 3.4.
A.4.1
Synthetic Experiments
In Figures A.1 and A.2, we report regret from 12 experiments with horizon T = 5000, where we vary K and d and use both linear and logistic rewards. For the linear setting, we compare meTS-Lin (Section 3.2.2), LinUCB (Abbasi-Yadkori et al., 2011), LinTS (Agrawal and Goyal, 2013a) and HierTS (Hong et al., 2022b). For the logistic setting, we compare meTS-GLM (Section 3.2.3), meTS-Lin (Section 3.2.2), UCB-GLM (Li et al., 2017), GLM-TS (Chapelle and Li, 2012) and HierTS (Hong et al., 2022b). We also include the factored approximation of meTS (meTS-Lin-Fa and meTS-GLM-Fa). In all experiments, we observe that meTS-Lin and meTS-Fa outperform other baselines that ignore the effect parameters or incorporate them partially. We also notice that the gain in performance becomes smaller when K/L decreases. 139
3500
Linear bandit: K = 100, L = 3, d = 2 meTS-Lin meTS-Lin-Fa HierTS LinUCB LinTS
3000 Regret
2500 2000 1500 1000 500 0
8000
0
1000 2000 3000 4000 5000
Linear bandit: K = 100, L = 3, d = 5
7000
Regret
6000 5000 4000 3000 2000 1000 0
0
1000 2000 3000 4000 5000
1800 1600 1400 1200 1000 800 600 400 200 0
4500 4000 3500 3000 2500 2000 1500 1000 500 0
Linear bandit: K = 50, L = 3, d = 2
800
Linear bandit: K = 20, L = 3, d = 2
700 600 500 400 300 200 100 0
1000 2000 3000 4000 5000
Linear bandit: K = 50, L = 3, d = 5
0
1000 2000 3000 4000 5000
0
1800 1600 1400 1200 1000 800 600 400 200 0
0
1000 2000 3000 4000 5000
Linear bandit: K = 20, L = 3, d = 5
0
1000 2000 3000 4000 5000
Figure A.1: Regret of meTS-Lin on synthetic linear bandit problems with varying feature dimension d ∈ {2, 5} and number of actions K ∈ {20, 50, 100}.
Logistic bandit: K = 100, L = 3, d = 2 1400
meTS-GLM meTS-GLM-Fa meTS-Lin GLM-TS UCB-GLM HierTS
Regret
1200 1000 800 600 400
700
400
1400
100 0
1000 2000 3000 4000 5000
Logistic bandit: K = 50, L = 3, d = 5 1200 1000
1200
800
1000 800
600
600
400
400
200
200 0
200
200 0
1000 2000 3000 4000 5000
300
300 100
0
400
500
0
0
1000 2000 3000 4000 5000
0
Logistic bandit: K = 20, L = 3, d = 2 600 500
600
200
Logistic bandit: K = 100, L = 3, d = 5 1600
Regret
Logistic bandit: K = 50, L = 3, d = 2 800
0
1000 2000 3000 4000 5000
0
0
1000 2000 3000 4000 5000
Logistic bandit: K = 20, L = 3, d = 5 900 800 700 600 500 400 300 200 100 0 0 1000 2000 3000 4000 5000
Figure A.2: Regret of meTS-GLM on synthetic logistic bandit problems with varying feature dimension d ∈ {2, 5} and number of actions K ∈ {20, 50, 100}.
A.4.2
MovieLens Experiments
We plot the regret of meTS and the baselines up to T = 5000 rounds in Figures A.3 and A.4. We observe that meTS outperforms the other baselines. This is despite the fact that we did not fine-tune the mixing weights, which attests to the robustness of our approach to model misspecification. Similarly to the synthetic problems, we observe that the gap in performance between meTS and other baselines is less significant when K/L is small. 140
1000
Linear bandit: K = 100, L = 5, d = 2 meTS-Lin meTS-Lin-Fa HierTS LinTS
Regret
800 600 400
250
150
200 100 50
50 0
1000 2000 3000 4000 5000
Linear bandit: K = 100, L = 5, d = 5
0
0
1000 2000 3000 4000 5000
Linear bandit: K = 50, L = 5, d = 5
700
0
600
300
1000
500
250
800
400
200
600
300
150
400
200
100
200
100
50
0
0
1000 2000 3000 4000 5000
0
1000 2000 3000 4000 5000
0
1000 2000 3000 4000 5000
Linear bandit: K = 20, L = 5, d = 5
350
1200
0
Linear bandit: K = 20, L = 5, d = 2
250 200
300
100
0
Regret
350
150
200
1400
Linear bandit: K = 50, L = 5, d = 2
400
0
0
1000 2000 3000 4000 5000
Regret
Figure A.3: Regret of meTS-Lin on the MovieLens dataset with linear rewards and varying feature dimension d ∈ {2, 5} and number of actions K ∈ {20, 50, 100}.
Logistic bandit: K = 100, L = 5, d = 2 250 meTS-GLM 200 meTS-GLM-Fa meTS-Lin 150 HierTS GLM-TS 100
Logistic bandit: K = 20, L = 5, d = 2
150
120 100
100
80 60 40 20
0
1000 2000 3000 4000 5000
0
300
300
250
250 Regret
160 140
Logistic bandit: K = 100, L = 5, d = 5 350
0
1000 2000 3000 4000 5000
Logistic bandit: K = 50, L = 5, d = 5
0
250
0
1000 2000 3000 4000 5000
Logistic bandit: K = 20, L = 5, d = 5
200
200
200
150
150
150
100
100
100
50
50
50 0
Logistic bandit: K = 50, L = 5, d = 2
50
50 0
200
0
1000 2000 3000 4000 5000
0
0
1000 2000 3000 4000 5000
0
0
1000 2000 3000 4000 5000
Figure A.4: Regret of meTS-GLM on the MovieLens dataset with logistic rewards and varying feature dimension d ∈ {2, 5} and number of actions K ∈ {20, 50, 100}.
A.4.3
Robustness to Model Misspecification
We conduct additional synthetic experiments where the hyper-parameters do not match the parameters of the bandit environment to assess the robustness of our approach to misspecification. We provide results for this experiment in Figure A.5. Here we consider the setting described in Section 3.4.1 except that the true hyper-parameters are misspecified as follows. At each run, we sample uniformly 4 misspecification constants c1 , c2 , c3 , and c4 from (0, 2) and set the hyper-parameters of meTS-Lin as c1 ΣΨ , c2 µΨ , c3 Σ0,a , and c4 ba for any a ∈ [K]; where ΣΨ , µΨ , Σ0,a , and ba for a ∈ [K] are the true hyper-parameters. Model 141
misspecification is only applied to meTS-Lin and we refer to it as meTS-Lin-mis. We compare it to meTS-Lin and the other baselines, all with the true hyper-parameters. Although the baselines are not misspecified, meTS-Lin-mis still performs better. meTS-Lin-mis also performs similarly to meTS-Lin (with true hyper-parameters). 3500
Linear bandit: K = 100, L = 3, d = 2 meTS-Lin meTS-Lin-mis HierTS LinUCB LinTS
3000 Regret
2500 2000 1500 1000 500 0
0
1000 2000 3000 4000 5000
1800 1600 1400 1200 1000 800 600 400 200 0
Linear bandit: K = 50, L = 3, d = 2
800
Linear bandit: K = 20, L = 3, d = 2
700 600 500 400 300 200 100 0
1000 2000 3000 4000 5000
0
0
1000 2000 3000 4000 5000
Figure A.5: Regret of misspecified meTS-Lin on synthetic bandit problems with a varying number of actions K. Here, the misspecified meTS, meTS-Lin-mis, is compared to baselines with true hyper-parameters.
A.4.4
Effect of Action Uncertainty
As we mentioned in Section 3.4.1 and predicted by our Bayes regret bound, learning the effect parameters is most beneficial when they are more uncertain than the action parameters. In this section, we support this claim by conducting an experiment where the initial uncertainty of action parameters is greater than the initial uncertainty of the effect parameters. Precisely, we consider the setting described in Section 3.4.1 except that we set ΣΨ = ILd and Σ0,a = 3Id for all a ∈ [K]. We report the results in Figure A.6. By comparing Figure A.6 to Figure 3.2, we observe that meTS-Lin still outperforms the baselines but the gap in performance shrinks when the action parameters are more uncertain than the effect parameters. 4000
Linear bandit: K = 100, L = 3, d = 2 meTS-Lin meTS-Lin-Fa HierTS LinUCB LinTS
3500 Regret
3000 2500 2000 1500 1000 500 0
0
1000 2000 3000 4000 5000
1800 1600 1400 1200 1000 800 600 400 200 0
Linear bandit: K = 50, L = 3, d = 2
800
Linear bandit: K = 20, L = 3, d = 2
700 600 500 400 300 200 100 0
1000 2000 3000 4000 5000
0
0
1000 2000 3000 4000 5000
Figure A.6: Regret of meTS-Lin on synthetic bandit problems with a varying number of actions K, where the action parameters are more uncertain than the effect parameters.
142
Chapter B
Supplementary Materials for Chapter 4
Contents A.1 Preliminaries . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 125 A.2 Posterior Derivations . . . . . . . . . . . . . . . . . . . . . . . . . . . 125 A.2.1
Effect Posterior Derivation . . . . . . . . . . . . . . . . . . . .
125
A.2.2
Action Posterior Derivation . . . . . . . . . . . . . . . . . . .
128
A.3 Regret Proofs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 129 A.3.1
Problem Reformulation for Regret Analysis . . . . . . . . . .
130
A.3.2
Derivation of cov [θ∗,a | Ht ] . . . . . . . . . . . . . . . . . . . .
130
A.3.3
Preliminary Eigenvalues Results . . . . . . . . . . . . . . . . .
131
A.3.4
Regret Proof . . . . . . . . . . . . . . . . . . . . . . . . . . . .
133
A.4 Additional Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . 139
B.1
A.4.1
Synthetic Experiments . . . . . . . . . . . . . . . . . . . . . .
139
A.4.2
MovieLens Experiments . . . . . . . . . . . . . . . . . . . . .
140
A.4.3
Robustness to Model Misspecification . . . . . . . . . . . . . .
141
A.4.4
Effect of Action Uncertainty . . . . . . . . . . . . . . . . . . .
142
Posterior for Linear Diffusion Models
Our posterior approximation builds on the simplified setting where the diffusion model is fully linear, i.e., each link function fℓ is linear in ψℓ . This linear case, studied in our earlier workshop paper (Aouali, 2023), serves as the analytical foundation for our posterior approximation used in the general non-linear case. In Section B.2, we show how the exact posteriors derived in this linear setting inspire our efficient approximation, which extends naturally to practical diffusion models that are typically highly non-linear. 143
B.1.1
Linear Diffusion Models
Here, we assume the link functions fℓ are linear such as fℓ (ψℓ ) = Wℓ ψℓ for ℓ ∈ [L], where Wℓ ∈ Rd×d are known mixing matrices. Then, Equation (4.1) becomes a linear Gaussian system (LGS) (Bishop, 2006) and can be summarized as follows (B.1)
ψL ∼ N (0, ΣL+1 ) , ψℓ−1 | ψℓ ∼ N (Wℓ ψℓ , Σℓ ) , θa | ψ1 ∼ N (W1 ψ1 , Σ1 ) , Rt | Xt , At , θ, (ψℓ )ℓ∈[L] ∼ p(· | Xt ; θAt ) ,
∀ℓ ∈ [L]/{1} , ∀a ∈ [K] , ∀t ∈ [T ] .
This model is important because it yields closed-form posteriors when the reward distribution is linear-Gaussian, i.e., p(· | x; θa ) = N (·; x⊤ θa , σ 2 ). This allows bounding the Bayes regret of sDM. For practice, the posterior expressions are used to motivate efficient approximations for the general case in Equation (4.1) as we show in Section 4.2.1.
B.1.2
Posterior Expressions for Linear Diffusion Models
Recall that the reward distribution is modeled as a generalized linear model (GLM) (McCullagh and Nelder, 1989), allowing for non-linear rewards even when the diffusion links are linear. This non-linearity in the reward distribution prevents closed-form posteriors. However, since the non-linearity arises only through the reward likelihood, we approximate it by a Gaussian, leading to efficient posterior updates that are exact whenever the reward model itself is Gaussian; a special case of the GLM framework. Precisely, let B̂t,a and Ĝt,a denote the MLE (see the remark below for practical considerations) and the Hessian of the negative log-likelihood, respectively: ∑︂ (︁ ∑︂ )︁ ġ Xi⊤ B̂t,a Xi Xi⊤ , B̂t,a = argmax log p(Ri | Xi ; θa ), Ĝt,a = (B.2) θa ∈Rd
i∈St,a
i∈St,a
where St,a = {i ∈ [t − 1] : Ai = a} is the set of rounds in which action a was taken up to round t. We approximate the likelihood as (︂ )︂ ∏︂ p(Ri | Xi ; θa ) ∝ exp − 21 (θa − B̂t,a )⊤ Ĝt,a (θa − B̂t,a ) , (B.3) i∈St,a
which makes all subsequent posteriors Gaussian. Once this approximation is done, all other derivations of the action posterior and latent posteriors are exact. Remark 8. The MLE may be ill-posed. In practice, we maximize an ℓ2 -regularized estimator. Action posterior. The conditional action posterior becomes (︁ )︁ p(θa | ψ1 , Ht,a ) ≈ N θa ; µ̂t,a , Σ̂t,a , with parameters −1 Σ̂−1 t,a = Σ1 + Ĝt,a ,
(︂
µ̂t,a = Σ̂t,a Σ−1 1 W1 ψ1 + Ĝt,a B̂t,a 144
)︂
.
(B.4)
Latent posteriors. For each ℓ ∈ [L] \ {1}, the conditional latent posterior is (︁ )︁ p(ψℓ−1 | ψℓ , Ht ) ≈ N ψℓ−1 ; µ̄t,ℓ−1 , Σ̄t,ℓ−1 , where −1 Σ̄−1 t,ℓ−1 = Σℓ + Ḡt,ℓ−1 ,
(︁ )︁ µ̄t,ℓ−1 = Σ̄t,ℓ−1 Σ−1 ℓ Wℓ ψℓ + B̄t,ℓ−1 .
(B.5)
The top-layer posterior is (︁ )︁ p(ψL | Ht ) ≈ N ψL ; µ̄t,L , Σ̄t,L , with −1 Σ̄−1 t,L = ΣL+1 + Ḡt,L ,
µ̄t,L = Σ̄t,L B̄t,L .
(B.6)
Recursive updates. The matrices Ḡt,ℓ and B̄t,ℓ for ℓ ∈ [L] are defined recursively. The base recursion is Ḡt,1 = W1⊤
K ∑︂ (︁ −1 )︁ −1 Σ1 − Σ−1 W1 , 1 Σ̂t,a Σ1
B̄t,1 = W1⊤ Σ−1 1
a=1
K ∑︂
Σ̂t,a Ĝt,a B̂t,a .
(B.7)
a=1
Then, for ℓ ∈ [L] \ {1}, the recursive step is (︁ )︁ −1 −1 Ḡt,ℓ = Wℓ⊤ Σ−1 − Σ Σ̄ Σ Wℓ , t,ℓ−1 ℓ ℓ ℓ
B̄t,ℓ = Wℓ⊤ Σ−1 ℓ Σ̄t,ℓ−1 B̄t,ℓ−1 .
(B.8)
Discussion. This completes the derivation of the linear posterior approximation. All posteriors are Gaussian and exact whenever the reward distribution follows a linear-Gaussian model, i.e. p(· | x; θa ) = N (·; x⊤ θa , σ 2 ). In this case, the above posterior updates coincide with the exact Bayesian updates, while for general GLMs they serve as efficient and accurate approximations.
B.2
Posterior for Non-Linear Diffusion Models
The general diffusion model (Equation (4.1), which is our case of interest) involves two sources of non-linearity: (i) the reward distribution p(· | x; θ), which may follow a nonlinear generalized linear model (GLM), and (ii) the diffusion links fℓ (ψℓ ), which can be arbitrary non-linear functions. Both sources make the posterior intractable, and therefore two approximations are needed. First approximation (likelihood). We first approximate the reward likelihood by a Gaussian density (as we did above in Equation (B.3)). After this substitution, the model becomes conditionally Gaussian given the latent variables. This step is exact when the reward model is linear-Gaussian, and approximate otherwise. Second approximation (diffusion hierarchy). Even after the likelihood is approximated, the diffusion hierarchy remains non-linear because of the non-linear mappings fℓ . To handle this, we reuse the exact Gaussian posteriors derived for the linear diffusion case (Section B.1.2) and generalize them as follows: 145
• Replace each linear mapping Wℓ ψℓ by its non-linear counterpart fℓ (ψℓ ), which represents the mean of the diffusion prior at layer ℓ. • Remove matrix multiplications involving Wℓ in the recursive updates. This step can be viewed as extending the linear-Gaussian posterior updates to a general non-linear setting. This allows fast sampling and updating of the posterior, without heavy standard posterior approximation techniques. Of course, this is a purely empirical and heuristic based approximation that does not come with guarantees but performs very well in practice. Resulting approximation. The two steps above yield a posterior where each conditional factor p(θa | ψ1 , Ht,a ) and p(ψℓ−1 | ψℓ , Ht ) remains Gaussian with updated means and covariances, while the overall model retains the hierarchical diffusion structure. The approximation satisfies two desirable properties: it exactly recovers the diffusion prior when no data is available, and as more data is observed, the likelihood terms dominate and the prior influence fades naturally.
B.3
Connection to Two-Level Hierarchies
The linear diffusion Equation (B.1) can be marginalized into a 2-level hierarchy using two different strategies. To simplify, we let Σℓ = σℓ2 Id . The first one yields, (B.9)
2 ψL ∼ N (0, σL+1 BL B⊤ L) , θa | ψL ∼ N (ψL , Ω1 ) ,
with Ω1 = σ12 Id +
2 ⊤ ℓ=1 σℓ+1 Bℓ Bℓ and Bℓ =
∑︁L−1
∀a ∈ [K] , i=1 Wi . The second strategy yields,
∏︁ℓ
(B.10)
ψ1 ∼ N (0, Ω2 ) , θa | ψ1 ∼ N (ψ1 , σ12 Id ) ,
∀a ∈ [K] ,
∑︁ 2 where Ω2 = Lℓ=1 σℓ+1 Bℓ B⊤ ℓ . Recently, HierTS (Hong et al., 2022b) was developed for such two-level graphical models, and we call HierTS under Equation (B.9) by HierTS-1 and HierTS under Equation (B.10) by HierTS-2. Then, we start by highlighting the differences between these two variants of HierTS. First, their regret bounds scale as ⌜ ⌜ ⃓ ⃓ L L ⃓ ∑︂ ∑︂ )︁ (︁ )︁ (︁⃓ ⎷ ⎷ 2 2 2 2 T d(K σℓ + LσL+1 , HierTS-2 : Õ T d(Kσ1 + σℓ+1 ) . HierTS-1 : Õ ℓ=1
ℓ=1
When K ≈ L, the regret bounds of HierTS-1 and HierTS-2 are similar. However, when K > L, HierTS-2 outperforms HierTS-1. This is because HierTS-2 puts more uncertainty on a single d-dimensional latent parameter ψ1 , rather than K individual ddimensional action parameters θa . More importantly, HierTS-1 implicitly assumes that action parameters θa are conditionally independent given ψL , which is not true. Consequently, HierTS-2 outperforms HierTS-1. Note that, under the linear diffusion model Equation (B.1), sDM and HierTS-2 have roughly similar regret bounds. Specifically, their regret bounds dependency on K is identical, where both methods involve multiplying K by σ12 , and both enjoy improved performance compared to HierTS-1. That said, note that 146
Theorem 7 and Proposition 7 provide an understanding of how sDM’s regret scales under linear link functions fℓ , and do not say that using sDM is better than using HierTS when the link functions fℓ are linear since the latter can be obtained by a proper marginalization of latent parameters (i.e., HierTS-2 instead of HierTS-1). While such a comparison is not the goal of this work, we still provide it for completeness next. When the mixing matrices Wℓ are dense (i.e., assumption (A5) is not applicable), sDM and HierTS-2 have comparable regret bounds and computational efficiency. However, under the sparsity assumption (A5) and with mixing matrices that allow for conditional independence of ψ1 coordinates given ψ2 , sDM enjoys a computational advantage over HierTS-2. This advantage explains why works focusing on multi-level hierarchies typically benchmark their algorithms against two-level structures akin to HierTS-1, rather than the more competitive HierTS-2. This is also consistent with prior works in Bayesian bandits using multi-level hierarchies, such as Tree-based priors (Hong et al., 2022a), which compared their method to HierTS-1. In line with this, we also compared sDM with HierTS-1 in our experiments. But this is only given for completeness as this is not the aim of Theorem 7 and Proposition 7. More importantly, HierTS is inapplicable in the general case in Equation (4.1) with non-linear link functions since the latent parameters cannot be analytically marginalized.
B.4
Formal Theory
We analyze sDM assuming that: (A0) The true environment parameters θ∗ and ψ∗,ℓ are drawn from the same prior distribution used by sDM, as is standard in Bayes regret analysis (Russo and Van Roy, 2014); we thus use θ∗ and θ (ψ∗,ℓ and ψℓ ) interchangeably throughout. (A1) The rewards are linear p(· | x; θa ) = N (·; x⊤ θa , σ 2 ). (A2) The link functions fℓ are linear such as fℓ (ψℓ ) = Wℓ ψℓ for ℓ ∈ [L], where Wℓ ∈ Rd×d are known mixing matrices. This leads to a structure with L layers of linear Gaussian relationships detailed in Section B.1.1. In particular, this leads to closed-form posteriors given in Section B.1.2 that inspired our approximation and enable theory similar to linear bandits (Agrawal and Goyal, 2013a). However, proofs are not the same, and technical challenges remain (explained in Section B.5). Although our result holds for milder assumptions, we make additional simplifications for clarity and interpretability. We assume that (A3) Contexts satisfy ∥Xt ∥22 = 1 for any t ∈ [T ]. Note that (A3) can be relaxed to any contexts Xt with bounded norms ∥Xt ∥2 . (A4) Mixing matrices and covariances satisfy λ1 (Wℓ⊤ Wℓ ) = 1 for any ℓ ∈ [L] and Σℓ = σℓ2 Id for any ℓ ∈ [L + 1]. In this section, we write Õ for the big-O notation up to polylogarithmic factors. We start by stating our bound for sDM. σ2
2 Theorem 7. Let σmax = maxℓ∈[L+1] 1 + σℓ2 . There exists a constant c > 0 such that for any δ ∈ (0, 1), the Bayes regret of sDM under (A1), (A2), (A3) and (A4) is bounded
147
as ⌜ ⃓ L )︂ ⃓ (︁ ∑︂ )︁ ⎷ lat act BR(T ) ≤ 2T R (T ) + Rℓ log(1/δ) + cT δ , ℓ=1
(︁ T σ12 )︁ σ12 (︂ )︂ , Ract (T ) = c0 dK log 1 + , c = 0 σ2 dσ 2 log 1 + 1 σ2
2ℓ 2 )︁ 2 (︁ σmax σℓ+1 σℓ+1 )︂ , (︂ = c d log 1 + Rlat , c = ℓ ℓ ℓ σ2 σℓ2 log 1 + ℓ+1
(B.11)
σ2
Equation (B.11) holds for any δ ∈√︂ (0, 1). In particular, the term cT δ is constant when )︂ (︂ ∑︁ 2 2ℓ ) , and this dependence σmax δ = 1/T . Then, the bound is Õ T (dKσ12 + d Lℓ=1 σℓ+1 on the horizon T aligns with prior Bayes regret bounds. The bound comprises L + 1 main terms, Ract (T ) and Rlat for ℓ ∈ [L]. First, Ract (T ) relates to action parameters learning, ℓ is associated with conforming to a standard form (Lu and Van Roy, 2019). Similarly, Rlat ℓ learning the ℓ-th latent parameter. To include more structure, we propose the sparsity assumption (A5) Wℓ = (W̄ℓ , 0d,d−dℓ ), where W̄ℓ ∈ Rd×dℓ for any ℓ ∈ [L]. Note that (A5) is not an assumption when dℓ = d for any ℓ ∈ [L]. Notably, (A5) incorporates a plausible structural characteristic that a diffusion model could capture. σ2
2 = maxℓ∈[L+1] 1 + σℓ2 . There exists a constant c > 0 Proposition 7 (Sparsity). Let σmax such that for any δ ∈ (0, 1), the Bayes regret of sDM under (A1), (A2), (A3), (A4) and (A5) is bounded as ⌜ ⃓ L )︂ ⃓ (︁ ∑︂ )︁ ⎷ lat act R̃ℓ log(1/δ) + cT δ , BR(T ) ≤ 2T R (T ) + ℓ=1
(︁ σ12 T σ12 )︁ )︂ , (︂ , c = R̃act (T ) = c0 dK log 1 + 0 σ12 dσ 2 log 1 + σ2
2 )︁ 2 2ℓ (︁ σℓ+1 σℓ+1 σmax (︂ )︂ , Rlat = c d log 1 + , c = ℓ ℓ ℓ ℓ σ2 σℓ2 log 1 + ℓ+1
(B.12)
σ2
From Proposition 7, our bounds scales as ⌜ L (︂⃓ )︂ ⃓ ∑︂ ⎷ 2 2 2ℓ ) . BR(T ) = Õ T (dKσ1 + dℓ σℓ+1 σmax
(B.13)
ℓ=1
B.5
Regret proof
Important notation clarification. Throughout this proof, we operate under the standard Bayesian bandit framework where the true environment parameters θ∗ are drawn 148
from the same prior distribution that sDM uses for posterior inference. Specifically, the true action parameters θ∗,a for a ∈ [K] and the true latent parameters ψ∗,ℓ for ℓ ∈ [L] are assumed to be sampled according to the generative process in Equation (B.1). As a consequence, the true parameters θ∗ and the model parameters θ used in our derivations follow the same distribution, and we use them interchangeably throughout the proof to simplify notation.
B.5.1
Proof Sketch
We start with the following standard lemma upon which we build our analysis (Aouali et al., 2023b). Lemma 4. Assume that p(θa | Ht ) = N (θa ; µ̌t,a , Σ̌t,a ) for any a ∈ [K], then for any δ ∈ (0, 1), ⌜ [︄ ]︄ ⃓ T ⃓ ∑︂ √︁ ∥Xt ∥2Σ̌ + cT δ , where c > 0 is a constant . BR(T ) ≤ 2T log(1/δ)⎷E t,At
t=1
(B.14) Applying Lemma 4 requires proving that the marginal action-posterior densities of θa | Ht in Equation (4.3) are Gaussian and computing their covariances, while we only know the conditional action-posteriors p(θa | ψ1 , Ht ) and latent-posteriors p(ψℓ−1 | ψℓ , Ht ). This is achieved by leveraging the preservation properties of the family of Gaussian distributions (Koller and Friedman, 2009) and the total covariance decomposition (Weiss, 2005) which leads to the next lemma. Lemma 5. Let t ∈ [T ] and a ∈ [K], then the marginal covariance matrix Σ̌t,a reads Σ̌t,a = Σ̂t,a +
∑︂
Pa,ℓ Σ̄t,ℓ P⊤ a,ℓ ,
where
Pa,ℓ = Σ̂t,a Σ−1 1 W1
ℓ−1 ∏︂
Σ̄t,i Σ−1 i+1 Wi+1 .
(B.15)
i=1
ℓ∈[L]
The marginal covariance matrix Σ̌t,a in Equation (B.15) decomposes into L + 1 terms. The first term corresponds to the posterior uncertainty of θa | ψ1 . The remaining L terms capture the posterior uncertainties of ψL and ψℓ−1 | ψℓ for ℓ ∈ [L]/{1}. These are then used to quantify the posterior information gain of latent parameters after one round as follows. Lemma 6 (Posterior information gain). Let t ∈ [T ] and ℓ ∈ [L], then −1 −2 −2ℓ ⊤ ⊤ Σ̄−1 t+1,ℓ − Σ̄t,ℓ ⪰ σ σmax PAt ,ℓ Xt Xt PAt ,ℓ ,
2 where σmax = max 1 + ℓ∈[L+1]
σℓ2 . σ2
(B.16)
Finally, Lemma 5 is used to decompose ∥Xt ∥2Σ̌t,A in Equation (B.14) into L + 1 terms. t Each term is bounded thanks to Lemma 6. This results in the Bayes regret bound in Theorem 7. 149
B.5.2
Proof of lemma 5
In this proof, we heavily rely on the total covariance decomposition (Weiss, 2005). Also, refer to (Hong et al., 2022b, Section 5.2) for a brief introduction to this decomposition. Now, from Equation (B.4), we have that (︂ )︂−1 cov [θa | Ht , ψ1 ] = Σ̂t,a = Ĝt,a + Σ−1 , 1 (︂ )︂ −1 E [θa | Ht , ψ1 ] = µ̂t,a = Σ̂t,a Ĝt,a B̂t,a + Σ1 W1 ψ1 . First, given Ht , cov [θa | Ht , ψ1 ] =
(︂
Ĝt,a + Σ−1 1
)︂−1
is constant. Thus
E [cov [θa | Ht , ψ1 ] | Ht ] = cov [θa | Ht , ψ1 ] =
(︂
Ĝt,a + Σ−1 1
)︂−1
= Σ̂t,a .
In addition, given Ht , Σ̂t,a , Ĝt,a and B̂t,a are constant. Thus [︂ (︂ )︂ ⃓ ]︂ ⃓ cov [E [θa | Ht , ψ1 ] | Ht ] = cov Σ̂t,a Ĝt,a B̂t,a + Σ−1 W ψ ⃓ Ht , 1 1 1 ⃓ ]︂ [︂ ⃓ = cov Σ̂t,a Σ−1 1 W1 ψ1 ⃓ Ht , ⊤ −1 = Σ̂t,a Σ−1 1 W1 cov [ψ1 | Ht ] W1 Σ1 Σ̂t,a , ⊤ −1 = Σ̂t,a Σ−1 1 W1 Σt,1 W1 Σ1 Σ̂t,a ,
where Σt,1 = cov [ψ1 | Ht ] is the marginal posterior covariance of ψ1 . Finally, the total covariance decomposition (Weiss, 2005; Hong et al., 2022b) yields that Σ̌t,a = cov [θa | Ht ] = E [cov [θa | Ht , ψ1 ] | Ht ] + cov [E [θa | Ht , ψ1 ] | Ht ] , ⊤ −1 = Σ̂t,a + Σ̂t,a Σ−1 1 W1 Σt,1 W1 Σ1 Σ̂t,a ,
(B.17)
However, Σt,1 = cov [ψ1 | Ht ] is different from Σ̄t,1 = cov [ψ1 | Ht , ψ2 ] that we already derived in Equation (B.5). Thus we do not know the expression of Σt,1 . But we can use the same total covariance decomposition trick to find it. Precisely, let Σt,ℓ = cov [ψℓ | Ht ] for any ℓ ∈ [L]. Then we have that (︁ )︁−1 Σ̄t,1 = cov [ψ1 | Ht , ψ2 ] = Σ−1 + Ḡ , t,1 2 (︂ )︂ µ̄t,1 = E [ψ1 | Ht , ψ2 ] = Σ̄t,1 Σ−1 W ψ + B̄ 2 2 t,1 . 2 (︁ )︁−1 First, given Ht , cov [ψ1 | Ht , ψ2 ] = Σ−1 is constant. Thus 2 + Ḡt,1 E [cov [ψ1 | Ht , ψ2 ] | Ht ] = cov [ψ1 | Ht , ψ2 ] = Σ̄t,1 . In addition, given Ht , Σ̄t,1 , Σ̃t,1 and B̄t,1 are constant. Thus [︂ (︂ )︂ ⃓ ]︂ ⃓ cov [E [ψ1 | Ht , ψ2 ] | Ht ] = cov Σ̄t,1 Σ−1 W ψ + B̄ 2 2 t,1 ⃓ Ht , 2 ⃓ ]︁ [︁ ⃓ = cov Σ̄t,1 Σ−1 2 W2 ψ2 Ht , ⊤ −1 = Σ̄t,1 Σ−1 2 W2 cov [ψ2 | Ht ] W2 Σ2 Σ̄t,1 , ⊤ −1 = Σ̄t,1 Σ−1 2 W2 Σt,2 W2 Σ2 Σ̄t,1 .
150
Finally, total covariance decomposition (Weiss, 2005; Hong et al., 2022b) leads to Σt,1 = cov [ψ1 | Ht ] = E [cov [ψ1 | Ht , ψ2 ] | Ht ] + cov [E [ψ1 | Ht , ψ2 ] | Ht ] , ⊤ −1 = Σ̄t,1 + Σ̄t,1 Σ−1 2 W2 Σt,2 W2 Σ2 Σ̄t,1 .
Now using the techniques, this can be generalized using the same technique as above to −1 ⊤ Σt,ℓ = Σ̄t,ℓ + Σ̄t,ℓ Σ−1 ℓ+1 Wℓ+1 Σt,ℓ+1 Wℓ+1 Σℓ+1 Σ̄t,ℓ ,
Then, by induction, we get that ∑︂ Σt,1 = P̄ℓ Σ̄t,ℓ P̄⊤ ℓ ,
∀ℓ ∈ [L − 1] .
∀ℓ ∈ [L − 1] ,
ℓ∈[L]
where we use that by definition Σt,L = cov [ψL | Ht ] = Σ̄t,L and set P̄1 = Id and P̄ℓ = ∏︁ℓ−1 −1 i=1 Σ̄t,i Σi+1 Wi+1 for any ℓ ∈ [L]/{1}. Plugging this in Equation (B.17) leads to ∑︂ ⊤ ⊤ −1 Σ̌t,a = Σ̂t,a + Σ̂t,a Σ−1 1 W1 P̄ℓ Σ̄t,ℓ P̄ℓ W1 Σ1 Σ̂t,a , ℓ∈[L]
= Σ̂t,a +
∑︂
−1 ⊤ Σ̂t,a Σ−1 1 W1 P̄ℓ Σ̄t,ℓ (Σ̂t,a Σ1 W1 ) ,
ℓ∈[L]
= Σ̂t,a +
∑︂
Pa,ℓ Σ̄t,ℓ P⊤ a,ℓ ,
ℓ∈[L] −1 where Pa,ℓ = Σ̂t,a Σ−1 1 W1 P̄ℓ = Σ̂t,a Σ1 W1
B.5.3
−1 i=1 Σ̄t,i Σi+1 Wi+1 .
∏︁ℓ−1
Proof of lemma 6
We prove this result by induction. We start with the base case when ℓ = 1. 1
2 Xt From the expression of Σ̄t,1 in Equation (B.5), we (I) Base case. Let u = σ −1 Σ̂t,A t have that (︂ )︂ −1 −1 ⊤ −1 −1 −2 ⊤ −1 −1 −1 −1 −1 Σ̄−1 − Σ̄ = W Σ − Σ ( Σ̂ + σ X X ) Σ − (Σ − Σ Σ̂ Σ ) W1 , t t,A t t+1,1 t,1 1 1 1 t 1 1 1 1 t,At (︂ )︂ −1 −2 ⊤ −1 −1 = W1⊤ Σ−1 W1 , 1 (Σ̂t,At − (Σ̂t,At + σ Xt Xt ) )Σ1 (︂ )︂ 1 1 1 1 −2 2 ⊤ 2 −1 −1 2 2 = W1⊤ Σ−1 Σ̂ (I − (I + σ Σ̂ X X Σ̂ ) ) Σ̂ Σ W1 , d d t 1 t t,At t,At t,At t,At 1 (︂ )︂ 1 1 ⊤ −1 −1 2 2 = W1⊤ Σ−1 W1 , 1 Σ̂t,At (Id − (Id + uu ) )Σ̂t,At Σ1 (︃ )︃ 1 1 uu⊤ (i) −1 2 ⊤ −1 2 W1 , = W1 Σ1 Σ̂t,At Σ̂ Σ 1 + u⊤ u t,At 1 Xt Xt⊤ (ii) −2 = σ W1⊤ Σ−1 Σ̂ Σ̂t,At Σ−1 (B.18) t,At 1 1 W1 . 1 + u⊤ u
−1 In (i) we use the Sherman-Morrison formula. Note that (ii) says that Σ̄−1 t+1,1 − Σ̄t,1 is one-rank which we will also need in induction step. Now, we have that ∥Xt ∥2 = 1. Therefore, 2 1 + u⊤ u = 1 + σ −2 Xt⊤ Σ̂t,At Xt ≤ 1 + σ −2 λ1 (Σ1 )∥Xt ∥2 = 1 + σ −2 σ12 ≤ σmax ,
151
2 2 ≥ 1 + σ −2 σ12 . in Lemma 6, we have that σmax where we use that by definition of σmax 1 −2 Therefore, by taking the inverse, we get that 1+u⊤ u ≥ σmax . Combining this with Equation (B.18) leads to −1 −2 −2 ⊤ −1 ⊤ −1 Σ̄−1 t+1,1 − Σ̄t,1 ⪰ σ σmax W1 Σ1 Σ̂t,At Xt Xt Σ̂t,At Σ1 W1
Noticing that PAt ,1 = Σ̂t,At Σ−1 1 W1 concludes the proof of the base case when ℓ = 1. −1 (II) Induction step. Let ℓ ∈ [L]/{1} and suppose that Σ̄−1 t+1,ℓ−1 − Σ̄t,ℓ−1 is one-rank and that it holds for ℓ − 1 that −1 −2 −2(ℓ−1) ⊤ 2 Σ̄−1 PAt ,ℓ−1 Xt Xt⊤ PAt ,ℓ−1 , where σmax = max 1 + σ −2 σℓ2 . t+1,ℓ−1 − Σ̄t,ℓ−1 ⪰ σ σmax ℓ∈[L+1]
−1 Then, we want to show that Σ̄−1 t+1,ℓ − Σ̄t,ℓ is also one-rank and that it holds that 2 where σmax = max 1 + σ −2 σℓ2 .
−1 −2 −2ℓ ⊤ ⊤ Σ̄−1 t+1,ℓ − Σ̄t,ℓ ⪰ σ σmax PAt ,ℓ Xt Xt PAt ,ℓ ,
ℓ∈[L+1]
This is achieved as follows. Define the precision increment at level ℓ − 1 by −1 ∆t,ℓ−1 := Σ̄−1 t+1,ℓ−1 − Σ̄t,ℓ−1 .
By the induction hypothesis, ∆t,ℓ−1 is rank-one PSD, hence there exists u ∈ Rd such that ∆t,ℓ−1 = uu⊤ . Using Equation (B.8), we have for any s ∈ {t, t + 1}: (︁ −1 )︁ −1 −1 −1 −1 ⊤ Σ̄−1 = Σ + Ḡ = Σ + W Σ − Σ Σ̄ Σ Wℓ . s,ℓ s,ℓ−1 ℓ s,ℓ ℓ+1 ℓ+1 ℓ ℓ ℓ Therefore, −1 Σ̄−1 t+1,ℓ − Σ̄t,ℓ = Ḡt+1,ℓ − Ḡt,ℓ (︁ )︁ = Wℓ⊤ Σ−1 Σ̄t,ℓ−1 − Σ̄t+1,ℓ−1 Σ−1 ℓ ℓ Wℓ . −1 ⊤ Since Σ̄−1 t+1,ℓ−1 = Σ̄t,ℓ−1 + uu , Sherman–Morrison yields
(︁ −1 )︁ Σ̄t,ℓ−1 uu⊤ Σ̄t,ℓ−1 ⊤ −1 . Σ̄t+1,ℓ−1 = Σ̄t,ℓ−1 + uu = Σ̄t,ℓ−1 − 1 + u⊤ Σ̄t,ℓ−1 u Hence, Σ̄t,ℓ−1 − Σ̄t+1,ℓ−1 =
Σ̄t,ℓ−1 uu⊤ Σ̄t,ℓ−1 , 1 + u⊤ Σ̄t,ℓ−1 u
and plugging this back gives −1 ⊤ −1 Σ̄−1 t+1,ℓ − Σ̄t,ℓ = Wℓ Σℓ Σ̄t,ℓ−1
uu⊤ Σ̄t,ℓ−1 Σ−1 ℓ Wℓ . ⊤ 1 + u Σ̄t,ℓ−1 u
In particular, this increment is rank-one PSD. 152
However, it follows from the induction hypothesis that −1 −2 −2(ℓ−1) ⊤ uu⊤ = Σ̄−1 PAt ,ℓ−1 Xt Xt⊤ PAt ,ℓ−1 . t+1,ℓ−1 − Σ̄t,ℓ−1 ⪰ σ σmax
Therefore, −1 ⊤ −1 Σ̄−1 t+1,ℓ − Σ̄t,ℓ = Wℓ Σℓ Σ̄t,ℓ−1
uu⊤ Σ̄t,ℓ−1 Σ−1 ℓ Wℓ , ⊤ 1 + u Σ̄t,ℓ−1 u −2(ℓ−1)
⪰ Wℓ⊤ Σ−1 ℓ Σ̄t,ℓ−1
σ −2 σmax
⊤ P⊤ At ,ℓ−1 Xt Xt PAt ,ℓ−1 Σ̄t,ℓ−1 Σ−1 ℓ Wℓ , 1 + u⊤ Σ̄t,ℓ−1 u
−2(ℓ−1)
σ −2 σmax −1 ⊤ = W⊤ Σ−1 Σ̄t,ℓ−1 P⊤ At ,ℓ−1 Xt Xt PAt ,ℓ−1 Σ̄t,ℓ−1 Σℓ Wℓ , 1 + u⊤ Σ̄t,ℓ−1 u ℓ ℓ −2(ℓ−1)
σ −2 σmax = P⊤ Xt Xt⊤ PAt ,ℓ . 1 + u⊤ Σ̄t,ℓ−1 u At ,ℓ Finally, we use that 1 + u⊤ Σ̄t,ℓ−1 u ≤ 1 + ∥u∥22 λ1 (Σ̄t,ℓ−1 ) ≤ 1 + σ −2 σℓ2 . Here we use that ∥u∥22 ≤ σ −2 , which can also be proven by induction, and that λ1 (Σ̄t,ℓ−1 ) ≤ σℓ2 , which follows from the expression of Σ̄t,ℓ−1 in Section B.1.2. Therefore, we have that −2(ℓ−1)
−1 Σ̄−1 t+1,ℓ − Σ̄t,ℓ ⪰
σ −2 σmax ⊤ P⊤ At ,ℓ Xt Xt PAt ,ℓ , ⊤ 1 + u Σ̄t,ℓ−1 u −2(ℓ−1)
σ −2 σmax ⊤ ⪰ P⊤ At ,ℓ Xt Xt PAt ,ℓ , 2 −2 1 + σ σℓ −2ℓ ⊤ ⪰ σ −2 σmax PAt ,ℓ Xt Xt⊤ PAt ,ℓ , 2 = maxℓ∈[L+1] 1 + σ −2 σℓ2 . This where the last inequality follows from the definition of σmax concludes the proof.
B.5.4
Proof of theorem 7
We start with the following standard result which we borrow from (Hong et al., 2022a; Aouali et al., 2023b), ⌜ [︄ ⃓ T ⃓ ∑︂ √︁ ⎷ BR(T ) ≤ 2T log(1/δ) E ∥Xt ∥2Σ̌ t=1
]︄
t,At
+ cT δ ,
where c > 0 is a constant . (B.19)
Then we use Lemma 5 and express the marginal covariance Σ̌t,At as Σ̌t,a = Σ̂t,a +
∑︂
Pa,ℓ Σ̄t,ℓ P⊤ a,ℓ ,
where Pa,ℓ = Σ̂t,a Σ−1 1 W1
ℓ−1 ∏︂ i=1
ℓ∈[L]
153
Σ̄t,i Σ−1 i+1 Wi+1 .
(B.20)
Therefore, we can decompose ∥Xt ∥2Σ̌t,A as t
∑︂ )︁ X ⊤ Σ̌t,A Xt (i) 2 (︁ −2 ⊤ = σ σ Xt Σ̂t,At Xt + σ −2 ∥Xt ∥2Σ̌t,A = σ 2 t 2 t Xt⊤ PAt ,ℓ Σ̄t,ℓ P⊤ At ,ℓ Xt , t σ ℓ∈[L] (ii)
≤ c0 log(1 + σ −2 Xt⊤ Σ̂t,At Xt ) +
∑︂
cℓ log(1 + σ −2 Xt⊤ PAt ,ℓ Σ̄t,ℓ P⊤ At ,ℓ Xt ) ,
(B.21)
ℓ∈[L]
where (i) follows from Equation (B.20), and we use the following inequality in (ii) (︃ )︃ x u x log(1 + x) ≤ max log(1 + x) = log(1 + x) , x= x∈[0,u] log(1 + x) log(1 + x) log(1 + u) which holds for any x ∈ [0, u], where constants c0 and cℓ are derived as c0 =
σ12 σ2
,
cℓ =
log(1 + σ12 )
2 σℓ+1 σ2
.
log(1 + σℓ+1 2 )
The derivation of c0 uses that −1 −1 −1 2 Xt⊤ Σ̂t,At Xt ≤ λ1 (Σ̂t,At )∥Xt ∥2 ≤ λ−1 d (Σ1 + Gt,At ) ≤ λd (Σ1 ) = λ1 (Σ1 ) = σ1 .
The derivation of cℓ follows from ⊤ 2 2 Xt⊤ PAt ,ℓ Σ̄t,ℓ P⊤ At ,ℓ Xt ≤ λ1 (PAt ,ℓ PAt ,ℓ )λ1 (Σ̄t,ℓ )∥Xt ∥ ≤ σℓ+1 .
Therefore, from Equation (B.21) and Equation (B.19), we get that T (︂ [︂ ∑︂ √︁ log(1 + σ −2 Xt⊤ Σ̂t,At Xt ) BR(T ) ≤ 2T log(1/δ) E c0 t=1
+
∑︂
cℓ
ℓ∈[L]
T ∑︂
]︂)︂ 12 log(1 + σ −2 Xt⊤ PAt ,ℓ Σ̄t,ℓ P⊤ X ) + cT δ At ,ℓ t
(B.22)
t=1
Now we focus on bounding the logarithmic terms in Equation (B.22). (I) First term in Equation (B.22) We first rewrite this term as 1
1
(i)
2 2 log(1 + σ −2 Xt⊤ Σ̂t,At Xt ) = log det(Id + σ −2 Σ̂t,A Xt Xt⊤ Σ̂t,A ), t t
−1 −1 −1 −2 ⊤ = log det(Σ̂−1 t,At + σ Xt Xt ) − log det(Σ̂t,At ) = log det(Σ̂t+1,At ) − log det(Σ̂t,At ) ,
where (i) follows from the Weinstein-Aronszajn identity. Then we sum over all rounds t ∈ [T ], and get a telescoping T ∑︂
log det(Id + σ
t=1
−2
1 2
1
2 Σ̂t,At Xt Xt⊤ Σ̂t,A )= t
T ∑︂
−1 log det(Σ̂−1 t+1,At ) − log det(Σ̂t,At ) ,
t=1
=
=
T ∑︂ K ∑︂ t=1 a=1 K ∑︂
−1 log det(Σ̂−1 t+1,a ) − log det(Σ̂t,a ) =
(i)
−1 log det(Σ̂−1 T +1,a ) − log det(Σ̂1,a ) =
a=1
K ∑︂ T ∑︂
a=1 t=1 K ∑︂
−1 log det(Σ̂−1 t+1,a ) − log det(Σ̂t,a ) ,
1
1
2 log det(Σ12 Σ̂−1 T +1,a Σ1 ) ,
a=1
154
where (i) follows from the fact that Σ̂1,a = Σ1 . Now we use the inequality of arithmetic and geometric means and get T ∑︂
1
1
2 2 log det(Id + σ −2 Σ̂t,A Xt Xt⊤ Σ̂t,A )= t t
K ∑︂ a=1 K ∑︂
t=1
1
1
2 log det(Σ12 Σ̂−1 T +1,a Σ1 ) ,
(︃
)︃ 1 1 1 −1 ≤ d log Tr(Σ12 Σ̂T +1,a Σ12 ) , (B.23) d a=1 (︃ )︃ (︃ )︃ K ∑︂ T σ12 T σ12 ≤ d log 1 + = Kd log 1 + . d σ2 d σ2 a=1 (II) Remaining terms in Equation (B.22) Let ℓ ∈ [L]. Then we have that ⊤ −2 ⊤ −2ℓ 2ℓ log(1 + σ −2 Xt⊤ PAt ,ℓ Σ̄t,ℓ P⊤ At ,ℓ Xt ) = σmax σmax log(1 + σ Xt PAt ,ℓ Σ̄t,ℓ PAt ,ℓ Xt ) , 2ℓ −2ℓ ≤ σmax log(1 + σ −2 σmax Xt⊤ PAt ,ℓ Σ̄t,ℓ P⊤ At ,ℓ Xt ) , (i)
1
1
2ℓ −2ℓ 2 ⊤ 2 = σmax log det(Id + σ −2 σmax Σ̄t,ℓ P⊤ At ,ℓ Xt Xt PAt ,ℓ Σ̄t,ℓ ) , (︂ )︂ −1 2ℓ −2 −2ℓ ⊤ ⊤ = σmax log det(Σ̄−1 + σ σ P X X P ) − log det( Σ̄ ) , At ,ℓ max At ,ℓ t t t,ℓ t,ℓ
where we use the Weinstein-Aronszajn identity in (i). Now we know from Lemma 6 that −1 −2ℓ ⊤ the following inequality holds σ −2 σmax PAt ,ℓ Xt Xt⊤ PAt ,ℓ ⪯ Σ̄−1 t+1,ℓ − Σ̄t,ℓ . As a result, we get −1 ⊤ −2 −2ℓ ⊤ that Σ̄−1 t,ℓ + σ σmax PAt ,ℓ Xt Xt PAt ,ℓ ⪯ Σ̄t+1,ℓ . Thus, (︂ )︂ −1 −1 2ℓ X ) ≤ σ log det( Σ̄ ) − log det( Σ̄ ) , log(1 + σ −2 Xt⊤ PAt ,ℓ Σ̄t,ℓ P⊤ t At ,ℓ max t+1,ℓ t,ℓ Then we sum over all rounds t ∈ [T ], and get a telescoping T ∑︂ t=1
2ℓ log(1 + σ −2 Xt⊤ PAt ,ℓ Σ̄t,ℓ P⊤ At ,ℓ Xt ) ≤ σmax
T ∑︂
−1 log det(Σ̄−1 t+1,ℓ ) − log det(Σ̄t,ℓ ) ,
(︂t=1 )︂ −1 −1 log det(Σ̄T +1,ℓ ) − log det(Σ̄1,ℓ ) , (︂ )︂ (i) 2ℓ −1 = σmax log det(Σ̄−1 ) − log det(Σ ) , T +1,ℓ ℓ+1 )︂ (︂ 1 1 2ℓ 2 2 Σ̄−1 Σ = σmax log det(Σℓ+1 T +1,ℓ ℓ+1 ) , 2ℓ = σmax
where we use that Σ̄1,ℓ = Σℓ+1 in (i). Finally, we use the inequality of arithmetic and geometric means and get that T ∑︂
(︂ )︂ 1 1 −1 2ℓ 2 2 log(1 + σ −2 Xt⊤ PAt ,ℓ Σ̄t,ℓ P⊤ X ) ≤ σ log det(Σ Σ̄ Σ ) , At ,ℓ t max ℓ+1 T +1,ℓ ℓ+1
t=1
)︃ 1 1 1 −1 2 2 Tr(Σℓ+1 Σ̄T +1,ℓ Σℓ+1 ) , d (︃ 2 )︃ σℓ+1 2ℓ ≤ dσmax log 1 + 2 , σℓ 2ℓ ≤ dσmax log
155
(︃
(B.24)
The last inequality follows from the expression of Σ̄−1 T +1,ℓ in Equation (B.5) that leads to 1
1
1
1
2 2 2 2 Σ̄−1 Σℓ+1 T +1,ℓ Σℓ+1 = Id + Σℓ+1 ḠT +1,ℓ Σℓ+1 , 1 1 (︁ )︁ −1 −1 2 2 = Id + Σℓ+1 Wℓ⊤ Σ−1 − Σ (B.25) Σ̄ Σ W Σ T +1,ℓ−1 ℓ ℓ ℓ+1 , ℓ ℓ 1 1 (︁ )︁ −1 −1 2 2 since ḠT +1,ℓ = Wℓ⊤ Σ−1 Wℓ . This allows us to bound d1 Tr(Σℓ+1 Σ̄−1 ℓ −Σℓ Σ̄T +1,ℓ−1 Σℓ T +1,ℓ Σℓ+1 ) as 1 1 1 1 (︁ −1 )︁ 1 1 −1 −1 ⊤ 2 2 2 2 ) = W Σ − Σ Tr(Σℓ+1 Σ̄−1 Σ Tr(I + Σ Σ̄ Σ W Σ d T +1,ℓ−1 ℓ ℓ ℓ+1 ℓ ℓ ℓ+1 ) , T +1,ℓ ℓ+1 ℓ d d 1 1 (︁ )︁ 1 −1 −1 2 2 = (d + Tr(Σℓ+1 − Σ Wℓ⊤ Σ−1 Σ̄ Σ W Σ T +1,ℓ−1 ℓ ℓ ℓ ℓ ℓ+1 ) , d d 1 1 (︁ )︁ 1 ∑︂ −1 −1 2 2 ≤1+ Wℓ⊤ Σ−1 , λ1 (Σℓ+1 Wℓ Σℓ+1 ℓ − Σℓ Σ̄T +1,ℓ−1 Σℓ d i=1
d
≤1+
(︁ )︁ 1 ∑︂ −1 −1 λ1 (Σℓ+1 )λ1 (Wℓ⊤ Wℓ )λ1 Σ−1 , ℓ − Σℓ Σ̄T +1,ℓ−1 Σℓ d i=1
d (︁ )︁ 1 ∑︂ λ1 (Σℓ+1 )λ1 (Wℓ⊤ Wℓ )λ1 Σ−1 , ≤1+ ℓ d i=1 d
2 2 σℓ+1 1 ∑︂ σℓ+1 =1+ 2 , ≤1+ d i=1 σℓ2 σℓ
(B.26)
2 where we use the assumption that λ1 (Wℓ⊤ Wℓ ) = 1 (A4) and that λ1 (Σℓ+1 ) = σℓ+1 and −1 2 2 λ1 (Σℓ ) = 1/σℓ . This is because Σℓ = σℓ Id for any ℓ ∈ [L + 1]. Finally, plugging Equations (B.23) and (B.24) in Equation (B.22) concludes the proof.
B.5.5
Proof of proposition 7
We use exactly the same proof in Section B.5.4, with one change to account for the sparsity assumption (A5). The change corresponds to Equation (B.24). First, recall that Equation (B.24) writes T ∑︂
)︂ (︂ 1 1 −1 2ℓ 2 2 Σ̄ Σ ) , log(1 + σ −2 Xt⊤ PAt ,ℓ Σ̄t,ℓ P⊤ X ) ≤ σ log det(Σ t At ,ℓ max ℓ+1 T +1,ℓ ℓ+1
t=1
where 1 1 1 1 )︁ (︁ −1 −1 −1 ⊤ 2 2 2 2 Wℓ Σℓ+1 Σℓ+1 Σ̄−1 , T +1,ℓ Σℓ+1 = Id + Σℓ+1 Wℓ Σℓ − Σℓ Σ̄t,ℓ−1 Σℓ )︁ (︁ −1 −1 2 Wℓ , = Id + σℓ+1 Wℓ⊤ Σ−1 ℓ − Σℓ Σ̄t,ℓ−1 Σℓ
(B.27)
2 where the second equality follows from the assumption that Σℓ+1 = σℓ+1 Id . But notice that in our assumption, (A5), we assume that Wℓ = (W̄ℓ , 0d,d−dℓ ), where W̄ℓ ∈ Rd×dℓ for any ℓ ∈ [L].(︃Therefore, we have that for any d × d matrix B ∈ Rd×d , the following holds, )︃ W̄ℓ⊤ BW̄ℓ 0dℓ ,d−dℓ Wℓ⊤ BWℓ = . In particular, we have that 0d−dℓ ,dℓ 0d−dℓ ,d−dℓ )︁ (︃ ⊤ (︁ −1 )︃ (︁ −1 )︁ W̄ℓ Σℓ − Σ−1 Σ̄t,ℓ−1 Σ−1 W̄ℓ 0dℓ ,d−dℓ −1 −1 ⊤ ℓ ℓ Wℓ Σℓ − Σℓ Σ̄t,ℓ−1 Σℓ Wℓ = . (B.28) 0d−dℓ ,dℓ 0d−dℓ ,d−dℓ
156
Therefore, plugging this in Equation (B.27) yields that (︁ −1 )︁ (︃ )︃ −1 −1 2 ⊤ 1 1 + σ I W̄ Σ − Σ Σ̄ Σ W̄ 0 d t,ℓ−1 ℓ d ,d−d −1 ℓ+1 ℓ 2 2 ℓ ℓ ℓ ℓ ℓ ℓ Σℓ+1 = . Σ̄T +1,ℓ Σℓ+1 0d−dℓ ,dℓ Id−dℓ
(B.29)
1 1 (︁ −1 )︁ −1 −1 2 ⊤ 2 2 As a result, det(Σℓ+1 ) = det(I + σ Σ̄−1 Σ W̄ Σ − Σ Σ̄ Σ W̄ℓ ). This d t,ℓ−1 ℓ+1 ℓ ℓ T +1,ℓ ℓ+1 ℓ ℓ ℓ allows us to move the problem from a d-dimensional one to a dℓ -dimensional one. Then we use the inequality of arithmetic and geometric means and get that
T ∑︂
)︂ (︂ 1 1 −1 2ℓ 2 2 ) , Σ̄ Σ log(1 + σ −2 Xt⊤ PAt ,ℓ Σ̄t,ℓ P⊤ X ) ≤ σ log det(Σ t At ,ℓ max ℓ+1 T +1,ℓ ℓ+1
t=1
(︁ )︁ −1 −1 2ℓ 2 = σmax log det(Idℓ + σℓ+1 W̄ℓ⊤ Σ−1 W̄ℓ ) , ℓ − Σℓ Σ̄t,ℓ−1 Σℓ (︃ )︃ (︁ −1 )︁ 1 −1 −1 2 ⊤ 2ℓ Tr(Idℓ + σℓ+1 W̄ℓ Σℓ − Σℓ Σ̄t,ℓ−1 Σℓ W̄ℓ ) , ≤ dℓ σmax log dℓ (︃ 2 )︃ σℓ+1 2ℓ . (B.30) ≤ dℓ σmax log 1 + 2 σℓ To get the last inequality, we use derivations similar to the ones we used in Equation (B.26). Finally, the desired result in obtained by replacing Equation (B.24) by Equation (B.30) in the previous proof in Section B.5.4.
B.6
Additional Experiments
B.6.1
Swiss roll data
Figure B.1 shows samples from the Swiss roll data and samples from generated by the pre-trained diffusion model for different pre-training sample sizes.
(a) Diffusion pre-trained on 50 (b) Diffusion pre-trained on (c) Diffusion pre-trained on 104 samples from the Swiss roll 103 samples from the Swiss roll samples from the Swiss roll dataset. dataset. dataset.
Figure B.1: True distribution of action parameters (blue) vs. distribution of pre-trained diffusion model (red).
B.6.2
Diffusion models pre-training
We used JAX for diffusion model pre-training, summarized as follows: 157
• Parameterization: Functions fℓ are parameterized with a fully connected 2-layer neural network (NN) with ReLU activation. The step ℓ is provided as input to capture the current sampling stage. Covariances are fixed (not learned) as Σℓ = σℓ2 Id with σℓ increasing with ℓ. • Loss: Offline data samples are progressively noised over steps ℓ ∈ [L], creating increasingly noisy versions of the data following a predefined noise schedule (Ho et al., 2020). The NN is trained to reverse this noise (i.e., denoise) by predicting the noise added at each step. The loss function measures the L2 norm difference between the predicted and actual noise at each step, as explained in Ho et al. (2020). • Optimization: Adam optimizer with a 10−3 learning rate was used. The NN was trained for 20,000 epochs with a batch size of min(2048, pre-training sample size). We used CPUs for pre-training, which was efficient enough to conduct multiple ablation studies. • After pre-training: The pre-trained diffusion model is used as a prior for sDM and compared to LinTS as the reference baseline. In our ablation study, we plot the cumulative regret of LinTS in the last round divided by that of sDM. A ratio greater than 1 indicates that sDM outperforms LinTS, with higher values representing a larger performance gap.
B.6.3
Quality of our posterior approximation
To assess the quality of our posterior approximation, we consider the scenario where the true distribution of action parameters is N (0d , Id ) with d = 2 and rewards are linear. We pre-train a diffusion model using samples drawn from N (0d , Id ). We then consider two priors: the true prior N (0d , Id ) and the pre-trained diffusion model prior. This yields two posteriors:
• P1 : Uses N (0d , Id ) as the prior. P1 is an exact posterior since the prior is Gaussian and rewards are linear-Gaussian. • P2 : Uses the pre-trained diffusion model as the prior. P2 is our approximate posterior.
The learned diffusion model prior matches the true Gaussian prior (as seen in Figure B.2a). Thus, if our approximation is accurate, their posteriors P1 and P2 should also be similar. This is observed in Figure B.2b where the approximate posterior P2 nearly matches the exact posterior P1 . 158
(b) Exact posterior P1 vs. approximate posterior P2 after T = 100 rounds of interactions.
(a) Gaussian distribution vs. diffusion model pre-trained on 103 samples drawn from it.
Figure B.2: Assessing the quality of our posterior approximation.
B.6.4
CIFAR Ablation
CIFAR. In Figure 4.3a in Section 4.4.2, we showed that with only 10 pre-training samples, sDM outperforms LinTS on the Swiss-roll benchmark. We now extend this analysis to the vision dataset CIFAR (Krizhevsky et al., 2009) (similar results were obtained on MNIST (Zhu, 2018)). Our setting is similar to that in Hong et al. (2022a) and we use sDM’s variant that uses a single shared parameter θ ∈ Rd (Remark 5 and Section 4.2.2) because it is more suited for this setting. These additional ablations on CIFAR confirm that sDM consistently benefits from offline pre-training, even when the true prior is not a diffusion model. Specifically, we vary the percentage of offline data used to train the prior and compare against both HierTS and LinTS. Table B.1: Regret improvement (%) of sDM on CIFAR. Offline Data (%) 1% 5% 25% 50%
B.6.5
vs. HierTS
vs. LinTS
69.11% 79.56% 80.65% 81.67%
87.74% 92.18% 92.48% 92.88%
Bound comparison
Here, we compare our bound in Theorem 7 to bounds of LinTS from the literature. 159
Linear Bandit: d=5, K=10000, L=5
600000
Bayes Bound for dTS (Ours) Frequentist Bound for LinTS
500000
Regret
Regret
400000 300000 200000 100000 0
0
200
400 600 Horizon n
800
1000
(a) Our bound vs. the frequentist bound of LinTS in Abeille and Lazaric (2017).
90000 80000 70000 60000 50000 40000 30000 20000 10000 0
Linear Bandit: d=5, K=10000, L=5 Bayes Bound for dTS (Ours) Bayes Bound for LinTS
0
200
400 600 Horizon n
800
1000
(b) Our bound vs. the standard Bayesian bound of LinTS.
Figure B.3: Comparing our Bayesian regret bound of dTS to the frequentist and Bayesian bounds of LinTS.
160
Chapter C
Supplementary Materials for Chapter 6
Contents B.1 Posterior for Linear Diffusion Models . . . . . . . . . . . . . . . . . . 143 B.1.1
Linear Diffusion Models . . . . . . . . . . . . . . . . . . . . .
144
B.1.2
Posterior Expressions for Linear Diffusion Models . . . . . . .
144
B.2 Posterior for Non-Linear Diffusion Models . . . . . . . . . . . . . . . . 145 B.3 Connection to Two-Level Hierarchies . . . . . . . . . . . . . . . . . . . 146 B.4 Formal Theory . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 147 B.5 Regret proof . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 148 B.5.1
Proof Sketch . . . . . . . . . . . . . . . . . . . . . . . . . . . .
149
B.5.2
Proof of lemma 5 . . . . . . . . . . . . . . . . . . . . . . . . .
150
B.5.3
Proof of lemma 6 . . . . . . . . . . . . . . . . . . . . . . . . .
151
B.5.4
Proof of theorem 7 . . . . . . . . . . . . . . . . . . . . . . . .
153
B.5.5
Proof of proposition 7
156
. . . . . . . . . . . . . . . . . . . . . .
B.6 Additional Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . 157
C.1
B.6.1
Swiss roll data . . . . . . . . . . . . . . . . . . . . . . . . . . .
157
B.6.2
Diffusion models pre-training . . . . . . . . . . . . . . . . . .
157
B.6.3
Quality of our posterior approximation . . . . . . . . . . . . .
158
B.6.4
CIFAR Ablation . . . . . . . . . . . . . . . . . . . . . . . . . .
159
B.6.5
Bound comparison . . . . . . . . . . . . . . . . . . . . . . . .
159
Posterior Derivations Under Standard Priors
Here we derive the posterior under the standard prior in Equation (6.2). These are standard derivations and we present them here for the sake of completeness. But first, we state the following standard assumption that allows posterior derivations. 161
Assumption 5 (Independence). (X, A) is independent of θ, and the θa , for a ∈ A are independent. Derivation of p(θa | Dn ) for the standard prior in Equation (6.2). We start by recalling the standard prior in Equation (6.2) θa ∼ N (µa , Σa ), ⊤
∀a ∈ A,
(C.1)
2
R | θ, X, A ∼ N (ϕ(X) θA , σ ), where N (µa , Σa ) is the prior on the action parameter θa . Let θ = (θa )a∈A ∈ RdK , ΣA = diag(Σa )a∈A ∈ RdK×dK and µA = (µa )a∈A ∈ RdK . Also, let ua ∈ {0, 1}K be the binary vector representing the action a. That is, ua,a = 1 and ua,a′ = 0 for all a′ ̸= a. Then we can rewrite the model in Equation (C.1) as θ ∼ N (µA , ΣA ),
(C.2)
R | θ, X, A ∼ N ((uA ⊗ ϕ(X))⊤ θ, σ 2 ). Then the joint action posterior p(θ | Dn ) decomposes as (i)
p(θ | Dn ) = p(θ | (Xi , Ai , Ri )i∈[n] ) ∝ p((Ri )i∈[n] | θ, (Xi , Ai )i∈[n] )p(θ | (Xi , Ai )i∈[n] ), (ii) (iii) ∏︂ = p((Ri )i∈[n] | θ, (Xi , Ai )i∈[n] )p(θ) = p(Ri | θ, Xi , Ai )p(θ), i∈[n] (iv) ∏︂ = N (Ri ; (uAi ⊗ ϕ(Xi ))⊤ θ, σ 2 )N (θ; µA , ΣA ), i∈[n]
(︃ (︂ )︂−1 )︃ ∝ N θ; µ̂A , Λ̂A .
(v)
In (i), we apply Bayes rule. In (ii), we use that θ is independent of (X, A), and (iii) follows from the assumption that Ri | θ, Xi , Ai are independent. Finally, ∑︁n in (iv),⊤ we replace the distribution by their Gaussian form, and in (v), we set Λ̂A = )︂ v i=1 (uAi uAi ⊗ (︁ ∑︁n −1 ⊤ ϕ(Xi )ϕ(Xi ) ) + ΛA , and µ̂A = Λ̂A v i=1 (uAi ⊗ ϕ(Xi ))Ri + ΛA µA where v = σ −2 −1 −1 and ΛA = Σ−1 A = diag(Σa )a∈A . Now notice that Λ̂A = diag(Σa + Ga )a∈A . Thus, p(θ | Dn ) = N (θ; µ̂A , Λ̂−1 A ) where µ̂A = (µ̂a )a∈A and Λ̂A = diag(Λ̂a )a∈A , with
Λ̂a µ̂a = Σ−1 a µa + B a .
Λ̂a = Σ−1 a + Ga ,
Since the covariance matrix of p(θ | Dn ) is diagonal by block, we know that the marginals θa | Dn also have a Gaussian density p(θa | Dn ) = N (θa ; µ̂a , Σ̂a ) where Σ̂a = Λ̂−1 a .
C.2
Posterior Derivations Under Structured Priors
Here we derive the posteriors under the structured prior in Equation (6.9). Precisely, we derive the latent posterior density of ψ | Dn , the conditional posterior density of θ | Dn , ψ. Then, we derive the marginal posterior θ | Dn . Posterior derivations rely on the following assumption. 162
Assumption 6 (Structured Independence). (i) (X, A) is independent of ψ and given ψ, (X, A) is independent of θ. (ii) Given ψ, the θa , for all a ∈ A are independent.
C.2.1
Latent Posterior
Derivation of p(ψ | Dn ). First, recall that our model in Equation (6.9) reads ψ ∼ N (µ, Σ), (︂ )︂ θa | ψ ∼ N Wa ψ, Σa ,
∀a ∈ A, (C.3)
R | ψ, θ, X, A ∼ N (ϕ(X)⊤ θA , σ 2 ). Then we first rewrite it as ψ ∼ N (µ, Σ), (︂ )︂ θ | ψ ∼ N WA ψ, ΣA , R | ψ, θ, X, A ∼ N ((uA ⊗ ϕ(X))⊤ θ, σ 2 ).
(C.4)
Then the latent posterior is p(ψ | (Xi , Ai , Ri )i∈[n] ) ∝ p((Ri )i∈[n] | ψ, (Xi , Ai )i∈[n] )p(ψ | (Xi , Ai )i∈[n] ), (i)
= p((Ri )i∈[n] | ψ, (Xi , Ai )i∈[n] )q(ψ), ∫︂ = p((Ri )i∈[n] , θ | ψ, (Xi , Ai )i∈[n] ) dθq(ψ), ∫︂θ = p((Ri )i∈[n] | ψ, θ, (Xi , Ai )i∈[n] )p(θ | ψ, (Xi , Ai )i∈[n] ) dθq(ψ), ∫︂θ (ii) = p((Ri )i∈[n] | ψ, θ, (Xi , Ai )i∈[n] )p(θ | ψ) dθq(ψ), θ
In (i), we use that (X, A) is independent of ψ, which follows from Assumption 6. Similarly, in (ii), we use that θ is conditionally independent of (X, A) given ψ. Now we know that given θ, Ri | Xi , Ai are i.i.d. and hence p((Ri )i∈[n] | ψ, θ, (Xi , Ai )i∈[n] ) = ∏︁ ). Moreover, θa for a ∈ A are conditionally independent given ψ. Thus a∈A La (θa∏︁ p(θ | ψ) = a∈A pa (θa ; fa (ψ)), where we also used that θa | ψ ∼ pa (·; fa (ψ)). This leads to ∫︂ ∏︂ p(ψ | (Xi , Ai , Ri )i∈[n] ) ∝ La (θa )pa (θa ; fa (ψ)) dθq(ψ), θ a∈A (i) ∏︂
∫︂ La (θa )N (θa ; Wa ψ, Σa ) dθa N (ψ; µ, Σ),
=
a∈A (ii) ∏︂
=
a∈A
θa
∫︂ (︂ ∏︂ θa
)︂ N (Ri ; ϕ(Xi ) θa , σ ) N (θa ; Wa ψ, Σa ) dθa N (ψ; µ, Σ). ⊤
i∈Ia
163
2
In (i), we notice that θ = (θa )a∈A and apply Fubini’s Theorem. In (ii), we let Ia = {i ∈ [n]; ∫︁Ai (︁=∏︁a} as the rounds where )︁action a appears in the sample set Dn . Now let ⊤ 2 ha (ψ) = θa i∈Ia N (Ri ; ϕ(Xi ) θa , σ ) N (θa ; Wa ψ, Σa ) dθa . Then we have that ∏︂ p(ψ | Dn ) ∝ ha (ψ)N (ψ; µ, Σ). (C.5) a∈A
We start by computing ha . To reduce clutter, let v = σ −2 and Λa = Σ−1 a . Then we compute ha as )︄ ∫︂ (︄ ∏︂ N (Ri ; ϕ(Xi )⊤ θa , σ 2 ) N (θa ; Wa ψ, Σa ) dθa , ha (ψ) = θa
i∈Ia
[︄
]︄ 1 ∑︂ 1 ∝ exp − v (Ri − ϕ(Xi )⊤ θa )2 − (θa − Wa ψ)⊤ Λa (θa − Wa ψ) dθa , 2 2 θa i∈Ia ∫︂ [︂ 1 (︂ ∑︂ = exp − v (Ri2 − 2Ri θa⊤ ϕ(Xi ) + (θa⊤ ϕ(Xi ))2 ) + θa⊤ Λa θa − 2θa⊤ Λa Wa ψ 2 θa i∈Ia )︂]︂ ⊤ + (Wa ψ) Λa (Wa ψ) dθa , )︄ )︄ (︄ (︄ ∫︂ [︂ 1 (︂ ∑︂ ∑︂ Ri ϕ(Xi ) + Λa Wa ψ ϕ(Xi )ϕ(Xi )⊤ + Λa θa − 2θa⊤ v θa⊤ v ∝ exp − 2 θa i∈Ia i∈Ia )︂]︂ + (Wa ψ)⊤ Λa (Wa ψ) dθa . ∑︁ ∑︁ Now recall that Ga = v i∈Ia ϕ(Xi )ϕ(Xi )⊤ and Ba = v i∈Ia Ri ϕ(Xi ) and let Va = (Ga + Λa )−1 , Ua = Va−1 , and βa = Va (Ba + Λa Wa ψ). Then have that Ua Va = Va Ua = Id , and thus [︃ ]︃ ∫︂ )︁ 1 (︁ ⊤ ⊤ ⊤ ha (ψ) ∝ exp − θa Ua θa − 2θa Ua Va (Ba + Λa Wa ψ) + (Wa ψ) Λa (Wa ψ) dθa , 2 θa ]︃ [︃ ∫︂ )︁ 1 (︁ ⊤ ⊤ ⊤ = exp − θa Ua θa − 2θa Ua βa + (Wa ψ) Λa (Wa ψ) dθa , 2 θa [︃ ]︃ ∫︂ )︁ 1 (︁ ⊤ ⊤ ⊤ exp − (θa − βa ) Ua (θa − βa ) − βa Ua βa + (Wa ψ) Λa (Wa ψ) dθa , = 2 θa [︃ ]︃ )︁ 1 (︁ ⊤ ⊤ ∝ exp − −βa Ua βa + (Wa ψ) Λa (Wa ψ) , 2 [︃ )︂]︃ 1 (︂ ⊤ ⊤ = exp − − (Ba + Λa Wa ψ) Va (Ba + Λa Wa ψ) + (Wa ψ) Λa (Wa ψ) , 2 ]︃ [︃ )︁)︁ (︁ ⊤ 1 (︁ ⊤ ⊤ ⊤ , ∝ exp − ψ Wa (Λa − Λa Va Λa ) Wa ψ − 2ψ Wa Λa Va Ba 2 [︃ ]︃ )︁ 1 (︁ ⊤ ⊤ ∝ exp − ψ Λ̄a ψ − 2ψ Λ̄a µ̄a , 2 ∫︂
where (︁ )︁ −1 −1 −1 −1 Wa , Λ̄a = Wa⊤ (Λa − Λa Va Λa ) Wa = Wa⊤ Σ−1 a − Σa (Ga + Σa ) Σa −1 −1 Λ̄a µ̄a = Wa⊤ Λa Va Ba = Wa⊤ Σ−1 a (Ga + Σa ) Ba .
164
(C.6)
∏︁ However, we know from Equation that p(ψ | D n) ∝ a∈A ha (ψ)N (ψ; µ, Σ). But [︁ 1 (︁ (C.5) )︁]︁ ha (ψ) is proportional to exp − 2 ψ ⊤ Λ̄a ψ − 2ψ ⊤ Λ̄a µ̄a for any a. Thus p(ψ | Dn ) can be seen as the product of K + 1 Gaussian kernels. Thus, p(ψ | Dn ) is a multivariate Gaussian distribution N (µ̄, Σ̄), with ∑︂ ∑︂ )︁ (︁ −1 −1 −1 −1 Wa , (C.7) Σ̄−1 = Σ−1 + Λ̄a = Σ−1 + Wa⊤ Σ−1 a − Σa (Ga + Σa ) Σa a∈A −1
−1
Σ̄ µ̄ = Σ µ +
∑︂
a∈A −1
Λ̄a µ̄a = Σ µ +
a∈A
C.2.2
∑︂
(C.8)
−1 −1 Wa⊤ Σ−1 a (Ga + Σa ) Ba .
a∈A
Conditional Posterior
Derivation of p(θa | ψ, Dn ). Let v = σ −2 , Λa = Σ−1 a . We consider the model rewritten in Equation (C.4), then the joint conditional action posterior p(θ | ψ, Dn ) decomposes as (i)
p(θ | ψ, Dn ) = p(θ | ψ, (Xi , Ai , Ri )i∈[n] ) ∝ p((Ri )i∈[n] | θ, ψ, (Xi , Ai )i∈[n] )p(θ | ψ, (Xi , Ai )i∈[n] ), (ii) (iii) ∏︂ = p((Ri )i∈[n] | θ, (Xi , Ai )i∈[n] )p(θ | ψ) = p(Ri | θ, Xi , Ai )p(θ | ψ), i∈[n] (iv) ∏︂ = N (Ri ; (uAi ⊗ ϕ(Xi ))⊤ θ, σ 2 )N (θ; WA ψ, ΣA ), i∈[n] n
1 (︂ ∑︂ 2 = exp − v (Ri − 2Ri (uAi ⊗ ϕ(Xi ))⊤ θ + ((uAi ⊗ ϕ(Xi ))⊤ θ)2 ) + θ⊤ ΛA θ − 2θ⊤ ΛA WA ψ 2 i=1 )︂]︂ ⊤ + ψ ⊤ WA Λ A WA ψ , [︂
n 1 (︂ ⊤ ∑︂ ∝ exp − θ (v (uAi ⊗ ϕ(Xi ))(uAi ⊗ ϕ(Xi ))⊤ + ΛA )θ 2 i=1 (︄ n )︄ )︂]︂ ∑︂ ⊤ − 2θ v (uAi ⊗ ϕ(Xi ))Ri + ΛA WA ψ ,
[︂
i=1
[︂
= exp −
1 (︂ 2
θ⊤ (v
n ∑︂
⊤ (uAi u⊤ Ai ⊗ ϕ(Xi )ϕ(Xi ) ) + ΛA )θ−
i=1
(︄ 2θ⊤ v
n ∑︂
)︄ (uAi ⊗ ϕ(Xi ))Ri + ΛA WA ψ
)︂]︂
,
i=1
(︃ (︂ )︂−1 )︃ ∝ N θ; µ̃A , Λ̃A ,
(v)
where we use Bayes rule in (i), (ii) uses two assumptions. First, Given θ, X, A, R is independent of ψ. Second, given ψ, θ is independent of (X, A). Moreover, (iii) follows from the assumption that Ri | θ, Xi , Ai are independent. Finally, replace the distribution ∑︁nin iv, we ⊤ by their Gaussian form, and in (v), we set Λ̃A = )︂v i=1 (uAi uAi ⊗ ϕ(Xi )ϕ(Xi )⊤ ) + ΛA , (︁ ∑︁n −1 −1 and µ̃A = Λ̃−1 A v i=1 (uAi ⊗ ϕ(Xi ))Ri + ΛA WA ψ , where ΛA = ΣA = diag(Σa )a∈A . 165
−1 Now notice that Λ̃A = diag(Σ−1 a + Ga )a∈A . Thus, p(θ | ψ, Dn ) = N (θ; µ̃A , Λ̃A ) where µ̃A = (µ̃a )a∈A and Λ̃A = diag(Λ̃a )a∈A , with
Λ̃a = Σ−1 a + Ga , Λ̃a µ̃a = Σ−1 a Wa ψ + B a . The covariance matrix of p(θ | ψ, Dn ) is diagonal by block. Thus θa | ψ, Dn for a ∈ A are independent and have a Gaussian density p(θa | ψ, Dn ) = N (θa ; µ̃a , Σ̃a ) where Σ̃a = Λ̃−1 a .
C.2.3
Action Posterior
Derivation of p(θa | Dn ). We know that θa | Dn , ψ ∼ N (µ̃a , Σ̃a ) and ψ | Dn ∼ N (µ̄, Σ̄). Thus the posterior density of θa | Dn is also Gaussian since Gaussianity is preserved after marginalization (Koller and Friedman, 2009). We let θa | Dn ∼ N (µ̂a , Σ̂a ). Then, we can compute µ̂a and Σ̂a using the total expectation and total covariance decompositions. Let Λa = Σ−1 a . Then we have that Σ̃a = (Ga + Λa )−1 E [θa | ψ, Dn ] = Σ̃a (Ba + Λa Wa ψ) First, given Dn , Σ̃a = (Ga + Λa )−1 and Ba are constant (do not depend on ψ). Thus [︂ ]︂ µ̂a = E [θa | Dn ] = E [E [θa | ψ, Dn ] | Dn ] = Eψ∼N (µ̄,Σ̄) Σ̃a (Ba + Λa Wa ψ) (︁ )︁ = Σ̃a Ba + Λa Wa Eψ∼N (µ̄,Σ̄) [ψ] , = Σ̃a (Ba + Λa Wa µ̄) . This concludes the computation of µ̂a . Similarly, given Dn , Σ̃a = (Ga + Λa )−1 and Ba are constant (do not depend on ψ), yields two things. First, [︂ ⃓ ]︂ ⃓ E [cov [θa | ψ, Dn ] | Dn ] = E Σ̃a ⃓ Dn = Σ̃a . Second, ⃓ ]︂ ⃓ cov [E [θa | ψ, Dn ] | Dn ] = cov Σ̃a Λa Wa ψ ⃓ Dn [︂
= Σ̃a Λa Wa cov [ψ | Dn ] Wa⊤ Λa Σ̃a = Σ̃a Λa Wa Σ̄Wa⊤ Λa Σ̃a . Finally, the total covariance decomposition (Weiss, 2005) yields that Σ̂a = cov [θa | Dn ] = E [cov [θa | ψ, Dn ] | Dn ] + cov [E [θa | ψ, Dn ] | Dn ] = Σ̃a + Σ̃a Λa Wa Σ̄Wa⊤ Λa Σ̃a . This concludes the proof.
166
C.3
Proofs
C.3.1
Main Result
In this section, we prove Theorem 2. Recall that we make the following well-specified prior assumption. Assumption 7 (Well-specified priors). Action parameters θ∗,a and rewards are drawn from Equation (6.9). Assumption 8 (Diagonal covariances for simplicity). We assume Σa = σ02 Id , Σ = τ 2 Id′ , ∥ϕ(x)∥2 ≤ 1, and the matrices Wa are normalized such that λ1 (Wa Wa⊤ ) = λd (Wa Wa⊤ ) = 1. Proof. First, given x ∈ X , by definition of the optimal policy, we know that it is deterministic. That is, there exists ax,θ∗ ∈ [K] such that π∗ (ax,θ∗ | x) = 1. To simplify the notation and since π∗ is deterministic, we let π∗ (x) = ax,θ∗ . Also, we know that the greedy policy is deterministic in âx = argmaxb∈A r̂(x, b). That is π̂g (âx | x) = 1. Similarly, we let π̂g (x) = âx . Moreover, we let Φ(x, a) = ea ⊗ ϕ(x) ∈ RdK where ea ∈ RK is the indicator vector of action a, such that ea,b = 0 for any b ∈ A/{a} and ea,a = 1. Also, recall that µ̂ = (µ̂a )a∈A is the concatenation of the posterior means. Bso(π̂g ) = E [V (π∗ ; θ∗ ) − V (π̂g ; θ∗ )] , = E [r(X, π∗ (X); θ∗ ) − r(X, π̂g (X); θ∗ )] , = E [r(X, π∗ (X); θ∗ ) − r(X, π̂g (X); µ̂) + r(X, π̂g (X); µ̂) − r(X, π̂g (X); θ∗ )] , ≤ E [r(X, π∗ (X); θ∗ ) − r(X, π∗ (X); µ̂) + r(X, π̂g (X); µ̂) − r(X, π̂g (X); θ∗ )] , ≤ E [r(X, π∗ (X); θ∗ ) − r(X, π∗ (X); µ̂)] + E [r(X, π̂g (X); µ̂) − r(X, π̂g (X); θ∗ )] . Now we start by proving that E [r(X, π̂g (X); µ̂) − r(X, π̂g (X); θ∗ )] = 0. This is achieved as follows E [r(X, π̂g (X); µ̂) − r(X, π̂g (X); θ∗ )] = E [E [r(X, π̂g (X); µ̂) − r(X, π̂g (X); θ∗ ) | X, Dn ]] , [︁ [︁ ]︁]︁ = E E ϕ(X)⊤ µ̂π̂g (X) − ϕ(X)⊤ θ∗,π̂g (X) | X, Dn , [︁ [︁ ]︁]︁ (i) = E E Φ(X, π̂g (X))⊤ µ̂ − Φ(X, π̂g (X))⊤ θ∗ | X, Dn , [︁ ]︁ (ii) = E Φ(X, π̂g (X))⊤ E [µ̂ − θ∗ | X, Dn ] , [︁ ]︁ (iii) = E Φ(X, π̂g (X))⊤ (µ̂ − E [θ∗ | X, Dn ]) , (iv)
= 0.
In (i), we used that by definition of Φ(x, a) = ea ⊗ ϕ(X) ∈ RdK , we have ϕ(x)⊤ µ̂a = Φ(x, a)⊤ µ̂ for any (x, a), and the same holds for θ∗ . In (ii), we used that Φ(X, π̂g (X)) is deterministic given X and Dn . In (iii), we used that µ̂ is deterministic given X and Dn . Finally, in (iv), we used that E [θ∗ | X, Dn ] = E [θ∗ | Dn ] = µ̂, which follows from the assumption that θ∗ does not depend on X and the assumption that θ∗ is drawn from the prior, and hence when conditioned on Dn , it is drawn from the posterior whose mean is µ̂. Therefore, E [r(X, π̂g (X); µ̂) − r(X, π̂g (X); θ∗ )] = 0 which leads to Bso(π̂g ) ≤ E [r(X, π∗ (X); θ∗ ) − r(X, π∗ (X); µ̂)] . 167
Now let δ ∈ (0, 1) and define the joint parameter event {︂ }︂ Eα = ∀a ∈ A : ∥θ∗,a − µ̂a ∥Σ̂−1 ≤ α . a Recall Za = r(X, a; θ∗ ) − r(X, a; µ̂) = ϕ(X)⊤ (θ∗,a − µ̂a ). By Cauchy-Schwarz, on Eα we have for all a, |Za | = |ϕ(X)⊤ (θ∗,a − µ̂a )| ≤ ∥ϕ(X)∥Σ̂a ∥θ∗,a − µ̂a ∥Σ̂−1 ≤ α∥ϕ(X)∥Σ̂a , a hence in particular |Zπ∗ (X) | ≤ α∥ϕ(X)∥Σ̂π (X) on Eα . Therefore, ∗
[︁ ]︁ Bso(π̂g ) ≤ E |Zπ∗ (X) | [︁ ]︁ [︁ {︁ }︁]︁ = E |Zπ∗ (X) |1{Eα } + E |Zπ∗ (X) |1 Ēα [︂ ]︂ [︁ {︁ }︁]︁ ≤ αE ∥ϕ(X)∥Σ̂π (X) + E |Zπ∗ (X) |1 Ēα . ∗
We now bound P(Ēα | Dn ). Under the well-specified Bayes assumption, conditional on Dn each marginal posterior satisfies θ∗,a | Dn ∼ N (µ̂a , Σ̂a ). This means that θ∗,a − µ̂a | −1
= Dn ∼ N (0, Σ̂a ). Thus, Σ̂a 2 (θ∗,a − µ̂a ) ∼ N (0, Id ). But notice that ∥θ∗,a − µ̂a ∥Σ̂−1 a −1
∥Σ̂a 2 (θ∗,a − µ̂a )∥. Thus we apply Laurent and Massart (2000, Lemma 1) and get that ⃓ )︂ (︂ δ ⃓ P ∥θ∗,a − µ̂a ∥Σ̂−1 ≥ 1 − , ≤ α D ⃓ n a K where α =
√︃
√︂ d + 2 d log Kδ + 2 log Kδ . This means that for every a, P(∥θ∗,a − µ̂a ∥Σ̂−1 > a
α | Dn ) ≤ δ/K, and by a union bound, (C.9)
P(Ēα | Dn ) ≤ δ.
and therefore P(Ēα ) = E[P(Ēα | Dn )] ≤ δ. Finally, control the bad-event contribution by Cauchy-Schwarz: √︃ [︂ √︃ [︂ ]︂√︂ ]︂√ [︁ {︁ }︁]︁ 2 P(Ēα ) ≤ E Zπ2∗ (X) δ. (C.10) E |Zπ∗ (X) |1 Ēα ≤ E Zπ∗ (X) Putting the pieces together gives [︂
]︂
Bso(π̂g ) ≤ αE ∥ϕ(X)∥Σ̂π (X) + ∗
with α2 = d + 2
√
√︃ [︂ ]︂ δ E Zπ2∗ (X) ,
(C.11)
√︁
d log(K/δ) + 2 log(K/δ). [︂ ]︂ We now upper bound E Zπ2∗ (X) in (C.11). We have that Zπ2∗ (X) ≤ max Za2 , a∈A
hence 168
E
[︁
Zπ2∗ (X)
]︁
[︃ ]︃ 2 ≤ E max Za . a∈A
(C.12)
Fix (X, Dn ). Under the well-specified Bayes assumption, the conditional joint posterior (θ∗,a )a∈A | Dn is Gaussian, and each marginal satisfies θ∗,a | Dn ∼ N (µ̂a , Σ̂a ) (with Σ̂a as derived in Section C.2.3). Therefore for each fixed a, (︁ )︁ Za | X, Dn ∼ N 0, s2a , s2a = ∥ϕ(X)∥2Σ̂a . Let s2max = maxa∈A s2a . Then for any t ≥ 0, by a union bound and the Gaussian tail bound, ⃓ (︃ )︃ ∑︂ (︃ )︃ ⃓ t2 ⃓ P max |Za | ≥ t⃓X, Dn ≤ P (|Za | ≥ t|X, Dn ) ≤ 2K exp − 2 . a∈A 2smax a∈A ∫︁ ∞ Using the identity E [W 2 ] = 0 P(W 2 ≥ u)du for W ≥ 0 and setting W = maxa |Za |, we obtain ⃓ ⃓ [︃ ]︃ ∫︂ ∞ (︃ )︃ ⃓ ⃓ √ 2⃓ ⃓ P max |Za | ≥ u⃓X, Dn du E max Za ⃓X, Dn = a∈A a∈A 0 {︃ (︃ )︃}︃ ∫︂ ∞ u min 1, 2K exp − 2 ≤ du 2smax 0 (︃ )︃ ∫︂ ∞ ∫︂ 2s2max log(2K) u 2K exp − 2 1du + du = 2smax 2s2max log(2K) 0 (︁ )︁ = 2s2max log(2K) + 2s2max = 2 log(2K) + 2 s2max . Taking expectation over (X, Dn ) yields [︃ ]︃ [︃ ]︃ (︁ )︁ 2 2 E max Za ≤ 2 log(2K) + 2 E max ∥ϕ(X)∥Σ̂a . a∈A
a∈A
(C.13)
Combining (C.12) and (C.13) gives E
[︁
Zπ2∗ (X)
]︁
[︃ ]︃ (︁ )︁ 2 ≤ 2 log(2K) + 2 E max ∥ϕ(X)∥Σ̂a . a∈A
(C.14)
Using the assumption that ∥ϕ(X)∥2 ≤ 1 yields that ∥ϕ(X)∥2Σ̂ ≤ λ1 (Σ̂a ), hence a
[︁ ]︁ (︁ )︁ E Zπ2∗ (X) ≤ 2 log(2K) + 2 max λ1 (Σ̂a ). a∈A
(C.15)
Now recall that to simplify, we also assumed that Σa = σ02 Id for any a ∈ A and that Σ = τ 2 Id′ . As a result, we have that: Σ̂a = Σ̃a + σ0−4 Σ̃a Wa Σ̄Wa⊤ Σ̃a , and we also have that λ1 (Σ̃a ) ≤ σ02 and that λ1 (Σ̂a ) ≤ σ02 + τ 2 ,
∀a ∈ A,
Plugging (C.14) (or (C.15)) into (C.11) yields [︂ ]︂ √︂ Bso(π̂g ) ≤ αE ∥ϕ(X)∥Σ̂π (X) + (2 log(2K) + 2)(σ02 + τ 2 )δ, ∗
169
(C.16)
(C.17)
√︁ with α2 = d + 2 d log(K/δ) + 2 log(K/δ) as defined above. Choosing δ = 1/n in (C.17) yields √︃ [︂ ]︂ (2 log(2K) + 2)(σ02 + τ 2 ) , (C.18) Bso(π̂g ) ≤ αn E ∥ϕ(X)∥Σ̂π (X) + ∗ n where √︁ αn2 = d + 2 d log(Kn) + 2 log(Kn). (C.19) This concludes the proof.
C.3.2
Explicit Bound
Additional simplification. To further simplify the exposition, we just set ϕ(x) = x for any x ∈ X . Assumption 9. Let G = E[XX ⊤ ] with g = λd (G). We assume that g > 0. Assumption 10 (Context-independent logging policy). A is independent of X, i.e., π0 (a | x) = pa for all x and a. Equivalently, (Xi ) are i.i.d. ∼ ν and independent of (Ai ), with P(A = a) = pa . Theorem 8 (Explicit Bound). Let π∗ (x) be the optimal action for context x. Then, the BSO of sDM under the structured prior Equation (6.9) satisfies √︄ [︃ ]︃ (︂ (︁ 2 )︁gmX /2 )︂ 1 τ2 2 /2 2 −nρ + (σ0 + τ 2 ) (d + 1)e X + d e + Bso(π̂g ) ≤ αn E λX σ04 λ2X √︃ (2 log(2K) + 2)(σ02 + τ 2 ) + , n where √︂ √︁ αn = d + 2 d log(Kn) + 2 log(Kn), and ρX = pπ∗ (X) ,
mX =
⌊︂ nρ ⌋︂ X
2
,
λX = σ −2 g
mX + σ0−2 . 2
Scaling with n. Recall mX = ⌊nρX /2⌋ and λX = σ −2 g m2X +σ0−2 , where ρX = π0 (π∗ (X)). For n large enough (so that mX is roughly nρX ), we have λX = Θ(nρX + 1) and the exponentially small tail terms can be ignored at the level of leading-order scaling. Consequently, up to absolute constants and polylogarithmic factors, )︄ (︄ √︄ [︃ ]︃ √︃ 1 log K , + Bso(π̂g ) = Õ αn E nρX + 1 n √ and using αn = Θ̃( d) this can be summarized as (︄√︄ [︃ )︄ ]︃ √︃ log K d Bso(π̂g ) = Õ E + , nρX + 1 n 170
Proof. Let’s focus on the main term of Theorem 2, and we start by bounding [︂ ]︂ E ∥X∥2Σ̂ , π∗ (X)
(C.20)
Now recall that to simplify, we also assumed that Σa = σ02 Id for any a ∈ A and that Σ = τ 2 Id′ . As a result, we have that: Σ̂a = Σ̃a + σ0−4 Σ̃a Wa Σ̄Wa⊤ Σ̃a , and we also have that λ1 (Σ̃a ) ≤ σ02 and that λ1 (Σ̂a ) ≤ σ02 + τ 2 ,
∀a ∈ A,
(C.21)
since λ1 (Wa Wa⊤ ) = 1. These are obtained using Weyl’s inequalities. Now let Na =
n ∑︂
1{Ai = a},
pa = P(A = a) = π0 (a),
ρx = pπ∗ (x) ,
ρX = pπ∗ (X) .
i=1
Define mx =
⌊︂ nρ ⌋︂ x
2
,
mX =
⌊︂ nρ ⌋︂ X
2
.
Then, under Assumption 10, Na ∼ Bin(n, pa ) and it is independent of context X. Therefore, Hoeffding gives, for t = nρX /2, (︃ )︃ (︂ )︂ nρX ⃓⃓ nρ2X P Nπ∗ (X) < . (C.22) ⃓X, θ∗ ≤ exp − 2 2 Let {︁ }︁ ΩX,1 = Nπ∗ (X) ≥ mX . Then we have that )︃ (︃ 2 ⃓ )︁ nρ X . P Ω̄X,1 ⃓X, θ∗ ≤ exp − 2
(C.23)
(︁
m Lemma 7 (Matrix Chernoff, Tropp (2012, Theorem ∑︁m 1.1)). Let (Yk )k=1 be independent PSD matrices with λ1 (Yk ) ≤ R a.s. Let µmin = λd ( k=1 E[Yk ]). Then for δ ∈ [0, 1], (︄ )︄ ]︃µmin /R [︃ m (︂ ∑︂ )︂ e−δ P λd Yk ≤ (1 − δ)µmin ≤ d . 1−δ (1 − δ) k=1
Define {︄ ΩX,2 =
λd
n (︂ ∑︂ i=1
1{Ai = π∗ (X)}Xi Xi⊤
)︂
1 ≥ Nπ∗ (X) g 2
We now bound P(Ω̄X,2 | X, θ∗ ). Recall that {︁ }︁ ΩX,1 = Nπ∗ (X) ≥ mX , 171
}︄
mX =
ΩX = ΩX,1 ∩ ΩX,2 .
,
⌊︂ nρ ⌋︂ X
2
.
Under Assumption 10, (Ai ) is independent of (Xi ), hence conditional on the index set Sπ∗ (X) = {i ∈ [n] : Ai = π∗ (X)}, the matrices {Xi Xi⊤ : i ∈ Sπ∗ (X) } are i.i.d. with the same law as XX ⊤ . Moreover, since the Chernoff bound below depends on Sπ∗ (X) only through |Sπ∗ (X) | = Nπ∗ (X) , the same bound holds when conditioning on Nπ∗ (X) . Moreover, we have that ϕ(x) = x and that ∥ϕ(x)∥ ≤ 1. Thus, ∥x∥ ≤ 1 and hence λ1 (Xi Xi⊤ ) ≤ 1. Applying Lemma 7 with R = 1, E[XX ⊤ ] = G, and (︁ )︁ µmin = λd Nπ∗ (X) G = Nπ∗ (X) g, and choosing δ = 21 , we obtain the conditional bound P(Ω̄X,2 | X, Nπ∗ (X) , θ∗ ) ≤ d
(︂√︁ )︂gNπ∗ (X) 2/e .
Taking conditional expectation Nπ∗ (X) given X and θ∗ , [︃(︂ ]︃ √︁ )︂gNπ∗ (X) [︁ ]︁ 2/e P(Ω̄X,2 | X, θ∗ ) = E P(Ω̄X,2 | X, Nπ∗ (X) , θ∗ ) | X, θ∗ ≤ dE | X, θ∗ . Splitting on ΩX,1 yields ]︃ [︃(︂ ]︃ [︃(︂ ]︃ [︃(︂ √︁ )︂gNπ∗ (X) √︁ )︂gNπ∗ (X) √︁ )︂gNπ∗ (X) 2/e 2/e 2/e | X, θ∗ = E 1{ΩX,1 } | X, θ∗ + E 1{Ω̄X,1 } | X, θ∗ E (︂√︁ )︂gmX 2/e + P(Ω̄X,1 | X, θ∗ ), ≤ (︂√︁ )︂gNπ∗ (X) (︂√︁ )︂gmX since on ΩX,1 we have Nπ∗ (X) ≥ mX and thus 2/e 2/e ≤ , while on (︂√︁ )︂gNπ∗ (X) 2/e ≤ 1. Therefore, Ω̄X,1 we have (︂√︁ )︂gmX 2/e + dP(Ω̄X,1 | X, θ∗ ).
P(Ω̄X,2 | X, θ∗ ) ≤ d
(C.24)
Combining (C.24) with (C.22) yields (︃ )︃ (︂√︁ )︂gmX nρ2X P(Ω̄X | X, θ∗ ) ≤ P(Ω̄X,1 | X, θ∗ ) + P(Ω̄X,2 | X, θ∗ ) ≤ (d + 1) exp − +d 2/e . 2 (C.25) Finally, let I1 = E
[︂
∥X∥2Σ̂ 1{ΩX } | X, θ∗ π∗ (X)
Then
[︂ E ∥X∥2Σ̂
]︂
,
π∗ (X)
I2 = E ]︂
[︂
∥X∥2Σ̂ 1{Ω̄X } | X, θ∗ π∗ (X)
= E[I1 ] + E[I2 ].
Using ∥X∥2 ≤ 1 and (C.16), ∥X∥2Σ̂
π∗ (X)
= X ⊤ Σ̂π∗ (X) X ≤ λ1 (Σ̂π∗ (X) )∥X∥22 ≤ σ02 + τ 2 . 172
]︂
.
Let c1 = σ02 + τ 2 . Then (︃ I2 ≤ c1 P(Ω̄X | X, θ∗ ) ≤ c1
nρ2X (d + 1) exp − 2 (︃
)︃
(︂√︁ )︂gmX )︃ +d 2/e .
(C.26)
Moreover, fix X and θ∗ , on ΩX , λd
n (︂ ∑︂
)︂ 1 1 1{Ai = π∗ (X)}Xi Xi⊤ ≥ Nπ∗ (X) g ≥ mX g, 2 2 i=1
so −2
λd (Ĝπ∗ (X) ) = σ λd
n (︂ ∑︂
1{Ai = π∗ (X)}Xi Xi⊤
)︂
i=1
Hence
1 ≥ σ −2 mX g. 2
1 λd (Ĝπ∗ (X) + σ0−2 Id ) ≥ σ −2 mX g + σ0−2 . 2
Let λX = σ −2 g
mX + σ0−2 . 2
Since Σ̃π∗ (X) = (Ĝπ∗ (X) + σ0−2 Id )−1 , we obtain on ΩX : λ1 (Σ̃π∗ (X) ) ≤
1 . λX
Moreover, on ΩX we have that λ1 (Σ̂π∗ (X) ) ≤ λ1 (Σ̃π∗ (X) ) + σ0−4 τ 2 λ1 (Σ̃π∗ (X) )2 ≤
1 σ −4 τ 2 + 02 . λX λX
Therefore, since ∥X∥2 ≤ 1, (︃ I1 ≤
1 σ −4 τ 2 + 02 λX λX
)︃ .
(C.27)
Combining (C.26) and (C.27), [︂ E
∥X∥2Σ̂ π∗ (X)
]︂
(︃ (︃ )︃ (︂√︁ )︂gmX )︃]︃ 1 nρ2X σ0−4 τ 2 2 2 ≤E + + (σ0 + τ ) (d + 1) exp − +d 2/e . λX λ2X 2 [︃
But from Theorem 2, we know that [︂
]︂
√︃
Bso(π̂g ) ≤ αn E ∥X∥Σ̂π (X) + ∗
√︂ √︁ where αn = d + 2 d log(Kn) + 2 log(Kn). 173
(2 log(2K) + 2)(σ02 + τ 2 ) , n
(C.28)
[︂
Finally, by Jensen’s inequality, E ∥X∥Σ̂π (X)
]︂
∗
√︃ [︂ E ∥X∥2Σ̂ ≤
π∗ (X)
]︂ . Substituting into
(C.28) and combining with (C.26) and (C.27), we obtain √︄ [︃ ]︃ (︂ (︁ 2 )︁gmX /2 )︂ 1 τ2 2 −nρ2X /2 2 +d e + (σ0 + τ ) (d + 1)e Bso(π̂g ) ≤ αn E + λX σ04 λ2X √︃ (2 log(2K) + 2)(σ02 + τ 2 ) , + n √︂ √︁ where αn = d + 2 d log(Kn) + 2 log(Kn).
C.3.3
Optimality of Greedy Policies
Here, we show that Greedy policy π̂g should be preferred to any other choice of policies when considering the BSO as our performance metric. This is because π̂g minimizes the BSO. To see this, note that by definition the Greedy policy π̂g is deterministic, that is for any context x ∈ X , there exists âg , such that π̂g (âg | x) = 1. Thus, for any context x ∈ X , we simplify the notation by letting π̂g (x) denote the action that has a mass equal to 1. Then, we have that EA∼π̂g (·|x) [Eθ∗ [r(x, A; θ∗ ) | Dn ]] = Eθ∗ [r(x, π̂g (x); θ∗ ) | Dn ] , ≥ E [r(x, a; θ∗ ) | Dn ]
∀x, a ∈ X × A.
(C.29) (C.30)
where this follows from the definition of r̂(x, a) = E [r(x, a; θ) | Dn ], the definition of π̂g and the fact that θ∗ is sampled from the prior, which leads to E [r(x, a; θ) | Dn ] = E [r(x, a; θ∗ ) | Dn ]. Now Equation (C.29) holds for any x ∈ X and a ∈ A, and hence it holds in expectation under X ∼ ν and A ∼ π(· | X) for any policy π. That is, EX∼ν,A∼π̂g (·|X) [Eθ∗ [r(x, A; θ∗ ) | Dn ]] ≥ EX∼ν,A∼π(·|X) [Eθ∗ [r(x, A; θ∗ ) | Dn ]] .
(C.31)
Taking another expectation w.r.t. the sample set Dn and using Fubini’s theorem and the tower rule leads to E [V (π̂g ; θ∗ )] ≥ E [V (π; θ∗ )] for any stationary policy π. Then, subtracting E [V (π∗ ; θ∗ )] from both sides of the previous inequality yields that the BSO is minimized by π̂g compared to any stationary policy π, in particular, compared to the policy πp induced by pessimism.
C.4
Additional Experiments
As mentioned in Section C.4, our experiments were conducted on internal machines with 30 CPUs and thus they required a moderate amount of computation. These experiments are also reproducible with minimal computational resources.
C.4.1
Implementation Details of Baselines
We implement the baselines as follows. 174
• IPS. n
π(Ai |Xi ) 1 ∑︂ Ri , argmax n i=1 max{π0 (Ai |Xi ), τ } π
(C.32)
where τ ∈ [0, 1] is a hyper-parameter. • snIPS. argmax ∑︁n π
n ∑︂ π(Ai |Xi )
1
π(Ai |Xi ) π (Ai |Xi ) i=1 π0 (Ai |Xi ) i=1 0
Ri ,
(C.33)
• MIPS. We cluster actions into L groups and let h(a) be the cluster of action a. Let Ci be the cluster of action Ai , then we use n
1 ∑︂ π(Ci |Xi ) argmax Ri , (C.34) n i=1 π0 (Ci |Xi ) π ∑︁ where Ci = h(Ai ) for any i ∈ [n], and π(c|x) = a∈A 1 [h(a) = c] π(a|x). • PC. We use the Knn implementation of PC. Let N (a, k) be the set of k-nearest neighbors of a, then n ∑︁ 1 [a ∈ N (Ai , k)] π(a|x) 1 ∑︂ ∑︁ a∈A Ri . (C.35) argmax n i=1 a∈A 1 [a ∈ N (Ai , k)] π0 (a|x) π • DM (Freq). This DM uses the linear-Gaussian likelihood model R | θ, X, A ∼ N (ϕ(X)⊤ θA , σ 2 ) and learn the parameters θa using the maximum likelihood principle leading to (C.36)
r̂(x, a) = ϕ(x)⊤ µ̂a ,
∑︁ where the MLE is µ̂a = (Ga + λId )−1 Ba , with Ga = i∈[n] I{Ai =a} ϕ(Xi )ϕ(Xi )⊤ and ∑︁ Ba = i∈[n] I{Ai =a} Ri ϕ(Xi ), and λ is a regularization hyper-parameter. • DM (Bayes). This DM uses the linear-Gaussian likelihood model combined with Gaussian priors as θa ∼ N (µa , Σa ), ⊤
∀a ∈ A,
(C.37)
2
R | θ, X, A ∼ N (ϕ(X) θA , σ ), Under this prior, each action a has an associated parameter θa . Given the prior in Equation (6.2), the posterior distribution of an action parameter follows a multivari−1 −1 ate Gaussian: θa | ∑︁ Dn ∼ N (µ̂a , Σ̂a ), where Σ̂−1 Σ̂−1 a = Σa +Ga and∑︁ a µ̂a = Σa µa +Ba . Here, Ga = σ −2 i∈[n] I{Ai =a} ϕ(Xi )ϕ(Xi )⊤ and Ba = σ −2 i∈[n] I{Ai =a} Ri ϕ(Xi ). Then, the reward estimate is r̂(x, a) = ϕ(x)⊤ µ̂a ,
(C.38)
• DR. n
1 ∑︂ π(Ai |Xi ) (Ri − r̂(Xi , Ai )) + EA∼π(·|Xi ) [r̂(Xi , A)] , argmax n i=1 max{π0 (Ai |Xi ), τ } π with τ ∈ [0, 1] and r̂ is the reward model obtained using DM (Freq). 175
(C.39)
C.4.2
Robustness to Likelihood Misspecification
We strengthened our evaluation by assessing sDM’s robustness to likelihood misspecification below (robustness to prior misspecification is provided in Section C.4.3). In these experiments, the true data-generating process (same as the synthetic experiments in Section 6.5) differed from sDM’s assumptions in two different ways: either the likelihood is misspecified
Avg. relative reward
Misspecified likelihood (Figure C.1). We also simulate when the true reward distribution differed from the likelihood assumed by sDM. For example, we simulated binary rewards using a Bernoulli-logistic model while sDM used a linear-Gaussian likelihood. Other DMs: DM (Bayes) and DM (Freq) also use a misspecified likelihood model and to emphasize this we add the suffix Lin to all DMs names. Overall, sDM still outperforms all methods by a large margin despite misspecification.
Miss. Likelihood - OPL K=100, d'=5, d=10
1.0
Miss. Likelihood - OPL K=1000, d'=5, d=10
Miss. Likelihood - OPL K=1000, d'=10, d=10
Miss. Likelihood - OPL K=1000, d'=20, d=10
0.9 0.8 0.7 0.6 0.5
0
500
1000 1500 20000
Number of samples n sDM (Ours)
DM (Bayes)
500
1000 1500 20000
Number of samples n DM (Freq)
500
1000 1500 20000
Number of samples n DR
IPS
500
1000 1500 2000
Number of samples n snIPS
MIPS
PC
Figure C.1: Effect of likelihood misspecification: The relative reward of the learned policy on synthetic problems using misspecified likelihood with varying n and K.
C.4.3
Robustness to Prior Misspecification
Misspecified prior means and covariances(Figure C.2). This is achieved by adding uniformly sampled noise from [v, v + 0.5] to both the true prior mean and covariance parameters µ, Σ, Wa , Σa , with v controlling the level of misspecification. We varied v ∈ {0.5, 1, 1.5} and analyzed its impact on sDM’s performance. For comparison, we included the well-specified sDM and the most competitive baseline, DM (Bayes), while omitting other baselines to reduce clutter. sDM’s performance decreases with increasing misspecification, yet sDM with misspecification still outperforms the most competitive baseline, especially when K is large. We also observe that the impact of prior covariance misspecification is less significant compared to prior mean misspecification. 176
Avg. relative reward
Mis. Prior - OPL K=100, d'=5, d=10
1.0
Mis. Prior - OPL K=1000, d'=5, d=10
Mis. Prior - OPL K=1000, d'=10, d=10
Mis. Prior - OPL K=1000, d'=20, d=10
0.9 0.8 0.7 0.6 0.5
0
500
1000 1500 20000
500
Number of samples n sDM
1000 1500 20000
Number of samples n sDM (v=0.5)
500
1000 1500 20000
Number of samples n
sDM (v=1)
sDM (v=1.5)
500
1000 1500 2000
Number of samples n DM (Bayes)
Figure C.2: Effect of prior mean and covariance misspecification: The average relative reward of the learned policy on synthetic problems using both misspecified prior means and covariances with varying n and K and d′ .
C.4.4
Comparison of Greedy and Pessimistic Policies
To validate our theory that a greedy policy should be preferred over the commonly adopted pessimistic policy in our Bayesian setting, we used a performance metric averaged over multiple bandit problems sampled from the prior. To verify this, we considered the same OPL synthetic setting as in Section 6.5 and compared sDM with a greedy policy to sDM with a pessimistic policy. Recall that a greedy policy with respect to our reward estimate writes (C.40)
π̂g (a | x) = 1{a = argmax r̂(x, b)}, b∈A
while a pessimistic one writes (C.41)
π̂p (a | x) = 1{a = argmax r̂(x, b) − u(x, a)}, b∈A
where u(x, a) = α(d, δ)∥ϕ(X)∥Σ̂a with α(d, δ) =
√︃ d+2
√︂ d log 1δ + 2 log 1δ . As predicted
Avg. relative reward
by our theory, the results show that the greedy policy has better average performance over multiple bandit instances sampled from the prior.
1.00 0.95 0.90 0.85 0.80 0.75 0.70 0.65 0.60
Synthetic problems K=100, d'=5, d=10
0
500
1000 1500 20000
Number of samples n
Synthetic problems K=1000, d'=5, d=10
500
1000 1500 20000
Number of samples n sDM (Greedy)
Synthetic problems K=1000, d'=10, d=10
500
1000 1500 20000
Number of samples n
Synthetic problems K=1000, d'=20, d=10
500
1000 1500 2000
Number of samples n
sDM (Pessimism)
Figure C.3: Comparison of sDM with greedy policy and sDM with pessimistic policy in OPL: The average MSE of an ϵ-greedy policy on synthetic problems with varying n, K, and d′ .
177
Chapter D
Supplementary Materials for Chapter 7
Contents C.1 Posterior Derivations Under Standard Priors . . . . . . . . . . . . . . 161 C.2 Posterior Derivations Under Structured Priors . . . . . . . . . . . . . 162 C.2.1
Latent Posterior . . . . . . . . . . . . . . . . . . . . . . . . . .
163
C.2.2
Conditional Posterior . . . . . . . . . . . . . . . . . . . . . . .
165
C.2.3
Action Posterior . . . . . . . . . . . . . . . . . . . . . . . . . .
166
C.3 Proofs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 167 C.3.1
Main Result . . . . . . . . . . . . . . . . . . . . . . . . . . . .
167
C.3.2
Explicit Bound . . . . . . . . . . . . . . . . . . . . . . . . . .
170
C.3.3
Optimality of Greedy Policies . . . . . . . . . . . . . . . . . .
174
C.4 Additional Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . 174 C.4.1
Implementation Details of Baselines . . . . . . . . . . . . . . .
174
C.4.2
Robustness to Likelihood Misspecification . . . . . . . . . . .
176
C.4.3
Robustness to Prior Misspecification . . . . . . . . . . . . . .
176
C.4.4
Comparison of Greedy and Pessimistic Policies . . . . . . . . .
177
D.1
Proofs for Oracle Policies
D.1.1
Oracle Policies for IPS-Based Objectives
(IPS), cIPS and ES. Recall the definition of the (logging propensity) clipped IPS estimator with τ ∈ [0, 1]: n
1 ∑︂ π(Ai |Xi ) V̂cips (π) = Ri . n i=1 max{π0 (Ai |Xi ), τ } 178
Taking n → ∞, one obtains: [︃
]︃ π(A|X) Vcips (π) = EX∼ν,A∼π0 (·|X) r(X, A) max{π0 (A|X), τ } [︃ ]︃ π0 (A|X) r(X, A) . = EX∼ν,A∼π(·|X) max{π0 (A|X), τ } As the objective is linear in the policy π, the optimal policy should put for any x ∈ X , all the mass on the action a that maximizes the weighted reward, giving: [︂ π0 (a′ |x)r(x, a′ ) ]︂ π∗cIPS (a|x) = 1 a = argmax . max{π0 (a′ |x), τ } a′ ∈A We recover the solution for IPS when we let τ → 0: [︂ ]︂ π∗IPS (a|x) = 1 a = argmax r(x, a′ )1 [π0 (a′ |x) > 0] . a′ ∈A
We also recover the solution of ES just by replacing the clipping function by an exponential function of factor α, obtaining: [︂ ]︂ π∗ES (a|x) = 1 a = argmax r(x, a′ )π0 (a′ |x)1−α . a′ ∈A
Doubly Robust (DR). The doubly robust estimator converges to the following quantity: ]︃ [︃ π0 (A|X) + r̂(X, A) . Vdr (π) = EX∼ν,A∼π(·|X) (r(X, A) − r̂(X, A)) max{π0 (A|X), τ } The objective is linear in π and is thus maximized by the following deterministic decision rule: [︃ ]︃ π0 (a′ |x) DR ′ ′ ′ π∗ (a|x) = 1 a = argmax r̂(x, a ) + (r(x, a ) − r̂(x, a )) max{π0 (a′ |x), τ } a′ ∈A Marginalized IPS (MIPS) with clusters. We adopt the same approach to look for the maximizer of MIPS. We generalize the clustering function h to also account for context. We write down the estimator: ∑︁ n n ′ ′ 1 ∑︂ π(Ci |Xi ) 1 ∑︂ ′ 1 [h(a , Xi ) = h(Ai , Xi )] π(a |Xi ) a ∑︁ Ri = V̂mips (π) = Ri , n i=1 a′′ 1 [h(a′′ , Xi ) = h(Ai , Xi )] π0 (a′′ |Xi ) n i=1 π0 (Ci |Xi ) with which, we recover when n → ∞: [︃ ∑︁ ]︃ ′ ′ 1 [h(a , X) = h(A, X)] π(a |X) ′ Vmips (π) = EX∼ν,A∼π0 (·|X) ∑︁ a r(X, A) ′′ ′′ a′′ 1 [h(a , X) = h(A, X)] π0 (a |X) [︄ ]︄ ∑︁ ′ ′ ∑︂ 1 [h(a , X) = h(a, X)] π(a |X) ′ = EX∼ν π0 (a|X) ∑︁ a r(X, a) ′′ ′′ a′′ 1 [h(a , X) = h(a, X)] π0 (a |X) a ]︄ [︄ ′ ∑︂ ∑︂ 1 [h(a , X) = h(a, X)] = EX∼ν π(a′ |X) π0 (a|X) ∑︁ r(X, a) ′′ ′′ a′′ 1 [h(a , X) = h(a, X)] π0 (a |X) a a′ [︄ ]︃]︄ [︃ ′ ∑︂ 1 [h(a , X) = h(A, X)] r(X, A) . = EX∼ν π(a′ |X)EA∼π0 (·|X) EA′′ ∼π0 (·|X) [1 [h(A′′ , X) = h(A, X)]] a′ 179
The objective is linear in π, and depends on the action a′ through its cluster h(a′ , ·) alone. This means that multiple solutions are maximizers as long as the policy chooses the best cluster c. We thus write down the oracle policy for MIPS in the cluster level, giving: ]︃ }︂]︂ [︃ [︂ {︂ r(x, A)1[h(A, x) = c′ ] MIPS π∗ (c|x) = 1 c = argmax EA∼π0 (·|x) EA′′ ∼π0 (·|x) [1 [h(A′′ , x) = h(A, x)]] c′ ∈C ]︃ }︂]︂ [︃ [︂ {︂ r(x, A)1[h(A, x) = c′ ] = 1 c = argmax EA∼π0 (·|x) EA′′ ∼π0 (·|x) [1 [h(A′′ , x) = c′ ]] c′ ∈C [︂ {︂ E ′ }︂]︂ A∼π0 (·|x) [r(x, A)1[h(A, x) = c ]] = 1 c = argmax , EA∼π0 (·|x) [1 [h(A, x) = c′ ]] c′ ∈C which ends the proof. Conjunct Effect Modeling (OffCEM). This estimator can be seen as the natural, doubly robust extension of the MIPS estimator. Combining similar techniques to the ones employed for MIPS and DR yields [︁ ]︁ }︂]︂ ′ [︂ {︂ E (r(x, Ā) − r̂(x, Ā)) 1 [h(x, Ā) = h(x, a )] Ā∼π (·|x) 0 π∗OffCEM (a|x) = 1 a = argmax r̂(x, a′ ) + . ′ π0 (h(x, a′ )|x) a ∈A Two Stage Decomposition (POTEC). This is an optimization strategy for OffCEM. It restricts the policy to a cluster-informed form, ∑︂ π rm (a | x, c)π cl (c | x), π(a | x) = c∈C
where π rm (a | x, c) = 1[a = argmaxa′ ∈c r̂(x, a′ )] is fixed, model-based policy that deterministically selects the best action within each cluster. Learning is then simplified to finding the optimal cluster-level policy π cl that maximizes the OffCEM objective: )︄ (︄ n cl ∑︂ ∑︂ 1 π (C | X ) i i V̂potec (π cl ) = π cl (c | Xi )r̂c∗ (Xi ) , (Ri − r̂(Xi , Ai )) + n i=1 π0 (Ci | Xi ) c∈C where r̂c∗ (x) = maxa∈c r̂(x, a) is the estimated reward of the best action in cluster c. This is exactly the Doubly Robust version of MIPS on the cluster level, the oracle policy on the cluster level can be followed in the same fashion: [︂ {︂ E }︂]︂ ′ A∼π0 (·|x) [(r(x, A) − r̂(x, A))1[h(A, x) = c ]] cl ∗ π∗ (c | x) = 1 c = argmax + r̂c′ (x) . EA∼π0 (·|x) [1 [h(A, x) = c′ ]] c′ ∈C The optimal policy for the POTEC optimization strategy unfolds as: ∑︂ π rm (a | x, c)π∗cl (c | x) . π∗POTEC (a|x) = c∈C
At first glance, it might be hard to see the connection between POTEC and OffCEM solutions, but they are equivalent. For ease of notation, let us denote by Dr̂,x (c): Dr̂,x (c) =
EA∼π0 (·|x) [(r(x, A) − r̂(x, A))1[h(A, x) = c]] . EA∼π0 (·|x) [1 [h(A, x) = c]] 180
and recall that the optimal policy of OffCEM finds the action a that maximizes: Ṽ (x, a) = r̂(x, a) + Dr̂,x (h(a, x)) . For any context x, the optimal action a∗ of POTEC verifies: • a∗ is in the optimal cluster: h(a∗ , x) = c∗ (x) with c∗ (x) = argmaxc∈C Dr̂,x (c) + r̂c∗ (x). • a∗ is optimal within that cluster: a = argmaxa∈c∗ (x) r̂(x, a). This means that for all actions a with h(a, x) ̸= c∗ (x), we have: Ṽ (x, a) = Dr̂,x (h(a, x)) + r̂(x, a) ∗ ≤ Dr̂,x (h(a, x)) + r̂h(a,x) (x) ≤ Dr̂,x (c∗ (x)) + r̂c∗∗ (x) (x) = Dr̂,x (h(x, a∗ )) + r̂(x, a∗ ) = Ṽ (x, a∗ ) . In addition, for all actions a with h(a, x) = c∗ (x), we have: Ṽ (x, a) = Dr̂,x (h(a, x)) + r̂(x, a) = Dr̂,x (c∗ (x)) + r̂(x, a) ≤ Dr̂,x (c∗ (x)) + r̂c∗∗ (x) (x) = Ṽ (x, a∗ ) . This means that the optimal action a∗ for POTEC is the maximizer of Ṽ (x, a), which is exactly the solution of OffCEM. Policy Convolution (PC). This estimator uses a nearest neighbors function to aggregate the propensities of similar actions, making the hypothesis that similar actions will result in similar reward signal. The estimator writes: n
V̂pc (π) =
1 ∑︂ π(Nϵ (Ai ) | Xi ) Ri , n i=1 π0 (Nϵ (Ai ) | Xi )
with π(Nϵ (a) | x) =
∑︂
π(a′ | x).
a′ ∈Nϵ (a)
This estimator is equivalent to the following when n → ∞:
1 [a′ ∈ Nϵ (A)]
′ a′ π(a |X)
[︃ ∑︁
PC
V (π) = EX∼ν,A∼π0 (·|X) [︄ = EX∼ν,A∼π(·|X)
]︃
r(X, A) π0 (Nϵ (A)|X) [︄ [︁ ]︁ ]︄]︄ r(x, Ā)1 A ∈ Nϵ (Ā) EĀ∼π0 (·|X) . π0 (Nϵ (Ā)|X)
The same argument of linearity applies here, giving us the corresponding oracle policy: π∗PC (a|x) =
[︃ ]︃ {︂ r(x, Ā)1[a′ ∈ Nϵ (Ā)] }︂]︂ 1 a = argmax EĀ∼π0 (·|x) . π0 (Nϵ (Ā)|x) a′ ∈A [︂
181
D.1.2
Oracle Policies for PWLL-Based Objectives
Our objectives can be written in the same form, only choosing for each a different function g: n
Ûg (π) =
1 ∑︂ g(Xi , Ai , Ri ) log π(Ai | Xi ) . n i=1
Since we are looking at oracle policies, we consider the expectation Ug (π) = EX∼ν, A∼π0 (·|X), R∼p(·|X,A) [g(X, A, R) log π(A | X)] . The maximization decomposes over contexts. Fix x and define the nonnegative weights wx (a) = ER∼p(·|x,a) [g(x, a, R)] ≥ 0. For each x, we thus consider ∑︂ max π0 (a | x) wx (a) log π(a | x) π(·|x)
a∈A
∑︂
s.t.
π(a | x) = 1,
∀a ∈ A, π(a | x) ≥ 0.
a∈A
Let vx (a) = π0 (a | x) wx (a) ≥ 0. The Lagrangian (with equality multiplier λ ∈ R and inequality multipliers {µ(a)}a∈A , µ(a) ≥ 0) is (︂ ∑︂ )︂ ∑︂ ∑︂ L(π, λ, µ) = vx (a) log π(a | x) + λ π(a | x) − 1 + µ(a) π(a | x). a∈A
a∈A
a∈A
By KKT conditions, at an optimum π∗g (· | x) we have for all a ∈ A: ∂L vx (a) = + λ + µ(a) = 0, ∂π(a | x) π(a | x)
and µ(a) π(a | x) = 0.
For any action with π∗g (a | x) > 0, we get that µ(a) = 0, and hence vx (a) . λ ∑︁ ∑︁ Normalizing with a π∗g (a | x) = 1 gives λ = − a′ vx (a′ ) and therefore π∗g (a | x) = −
π0 (a | x) ER∼p(·|x,a) [g(x, a, R)] vx (a) = ∑︁ . ′ ′ ′ a′ ∈A vx (a ) a′ ∈A π0 (a | x) ER∼p(·|x,a′ ) [g(x, a , R)]
π∗g (a | x) = ∑︁ This concludes the proof.
D.2
Proofs for Optimization Properties
In this section, we prove the propositions about the optimization landscape of IPS-based and PWLL learning approaches. We start by stating the following lemmas, that will be helpful to prove our propositions. 182
Lemma 8. (Mei et al., 2020b, Lemma 2) Consider the single context case. With a slight abuse of notation, we drop the dependence on x and write r(a) instead of r(x, a), πθ (a) instead of πθ (a|x), and r̂(a) instead of r̂(x, a). Let πθ be a softmax policy parameterized by θ. Then, for any r̂ ∈ [0, 1]K , and any estimator V̂ linear in πθ , the mapping θ ↦→ V̂ (πθ ) = ⟨r̂, πθ ⟩ is 5/2-smooth. Lemma 9. All the action level estimators EST in (IPS, cIPS, DR, PC) can be written, for any policy π, in the form: n
V̂est (π) =
1 ∑︂ EA∼π(·|Xi ) [r̂est,i (A, Xi )] , n i=1
(D.1)
For the cluster level estimators/approaches EST-C in (MIPS, OffCEM, POTEC), we also have n
1 ∑︂ V̂est-c (π) = EC∼π(·|Xi ) [r̂est-c,i (C, Xi )] , n i=1
(D.2)
meaning that all these estimators are linear in π. Proof. This is straightforward to prove. We begin by the action level estimators and take DR as a representative. For DR, we have the following: r̂DR,i (a, Xi ) = r̂(a, Xi ) + I[a = Ai ]
Ri − r̂(Ai , Xi ) max(τ, π0 (Ai |Xi ))
verifies the equation. Solutions for cIPS and IPS can be recovered directly, and PC follows the same construction. For the cluster level approaches, we take POTEC as a representative, and we have: r̂POTEC,i (c, Xi ) = r̂c⋆ (Xi ) + I[c = Ci ]
Ri − r̂(Ai , Xi ) , π0 (Ci |Xi )
The r̂MIPS,i follows as a special case when r̂ = 0. Lemma 10. Consider the single-context case and assume a finite action set A. For any estimator EST in (IPS, cIPS, DR, OffCEM, MIPS, PC), there exists a problem instance (i.e., a choice of r and π0 ; and when relevant, a choice of auxiliary objects such as r̂, h, Nϵ ) such that, in the large-n limit, r̂est (a) = 1[a = aK ] for some optimal action aK . Similarly, for cluster-based approaches (e.g., POTEC and MIPS), there exists an instance such that r̂est-c (c) = 1[c = c|C| ] for some optimal cluster c|C| . Proof. We give explicit constructions for cIPS (action-level) and POTEC (cluster-level). The other estimators follow by the same idea: choose a setting where the estimator becomes linear in π with some deterministic coefficient, and pick r (and possibly r̂, h, Nϵ ) so that the resulting linearized reward is one-hot. 183
Action-level: cIPS. Fix τ ∈ [0, 1).1 Choose a logging policy π0 with full support and such that max π0 (a) ≥ τ, (D.3) a∈A
Let
π0 (a) . a∈A max{π0 (a), τ }
aK ∈ arg max
Under Equation (D.3), there exists at least one action with π0 (a) ≥ τ , for which the ratio equals 1, hence the maximizer satisfies π0 (aK ) ≥ τ and therefore π0 (aK ) = 1. max{π0 (aK ), τ } Now define the reward function r(a) = 1[a = aK ]
max{π0 (a), τ } . π0 (a)
This satisfies r(a) ∈ [0, 1] for all a because r(a) = 0 for a ̸= aK , and r(aK ) =
max{π0 (aK ), τ } = 1 (since π0 (aK ) ≥ τ ). π0 (aK )
For cIPS, the large-n linearized reward is r̂cips (a) = hence r̂cips (a) =
π0 (a) r(a), max{π0 (a), τ }
π0 (a) max{π0 (a), τ } 1[a = aK ] = 1[a = aK ], max{π0 (a), τ } π0 (a)
as desired. Cluster-level: POTEC. We work in the single-context case and consider a clustering map h : A → C. Choose h so that aK forms a singleton cluster: c|C| = h(aK ) = {aK },
h(a) ̸= c|C| ∀a ̸= aK .
Let rewards be r(aK ) = 1 and r(a) = 0 for a ̸= aK . Pick any ε ∈ (0, 1/2] and define a reward model r̂(aK ) = 1 − ε, r̂(a) = ε ∀a ̸= aK . For POTEC, the induced (cluster-level) linearized reward takes the form (︁ )︁ ∑︁ π (a) r(a) − r̂(a) 0 r̂POTEC (c) = max r̂(a) + a∈c . a∈c π0 (c) 1
If τ = 1 and |A| > 1, the simplifying assumption π0 (a) > 0 for all a is incompatible with having some π0 (a) ≥ τ . In practice τ ≪ 1.
184
For the singleton cluster c|C| = {aK }, we get (︁ )︁ π0 (aK ) 1 − (1 − ε) r̂POTEC (c|C| ) = (1 − ε) + = (1 − ε) + ε = 1. π0 (aK ) For any other cluster c ̸= c|C| , all its actions satisfy r(a) = 0 and r̂(a) = ε, hence ∑︁ π0 (a)(−ε) r̂POTEC (c) = ε + a∈c = ε − ε = 0. π0 (c) Therefore r̂POTEC (c) = 1[c = c|C| ]. This concludes the constructions for cIPS and POTEC. The remaining estimators can be handled analogously by choosing π0 (and when relevant, h or Nϵ ) so that the estimator’s linear coefficient on r(a) equals 1 at a chosen aK (and equals something finite elsewhere), and then defining r (and possibly r̂) to make the resulting r̂est one-hot. Now we restate Proposition 3 and proceed to its proof. Proposition 8 (Plateau for linear-in-π objectives under softmax). Consider the singlecontext case. Let V̂ (π) be any objective linear in π and let π ⋆ ∈ arg maxπ V̂ (π) denote a maximizer over the probability simplex on the effective action space Aeff (of size Keff = |Aeff |). Let {πθt }t≥1 be the∑︁iterates of gradient ascent on θ ↦→ V̂ (πθ ) with a linear softmax policy πθ (a) = exp(θa )/ a′ ∈Aeff exp(θa′ ) and step sizes ηt ∈ (0, 1]. Then there exists a problem instance such that gradient ascent cannot escape a suboptimal region before t0 = C Keff = O(Keff ) iterations, in the sense that V̂ (π ⋆ ) − V̂ (πθt ) ≥ 0.9.
∀t ≤ t0 :
Proof. The proof follows the same technique as (Mei et al., 2020a, Theorem 1). By Lemma 10, there exists an instance (single context) for which the linearized reward is one-hot: r̂est (a) = 1[a = aK ] for some aK ∈ Aeff . Hence, for any policy π supported on Aeff , ∑︂ V̂ (π) = π(a)r̂est (a) = π(aK ). a∈Aeff
The maximizer over the simplex is therefore π ⋆ = δaK , and V̂ (π ⋆ ) = 1,
V̂ (π ⋆ ) − V̂ (πθ ) = 1 − πθ (aK ).
(Notice that supθ V̂ (πθ ) = 1 as well, although the supremum is not attained by any finite θ when Keff ≥ 2). We now upper bound the gradient norm. For the softmax parametrization, ⃦ ⃦ √ (︁ )︁ ⃦ ⃦ 2 πθ (aK ) 1 − πθ (aK ) , ∇ V̂ (π ) ⃦ θ θ ⃦ ≤ 2
where the bound follows by a direct computation (as in (Mei et al., 2020a)). 185
Define the update θt+1 = θt + ηt ∇θ V̂ (πθt ) and split iterations into tgood = {t ≥ 1 : πθt+1 (aK ) > πθt (aK )}, For t ∈ tbad ,
tbad = {t ≥ 1 : πθt+1 (aK ) ≤ πθt (aK )}.
1 1 − ≤ 0. πθt (aK ) πθt+1 (aK )
For t ∈ tgood , using Lemma 8 (the 5/2-smoothness of θ ↦→ V̂ (πθ )) and ηt ∈ (0, 1], we obtain 9 πθt+1 (aK ) − πθt (aK ) ≤ πθt (aK )2 , 2 and therefore (since πθt+1 (aK ) ≥ πθt (aK ) > 0), πθ (aK ) − πθt (aK ) 1 1 9 = t+1 ≤ . − πθt (aK ) πθt+1 (aK ) πθt+1 (aK )πθt (aK ) 2 Summing over s = 1, . . . , t − 1 yields )︃ t−1 (︃ ∑︂ 1 1 1 1 9 = − − ≤ t. πθ1 (aK ) πθt (aK ) πθs (aK ) πθs+1 (aK ) 2 s=1 Assume a standard symmetric initialization so that πθ1 (aK ) = 1/Keff . Pick any constant 2 Keff , then c ≥ 11 and take Keff large enough so that πθ1 (aK ) ≤ 1/c. If t ≤ 9c )︃ (︃ 1 1 1 9 1 ≥ − t≥ ≥ c − 1 ≥ 10, 1− πθt (aK ) πθ1 (aK ) 2 πθ1 (aK ) c hence πθt (aK ) ≤ 1/10, and thus V̂ (π ⋆ ) − V̂ (πθt ) = 1 − πθt (aK ) ≥ 0.9. 2 This proves the claim with t0 = 9c Keff .
Proposition 9. Even for a single context x, deterministic rewards, there is problem where IPS-based learning with a linear softmax policy πθ (a) ∝ exp(⟨θ, ϕ(x, a)⟩)I[a ∈ Aeff ] can have a number of local maxima exponential in the number of effective actions Keff . Proof. Let EST an off-policy estimators considered in the paper with an action-level policy. By Lemma 9, we have: n
1 ∑︂ V̂est (π) = Ea∼π(·|xi ) [r̂est,i (a, xi )] , n i=1
(D.4)
In a single context setting, it becomes: [︄ V̂est (πθ ) = Ea∼πθ (·)
]︄ n 1 ∑︂ r̂est,i (a) , n i=1
(D.5)
n
1 ∑︂ =⟨ r̂est,i , πθ ⟩ . n i=1 186
(D.6)
This also holds for estimators with policies in the cluster level, as we still have: ]︄ [︄ n 1 ∑︂ r̂est-c,i (c) , V̂est-c (πθ ) = Ec∼πθ (·) n i=1
(D.7)
n
=⟨
1 ∑︂ r̂est-c,i , πθ ⟩ . n i=1
(D.8)
These softmax policies are all defined on the effective action space Aeff , be it a subset of the action space A or the discrete cluster space C. Using the linearity of the objective, we can directly apply Theorem 1 from Chen et al. (2019) and obtain our result. Finally, we also restate Proposition 5, and provide its proof. Proposition 10. For an ℓ2 regularized (substituting λ2 ||θ||2 , with λ > 0), linear softmax policy πθ , the PWLL objective Û g (πθ ) defined as: n
1 ∑︂ Û (π) = g(Ri , π0 (Ai | Xi )) log π(Ai | Xi ) , n i=1 g
is λ-strongly concave. Without regularization, the objective is concave. Proof. For any x and a ∈ Aeff (x), we have: exp(⟨θ, ϕ(x, a)⟩) , ′ a′ ∈Aeff (x) exp(⟨θ, ϕ(x, a )⟩)
πθ (a|x) = ∑︁
optimizing an ℓ2 regularized linear softmax, giving: λ L̂g,λ (π) = Û g (π) − ||θ||2 , 2 with λ > 0 and recall that g ≥ 0. For strong concavity, we need to show that the Hessian ∇2θ Û g (πθ ) is negative definite with eigenvalues bounded away from zero. ∑︁ The gradient with respect to θ is: ∇θ Û g (πθ ) = n1 ni=1 g(Ri , π0 (Ai | Xi ))∇θ log πθ (Ai |Xi )− λθ For the softmax policy: ∇θ log πθ (a|x) = ϕ(x, a) −
∑︂
πθ (a′ |x)ϕ(x, a′ ) = ϕ(x, a) − EA∼πθ (·|x) [ϕ(x, A)]
a′
Therefore: ∇θ Û g (πθ ) = n1
∑︁n
i=1 g(Ri , π0 (Ai | Xi ))
(︁
)︁ ϕ(Xi , Ai ) − EA∼πθ (·|Xi ) [ϕ(Xi , A)] − λθ
∑︁ Taking the second derivative: ∇2θ Û g (πθ ) = − n1 ni=1 g(Ri , π0 (Ai | Xi ))∇θ EA∼πθ (·|Xi ) [ϕ(Xi , A)]− λId , where Id is the d × d identity matrix. The gradient of the expectation is: ∑︂ ∇θ EA∼πθ (·|x) [ϕ(x, A)] = ∇θ πθ (a|x)ϕ(x, a) a
187
Using ∇θ πθ (a|x) = πθ (a|x)(ϕ(x, a) − EA∼πθ (·|x) [ϕ(x, A)]): ∇θ EA∼πθ (·|x) [ϕ(x, A)] =
∑︂
πθ (a|x)(ϕ(x, a) − EA∼πθ (·|x) [ϕ(x, A)])ϕ(x, a)⊤
a
This simplifies to: ∇θ EA∼πθ (·|x) [ϕ(x, A)] = CovA∼πθ (·|x) [ϕ(x, A)] where CovA∼πθ (·|x) [ϕ(x, A)] = EA∼πθ (·|x) [ϕ(x, A)ϕ(x, A)⊤ ]−EA∼πθ (·|x) [ϕ(x, A)]EA∼πθ (·|x) [ϕ(x, A)]⊤ Therefore: 1 ∇2θ Û g (πθ ) = −
n ∑︂
n i=1
g(Ri , π0 (Ai | Xi ))CovA∼πθ (·|Xi ) [ϕ(Xi , A)] − λId
We can write this as: ∇2θ Û g (πθ ) = −H − λId ∑︁ where H = n1 ni=1 g(Ri , π0 (Ai | Xi ))CovA∼πθ (·|Xi ) [ϕ(Xi , A)] is positive semi-definite. To see this explicitly, for any vector v ∈ Rd : v ⊤ CovA∼πθ (·|Xi ) [ϕ(Xi , A)]v = VarA∼πθ (·|Xi ) [v ⊤ ϕ(Xi , A)] ≥ 0 , with the positivity of g, this ensures H is positive semi-definite. Then we have: v ⊤ ∇2θ Û g (πθ )v = −v ⊤ Hv − λv ⊤ v = −v ⊤ Hv − λ∥v∥2 , meaning that when v ̸= 0, we get v ⊤ ∇2θ Û g (πθ )v ≤ −λ∥v∥2 < 0 . This shows the Hessian is negative definite with all eigenvalues bounded above by −λ < 0. Therefore, ℓ2 regularized Û g (πθ ) is λ-strongly concave. In addition, when λ = 0, the hessian is negative semi-definite, giving simple concavity.
D.3
Stochastic Optimization Convergence Guarantees for PWLL
We analyze the convergence rates of stochastic gradient methods on the PWLL objective. We formulate this as the minimization of the finite-sum loss f (θ) = −Ûg (πθ ): n
f (θ) =
1 ∑︂ fi (θ), n i=1
where fi (θ) = −gi log πθ (Ai | Xi ),
(D.9)
where gi = g(Ri , π0 (Ai |Xi )). We adopt the linear softmax policy parametrization in Equation (7.16) with sθ (x, a) = ϕ(x, a)⊤ θ (lightweight parametrization in Equation (7.17)). We note that our analysis extends naturally to the heavyweight parametrization in Equation (7.17). 188
D.3.1
Assumptions and Regularity
To establish problem-dependent convergence bounds, we rely on the following structural assumptions regarding the feature space and the importance weights. Assumption 11 (Bounded features). For all context-action pairs (x, a) ∈ X × A, the feature representations are bounded in Euclidean norm: ∥ϕ(x, a)∥2 ≤ H. Assumption 12 (Bounded weighting function). The weights gi = g(Ri , π0 (Ai |Xi )) computed on the static dataset are strictly positive and bounded. That is, for all i ∈ {1, . . . , n}: 0 < gi ≤ Gmax . Assumptions 11 and 12 are sufficient to establish the smoothness and bounded variance of the objective f (θ). We formally derive these properties in the following proposition. Proposition 11 (Regularity and Variance Bounds). Under Assumptions 11 and 12, the objective f (θ) satisfies the following properties: 1. Global Smoothness: The objective is L̄-smooth with L̄ = Gmax H 2 . 2. Bounded Single-Sample Variance: The variance of the stochastic gradient for a single sample is bounded by σ̄ 2 = 4G2max H 2 . 3. Bounded Mini-Batch Variance: For a mini-batch of size b, the variance is 2 H2 . bounded by σ̄b2 = 4Gmax b Proof. 1. Smoothness: The Hessian of the objective is the weighted sum of the feature covariance matrices under the policy πθ : n
1 ∑︂ ∇ f (θ) = gi CovA∼πθ (·|Xi ) [ϕ(Xi , A)]. n i=1 2
The spectral norm of a covariance matrix is bounded by the maximum squared ∑︁n norm2 of 1 2 its random vectors. Thus, using Assumption 11 we get that ∥∇ f (θ)∥op ≤ n i=1 gi H ≤ Gmax H 2 . 2. Single-Sample Variance: We first bound the norm of the gradient for an arbitrary sample i. The gradient is ∇fi (θ) = −gi (ϕ(Xi , Ai ) − EA∼πθ (·|Xi ) [ϕ(Xi , A)]). Using the triangle inequality and Assumption 11: (︁ )︁ ∥∇fi (θ)∥2 ≤ gi ∥ϕ(Xi , Ai )∥2 + ∥EA∼πθ (·|Xi ) [ϕ(Xi , A)]∥2 ≤ Gmax (H + H) = 2Gmax H. Let ξ = ∇fI (θ) be the stochastic gradient sampled uniformly from the dataset. The variance is bounded by the second moment: n
Var(ξ) ≤ E[∥ξ∥2 ] =
1 ∑︂ ∥∇fi (θ)∥2 ≤ (2Gmax H)2 = 4G2max H 2 . n i=1 189
∑︁ 3. Mini-Batch Variance: Let the mini-batch gradient be ḡt = 1b bj=1 ∇fij (θ), where indices are sampled independently with replacement. Using the standard variance reduction property for independent variables: 1 4G2max H 2 E[∥ḡt − ∇f (θ)∥2 ] = E[∥∇fI (θ) − ∇f (θ)∥2 ] ≤ . b b
Based on Proposition 11, we define the following global problem-dependent constants on which our convergence rates depend: • L̄ = Gmax H 2 : Smoothness constant. • σ̄ 2 = 4G2max H 2 : Upper bound on the gradient variance for a single sample. 2 H2 : Upper bound on the gradient variance for a mini-batch of size b. • σ̄b2 = 4Gmax b
D.3.2
PWLL without ℓ2 regularization
We begin by analyzing the standard unregularized PWLL objective. Here, the objective f (θ) is convex but not necessarily strongly convex. This implies the loss landscape may contain multiple minimizers rather than a unique global minimum. Consequently, we characterize convergence in terms of Ûg (πθopt ) − Ûg (πθ̄T ) (instead of ∥θt − θnopt ∥). Here, θopt ∈ arg maxθ Ûg (πθ ) is an optimal parameter and θ̄T is the average of the SGA iterates. Proposition 12. Let θopt ∈ arg maxθ Ûg (πθ ) be an optimal parameter. If the learning rate satisfies 0 < η ≤ 41L̄ , then by (Garrigos and Gower, 2023, Theorem 6.9), the iterates of mini-batch SGA satisfy: [︂
E Ûg (π
θopt
]︂ ∥θ − θopt ∥2 8ηG2 H 2 0 max ) − Ûg (πθ̄T ) ≤ + ηT b
where θ̄T is the average of the iterates. Proposition 12 highlights the trade-off inherent to constant step-size SGA: a larger η accelerates the decay of the initial error (first term) but increases the asymptotic noise √ T) floor (second term). For a fixed horizon T , one can recover a convergence rate of O(1/ √ by setting η ∝ 1/ T , which balances both terms.
D.3.3
PWLL with ℓ2 regularization
We now move to the ℓ2 -regularized case where the PWLL objective is strongly concave (Proposition 5). Precisely, we consider the regularized objective Ũ λ (θ) = Ûg (πθ ) − λ2 ∥θ∥2 . Strong convexity implies the existence of a unique global minimizer. This allows us to guarantee convergence of the parameters θt themselves, which is a stronger condition than value convergence. 190
opt Proposition 13. Let θn,λ = arg maxθ Ũ λ (θ) be the unique optimal parameter. If the learning rate satisfies 0 < η ≤ 2(Gmax1H 2 +λ) , then by (Garrigos and Gower, 2023, Theorem 6.12): [︁ ]︁ 8ηG2max H 2 opt 2 opt 2 E ∥θt − θn,λ ∥ ≤ (1 − ηλ)t ∥θ0 − θn,λ ∥ + λb
The regularized case demonstrates a convergence rate that is significantly faster than the rate of the unregularized case.
D.4
Additional Experiments
D.4.1
Detailed Experimental Setting
Experimental Setting. Table D.1: Statistics of Post Processed Datasets Dataset
Num. of actions
Num. of samples
MovieLens
60, 000
132, 744
Twitch
200, 000
400, 000
GoodReads
1, 000, 000
400, 000
Our experimental setup is designed to study the behavior of the different policy learning paradigms in large action spaces. To this end, we use three large action spaces collaborative filtering datasets: Movielens (Lam and Herlocker, 2016), Twitch (Rappaz et al., 2021) and GoodReads (Wan et al., 2019) that are preprocessed to obtain a user-item interaction matrix. We follow the exact procedure of Sakhi et al. (2023) to pre-process the datasets. The statistics of the obtained datasets are described in Table D.1. For each user, we keep half of its history as the context x, and use the other half of the history as the products with positive reward, which align the learned policies to recommend new and relevant items. We direct the interested readers to Sakhi et al. (2023) for a detailed description of the experimental setup. The large action space scenario restricts the policies used to the inner product parametrization (Aouali et al., 2022a). This parametrization is essential to leverage Maximum Inner Product Search algorithms (Shrivastava and Li, 2014) for fast query response. In particular, we adopt policies of the following form: πθ (a|x) ∝ exp(⟨ϕΓ (x), βa ⟩) , with the learnable parameter θ = [Γ, β], ϕΓ : X → Rℓ defines the context embedding function in Rℓ and β the actions embeddings of size K × ℓ. To define our policies, we start by extracting action embeddings β0 using an SVD decomposition of the user-item matrix. These embeddings help us define the context embedding function ϕΓ and our 191
logging policy π0 . ϕ0 is set to the average embeddings of the observed actions in the contexts and is fixed for the logging policy π0 . Using the SVD action embeddings β0 , we define our logging policy π0 as: (︃ )︃ ]︁ [︁ 1 π0 (a|x) ∝ exp ⟨ϕ0 (x), β0,a ⟩ I a ∈ topk0 (x) , t with t the temperature of the logging policy, and k0 define the support of the logging policy, concentrating on the top k0 actions with: topk0 (x) = argsorta1 ,··· ,ak ⟨ϕ0 (x), β0,a ⟩. 0
If not explicitly stated, k0 is set to 100 and the temperature at t = 1 in all experiments. This policy is used to collect the offline dataset Dn = {Xi , Ai , Ri }i∈[n] on which all trainings are conducted. For each i ∈ [n] in the processed dataset, Xi is the user history, Ai is the action played by the logging policy π0 (·|Xi ) and Ri = 1[Ai ∈ Hi ] the observed reward, which is if the action played is in the hidden items of user i. Trained Policies Parameterizations. We adopt two parameterizations of the trained policies. The first one is a heavyweight parametrization, and focuses on learning the embeddings of the actions β (be it A of size K or C of size |C|), meaning that θ in this case is β. For action-level policies, this gives β ∈ RK×ℓ and for any x: πβ (a|x) = ∑︁
exp(⟨ϕ0 (x), βa ⟩) , a′ ∈Aeff (x) exp(⟨ϕ0 (x), βa′ ⟩)
with Aeff (x) ⊂ A, which depends on the choice of the practitioner, for example Aeff (x) = S0 (x), the support of π0 for context x when we optimize IPS objectives. For cluster-level policies, this gives a β ∈ R|C|×ℓ and for any x: exp(⟨ϕ0 (x), βc ⟩) . c′ ∈C exp(⟨ϕ0 (x), βc′ ⟩)
πβ (c|x) = ∑︁
This is used by default if nothing is explicitly stated. We have also define a lightweight parametrization, where only a small projection W ∈ Rℓ×ℓ is learned, giving in action level policies: πW (a|x) = ∑︁
exp(⟨ϕ0 (x)W, βa,0 ⟩) , a′ ∈Aeff (x) exp(⟨ϕ0 (x)W, βa′ ,0 ⟩)
using β0 , ∑︁ the embeddings of π0 . For cluster level policies, we first define β̄0 ∈ R|C|×ℓ with 1 β̄0,c = |c| a∈c β0,a , and use it to define the cluster level policy: exp(⟨ϕ0 (x)W, β̄c,0 ⟩) . c′ ∈C exp(⟨ϕ0 (x)W, β̄c′ ,0 ⟩)
πW (c|x) = ∑︁
Reward Model. The reward model used r̂ is learned using regularized linear regression the collected interaction data, with r̂(x, a) = ⟨ϕ(x), θa ⟩. Clustering and ϵ used. We use the embeddings β0 , combined with K-means clustering to find our clusters. The number of clusters is set to 2000 for all datasets and experiments. For PC, the ℓ2 threshold ϵ is set to 0.1. 192
D.4.2
Additional results
Benefits of objective-aware parametrization. Figure D.1 shows the effect of objectiveaware policy parameterizations for two different objectives and three large action space datasets.
Effect of Objective-Aware Parametrization on Performance Reward - MovieLens (K=60K)
LR Schedule: None
Reward - Twitch (K=200K)
LR Schedule: One Cycle
0.35
0.35
0.30
0.30
0.30
0.25
0.25
0.25
0.20
0.20
0.20
0.15
0.15
0.15
0.10 24
0.10 25
26
27
28
24
0.10 25
27
26
28
24
0.30
0.30
0.30
0.25
0.25
0.25
0.20
0.20
0.20
0.15
0.15
0.15
0.10
0.10
0.10
0.05
0.05
0.05
0.00
0.00
24
Reward - GoodReads (K=1M)
LR Schedule: Warmup Cosine
0.35
25
26
27
28
24
27
26
28
24
0.35
0.35
0.30
0.30
0.30
0.25
0.25
0.25
0.20
0.20
0.20
0.15
0.15
0.15
0.10
0.10
0.10
0.05
0.05
0.05
0.00
0.00 25
26
Batch Size IPS (Objective-Aware Defined on Support)
27
28
24
26
27
28
25
26
27
28
26
27
28
0.00 25
0.35
24
25
0.00 25
27
26
Batch Size IPS (Whole Action Space)
28
24
25
Batch Size
cLPI (Objective-Aware Defined on Support)
cLPI (Whole Action Space)
Figure D.1: The effect of objective-aware parametrization for IPS and cLPI on three large-scale datasets
Average MSE. Figure D.2 shows the average MSE by dataset and method. Several methods are excluded from the figure, as their high MSE values would distort the scale and obscure the comparison. 193
Average MSE by Dataset and Method
0.020
IPS ES DR
Average MSE
0.015
MIPS PC OffCEM
0.010
0.005
0.000
)
0K
s
Mo
n Le vie
=6 (K
K)
00
tch
i Tw
2 K=
M)
=1
K s(
(
d ea dR
o Go
Figure D.2: Average MSE by Dataset and Method. Several methods are excluded from the figure, as their high MSE values would distort the scale and obscure the comparison. MSE progress during training. Figures D.3a to D.3c show the progress of the MSE over 10 epochs on all three datasets. Several methods are excluded from the figure, as their high MSE values would distort the scale and obscure the comparison. MSE Evolution - MovieLens (K=60K)
0.08
IPS ES DR
0.07 0.06
0.030
0.05
0.020
0.03
0.015
0.02
0.010
0.01
0.005 1
2
3
4
5
6
7
Epoch
(a) MovieLens
8
9
10
0.000
IPS ES DR
0.06
MSE
0.04
MSE Evolution - GoodReads (K=1M)
0.07
MIPS PC OffCEM
0.025
MSE
MSE
IPS ES DR
0.035
0.05
0.00
MSE Evolution - Twitch (K=200K)
0.040
MIPS PC OffCEM
MIPS PC OffCEM
0.04 0.03 0.02 0.01
1
2
3
4
5
6
7
8
9
Epoch
(b) Twitch
10
0.00
1
2
3
4
5
6
7
8
9
10
Epoch
(c) GoodReads
Figure D.3: MSE progression over 10 epochs across datasets.
D.4.3
Results Averaging Different Seeds
In this experiment, we analyze the reward evolution of representative PWLL and IPSbased methods on the three considered datasets. We compare two distinct optimization configurations: (i) a standard off-the-shelf Adam optimizer, and (ii) a carefully tuned setup using Adam with an optimized batch size and a one-cycle learning-rate scheduler. This comparison enables us to isolate the effect of optimization on stability and convergence. Each method is evaluated over 5 random seeds, and we report the mean reward along with a shaded standard deviation region to visualize sensitivity to optimization randomness. In Figure D.4, across all datasets and optimization settings, we observe that IPS-based methods (cIPS, IX, and even POTEC) not only reach inferior performance but also suffer 194
from considerably higher variance. Their uncertainty bands are significantly wider, indicating unstable optimization. In contrast, PWLL-based methods exhibit near-invisible variance bands, with standard deviations roughly an order of magnitude smaller on average.
GoodReads
0.35 0.30
0.25
0.25
0.20
0.20
0.15
0.15
0.10
0.10
0.05
0.05
0.00
0.00
−0.05
0.30
0.30
0.35
0.25
0.25
0.20
0.20
0.15
0.15
0.10
0.10
0.05
0.05
0.25 0.20 0.15 0.10 0.05 0.00
0.30
Off-the-shelf Optimizer
Validation reward Validation reward
Twitch
0.30
Optimizer with Best Scheduler
MovieLens
0.30
0.25 0.20 0.15 0.10 0.05
0.00
1
2
3
4
5
6
7
8
9
10
0.00
0.00 1
2
Epoch
3
4
5
6
7
8
9
10
−0.05
Epoch IPS
DR
ES
1
2
3
4
5
6
7
8
9
10
Epoch POTEC
cLPI
Figure D.4: cLPI vs IPS-Based methods: Evolution of rewards averaged over 5 different seeds. cLPI is more stable to optimize and reaches better policies.
Finally, in Figure D.5, we observe that adopting an Objective-Aware parametrization yields further performance and stability improvements. For example, cIPS with ObjectiveAware parametrization surpasses cIPS while maintaining lower variability, and cLPI in its Objective-Aware form consistently achieves the best overall performance. These results demonstrate that the combination of PWLL objectives and clever parametrization leads to more robust and more effective learned policies, while being very simple to implement. 195
GoodReads
0.35 0.30
0.25
0.25
0.20
0.20
0.15
0.15
0.10
0.10
0.05
0.05
0.05
0.00
0.00
0.00
0.30
0.30
0.35
0.25
0.30
0.25 0.20 0.15
0.25
0.10
Off-the-shelf Optimizer
Validation reward Validation reward
Twitch
0.30
Optimizer with Best Scheduler
MovieLens
0.30
0.25
0.20 0.20
0.20 0.15 0.15
0.15 0.10 0.10
0.05
0.10
0.05
1
2
3
4
5
6
7
8
9
10
0.00
0.05 1
2
3
4
Epoch
5
6
7
8
9
10
0.00
1
2
3
4
Epoch
IPS (Objective-Aware Defined on Support)
IPS (Whole Action Space)
cLPI (Objective-Aware Defined on Support)
5
6
7
8
9
10
Epoch
cLPI (Whole Action Space)
Figure D.5: Objective Aware Parametrisation: Evolution of rewards averaged over 5 different seeds. Objective Aware Parametrization stabilizes and improves performance for PWLL and IPS methods.
D.4.4
Ablation - Sensitivity to Reward Noise
In the original evaluation setup (see Section D.4), the observed reward is deterministic; for user i, we have Ri = 1[Ai ∈ Hi ], meaning that a positive reward is returned only when the selected action belongs to the user’s hidden set Hi . In this section, we investigate robustness to reward noise by introducing stochasticity in the form: Ri ∼ 1[Ai ∈ Hi ] (1 − B(ϵ)) + B(ϵ) s, where B(ϵ) is a Bernoulli random variable with parameter ϵ, and s ∈ [0, 1] is a shift. This results in noisy rewards supported on [0, 1]. Note that any reward scaling can be normalized to this range via R/Rmax when Rmax > 1. We evaluate six configurations defined by noise parameters ϵ ∈ {0.1, 0.2, 0.3} and reward shifts s ∈ {0, 0.5}. All methods are trained using the best-performing optimization schedule (one-cycle) to isolate the effect of noise. Results are reported in Figure D.6. We observe that increasing the noise level when s = 0 consistently harms all methods, as expected from a more stochastic reward signal. In contrast, when s = 0.5, higher noise tends to increase the overall reward level, since the shift raises the baseline reward. Across all noise–shift conditions, PWLL-based objectives maintain a clear advantage over IPS-based methods. When s = 0, RegKL and cLPI perform similarly, confirming that both benefit from the logarithmic reparameterization. However, as both noise and shift increase, RegKL begins to outperform cLPI, suggesting that, with an appropriately chosen regularization weight β, RegKL remains highly competitive even under reward high stochasticity. Conclusion. PWLL methods demonstrate robustness to reward noise, leading to improved stability and performance compared to traditional IPS-based objectives, even in 196
challenging noise regimes.
Movielens — Noise=0.1
Shift = 0.0
Validation Reward
0.25
Movielens — Noise=0.2
Movielens — Noise=0.3
Twitch — Noise=0.1
Twitch — Noise=0.2
Twitch — Noise=0.3
0.20 0.15 0.10 0.05 0.00
Shift = 0.5
Validation Reward
0.35 0.30 0.25 0.20 0.15 0.10 0.05
1
2
3
4
5
6
Epoch
7
8
9
10
1
2
3
4
5
6
7
8
9
10
1
2
3
Epoch
4
5
6
7
8
9
10
1
Epoch IPS
DR
ES
2
3
4
5
6
7
8
9
10
Epoch POTEC
cLPI
1
2
3
4
5
6
7
Epoch
8
9
10
1
2
3
4
5
6
7
8
9
10
Epoch
RegKL
Figure D.6: Ablation - Sensitivity to reward noise
D.4.5
Ablation Study on Hyperparameters and Log Transform
In this section, we evaluate the impact of hyperparameter choices on cIPS, cLPI, RegKL, and RegKL-LIN (the non-logarithmic variant of RegKL). All methods are run using the best-performing optimization configuration (optimizer + learning rate scheduler), ensuring that differences are driven solely by hyperparameter values and by whether the policy transformation is linear or logarithmic. The results are shown in Figure D.7.
• cIPS consistently fails to reach competitive performance across all values of τ , especially in large action spaces. Its PWLL counterpart, cLPI, dominates for every τ , converging faster and achieving superior results. • For the KL-based objectives, we restrict to β ≥ 0.1 in order to avoid numerical instability from the exponential term (exp(1/β) > 2 · 105 for β < 0.1). The same trend is observed: the PWLL variant (RegKL) reliably outperforms its linear analogue (RegKL-LIN) across all β, exhibiting more stable training dynamics, faster convergence and better performance.
PWLL dominates. Across both objective families, replacing linear weights with logtransformed policy weights (PWLL) consistently provides greater robustness to hyperparameters, faster optimization, and higher final performance, even more in challenging large-action-space settings. 197
Reward trajectories: RegKL (solid) vs RegKL-LIN (dashed) across ¯ Reward trajectories: cLPI (solid) vs cIPS (dashed) across ¿
Validation reward
Twitch
Validation reward
Movielens
0.35
Goodreads
0.30 0.25 0.20 0.15 0.10
1 2 3 4 5 6 7 8 9 10
1 2 3 4 5 6 7 8 9 10
1 2 3 4 5 6 7 8 9 10
Epoch
Epoch
Epoch
¿ = 0:001
¿ = 0:01
Goodreads
1 2 3 4 5 6 7 8 9 10
1 2 3 4 5 6 7 8 9 10
1 2 3 4 5 6 7 8 9 10
Epoch
Epoch
Epoch
0.25 0.20 0.15 0.10 0.05 0.00
cLPI (solid)
Twitch
0.30
0.05 0.00
Movielens
0.35
RegKL (solid) ¯ = 0:1
cIPS (dashed)
¿ = 0:05
¿ = 0:1
¿ = 0:5
¿=1
RegKL-LIN (dashed)
¯ = 0:2
¯ = 0:4
¯ = 0:8
¯=1
(b) RegKL-LIN (linear) vs RegKL (PWLL) w.r.t β
(a) cIPS (linear) vs cLPI (PWLL) w.r.t τ
Figure D.7: Ablation Study on hyper-parameters and Log Transform
D.4.6
Ablation: PWLL in Smaller Action Spaces
We have shown that PWLL provides a more benign optimization landscape and yields stronger policies than IPS-based objectives in large action spaces. Here, we examine whether these benefits also extend to smaller action spaces. We construct a reduced version of Movielens by subsampling the action space to K ∈ {100, 500, 1000, 5000} items. Figure D.8 reports performance across varying K for cIPS (linear) and its PWLL-enhanced counterpart cLPI (log). In small action space settings (K ≤ 500), cIPS convergences faster than cLPI, but cLPI identifies a better maxima by the end of the 10 epochs. For medium action spaces (K ≥ 500), cLPI consistently outperforms cIPS, converging faster and identifying a better maximum. These results indicate that the optimization advantages of PWLL can still be beneficial in medium sized action space settings. MovieLens — Comparison IPS (OPE) vs cLPI (PWLL) Varying Action Space Size K K = 100
Validation reward
0.7
K = 500
0.30
0.6
K = 1000
0.30
K = 5000
0.30
0.25
0.25
0.25
0.4
0.20
0.20
0.20
0.3
0.15
0.15
0.15
0.10
0.10
0.10
0.05
0.05
0.5
0.2 0.1 0.0
1
2
3
4
5
6
Epoch
7
8
9
10
1
2
3
4
5
6
7
8
9
10
1
Epoch
2
3
4
5
6
7
8
9
10
0.05
1
2
3
Epoch IPS
4
5
6
7
8
9
10
Epoch
cLPI
Figure D.8: PWLL (cLPI) vs IPS-based (IPS) in smaller action spaces.
D.4.7
Ablation - Sensitivity to the number of clusters
In this study, we compare our simple PWLL objective cLPI against MIPS and POTEC, two more complex IPS-based methods specifically designed for large action spaces. These baselines rely on a clustering function to reduce variance, and POTEC additionally leverages a reward model r̂. We examine how the number of clusters affects their optimization performance. Figure D.9 reports the results. 198
POTEC generally outperforms MIPS for all numbers of clusters. However, both methods exhibit optimization instability across settings. While POTEC can occasionally match the final performance of cLPI on Movielens for a carefully selected number of clusters (1000), it consistently falls short on Twitch regardless of the cluster configuration.
Conclusion. These findings demonstrate that focusing on optimization properties pays off: despite its simplicity, the PWLL objective cLPI can consistently outperform intricate IPS-based approaches tailored to large action spaces, even with the best finetuning.
Movielens
0.30
Twitch
Validation Reward
0.25 0.20 0.15 0.10 0.05 0.00
0
1
2
3
4
5
6
7
8
9
0
1
2
3
Epoch
5
6
7
8
9
Epoch MIPS
Nb Clusters = 10
4
POTEC
Nb Clusters = 100
cLPI
Nb Clusters = 1000
Nb Clusters = 10000
Figure D.9: PWLL (cLPI) vs POTEC and MIPS, changing the number of clusters.
D.4.8
Ablation - Different Logging Supports
We conduct experiments to quantify how increasing or restricting the support of the logging policy affects policy learning, comparing PWLL and IPS-based methods. Figure D.10 compiles the results and show that PWLL is still better than IPS-based approaches for different logging support sizes. 199
MovieLens
klog = 10 Validation reward
0.30
Twitch
GoodReads
0.25 0.20 0.15 0.10 0.05 0.00 0.35
klog = 100 Validation reward
0.30 0.25 0.20 0.15 0.10 0.05 0.00
klog = 1000 Validation reward
0.30 0.25 0.20 0.15 0.10 0.05 0.00
1
2
3
4
5
6
7
8
9
10
1
2
3
Epoch
4
5
6
7
8
9
10
1
2
3
4
Epoch IPS
5
6
7
8
9
10
Epoch
cLPI
Figure D.10: PWLL vs IPS: Different Logging Support sizes klog
D.4.9
Pessimism Does not Solve Optimization Problems
Pessimism in face of uncertainty Jin et al. (2021) is motivated through a pure statistical learning rationale and is used to provide better statistical guarantees and more controlled excess risk. In the context of OPL, pessimistic strategies are derived combining concentration bounds with class complexity measures, be it VC dimension (Swaminathan and Joachims, 2015a) or PAC-Bayesian tools (London and Sandler, 2019; Aouali et al., 2023a; Sakhi et al., 2024). For example, in its PAC-Bayesian formulation, the pessimistic objectives are all written in the following form:
arg max πθ
V̂ (πθ ) −
λ ||θ − θ0 ||22 , n
Adding an ℓ2 regularization term that pulls the parameters θ towards the behavior policy parameters θ0 (defining π0 ) induces pessimism by encouraging the learned policy to stay close to π0 in parameter space. However, the optimization landscape of this objective becomes concave only when the regularization weight λ is sufficiently large for the ℓ2 term to dominate. In that regime, the objective is indeed easier to optimize, but becomes overly conservative, yielding policies that remain too close to π0 and under-exploit potential improvements. Figure D.11 confirms this empirically: pessimistic approaches, whether based on Sample Variance Penalisation (SVP) (Swaminathan and Joachims, 2015a), PAC-Bayesian learning with clipped IPS (London and Sandler, 2019), Exponential Smoothing (Aouali et al., 2023a), or Logarithmic Smoothing (Sakhi et al., 2024), fail to outperform cLPI for any value of λ. 200
MovieLens
Validation reward
0.30
Twitch
0.30
0.25
0.25
0.20
0.20
0.15
0.15
0.10
0.10
0.05
0.05
GoodReads-1M
0.35 0.30 0.25 0.20 0.15
0.00
1
2
3
4
5
6
7
8
9
10
0.00
0.10 0.05 1
Epoch
2
3
4
5
6
7
8
9
10
0.00
1
Epoch
2
3
4
5
6
7
8
9
10
Epoch
IPS (PAC-Bayes) (¸ = 0:001)
IPS (PAC-Bayes) (¸ = 10)
LS (PAC-Bayes) (¸ = 1)
ES (PAC-Bayes) (¸ = 0:1)
IPS (PAC-Bayes) (¸ = 0:01)
LS (PAC-Bayes) (¸ = 0:001)
LS (PAC-Bayes) (¸ = 10)
ES (PAC-Bayes) (¸ = 1)
IPS (PAC-Bayes) (¸ = 0:1)
LS (PAC-Bayes) (¸ = 0:01)
ES (PAC-Bayes) (¸ = 0:001)
ES (PAC-Bayes) (¸ = 10)
IPS (PAC-Bayes) (¸ = 1)
LS (PAC-Bayes) (¸ = 0:1)
ES (PAC-Bayes) (¸ = 0:01)
IPS (SVP) (¸ = 0:001) IPS (SVP) (¸ = 0:01) IPS (SVP) (¸ = 0:1)
IPS (SVP) (¸ = 1) IPS (SVP) (¸ = 10) cLPI
Figure D.11: cLPI outperforms the pessimistic approaches. Some methods do not appear in the plot because their curves overlap.
201
Chapter E
Supplementary Materials for Chapter 8
Contents D.1 Proofs for Oracle Policies . . . . . . . . . . . . . . . . . . . . . . . . . 178 D.1.1
Oracle Policies for IPS-Based Objectives . . . . . . . . . . . .
178
D.1.2
Oracle Policies for PWLL-Based Objectives . . . . . . . . . .
182
D.2 Proofs for Optimization Properties . . . . . . . . . . . . . . . . . . . . 182 D.3 Stochastic Optimization Convergence Guarantees for PWLL . . . . . 188 D.3.1
Assumptions and Regularity . . . . . . . . . . . . . . . . . . .
189
D.3.2
PWLL without ℓ2 regularization . . . . . . . . . . . . . . . . .
190
D.3.3
PWLL with ℓ2 regularization
190
. . . . . . . . . . . . . . . . . .
D.4 Additional Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . 191 D.4.1
Detailed Experimental Setting . . . . . . . . . . . . . . . . . .
191
D.4.2
Additional results . . . . . . . . . . . . . . . . . . . . . . . . .
193
D.4.3
Results Averaging Different Seeds . . . . . . . . . . . . . . . .
194
D.4.4
Ablation - Sensitivity to Reward Noise . . . . . . . . . . . . .
196
D.4.5
Ablation Study on Hyperparameters and Log Transform . . .
197
D.4.6
Ablation: PWLL in Smaller Action Spaces . . . . . . . . . . .
198
D.4.7
Ablation - Sensitivity to the number of clusters . . . . . . . .
198
D.4.8
Ablation - Different Logging Supports
. . . . . . . . . . . . .
199
D.4.9
Pessimism Does not Solve Optimization Problems . . . . . . .
200
Notation Clarification: Value vs. Risk Formulation Important: Throughout the main chapter, we present our results using the value formulation, where the goal is to maximize the expected reward V (π) = EX∼ν,A∼π(·|X) [r(X, A)] 202
with rewards R ∈ [0, 1]. In contrast, the appendix uses the equivalent risk (or cost) formulation, where the goal is to minimize the expected cost L(π) = EX∼ν,A∼π(·|X) [c(X, A)] with costs C ∈ [−1, 0]. These two formulations are related by a simple sign change: L(π) = −V (π) and C = −R , where r(x, a) = −c(x, a) for any (x, a) ∈ X × A. Consequently: • Maximizing the value V (π) is equivalent to minimizing the risk L(π). • Upper bounds on |V (π)− V̂ (π)| translate directly to upper bounds on |L(π)− L̂(π)|. • All theoretical guarantees derived in the appendix using the risk formulation apply equivalently to the value formulation presented in the corresponding chapter. We adopt the risk formulation in the appendix as it aligns with the standard convention in statistical learning theory, where one typically minimizes a loss or risk function. The reader should keep this equivalence in mind when relating the appendix results to the main paper.
E.1
Bias and Variance Trade-Off
Notation reminder: This section uses the risk formulation with costs C ∈ [−1, 0], which corresponds to the value formulation with rewards R = −C ∈ [0, 1] in the main paper. The estimator L̂αn (π) here corresponds to −V̂ α (π) in the main paper. In this section, we provide additional results on how α controls the bias and variance of L̂αn (·).
E.1.1
Bias and Variance of IPS-α
The following proposition states the bias-variance trade-off for L̂αn (·). Proposition 14 (Bias and variance of IPS-α). Let α ∈ [0, 1], the following holds for any evaluation policy π ∈ Π that is absolutely continuous with respect to π0 [︁ ]︁ |B(L̂αn (π))| ≤ EX∼ν,A∼π(·|X) 1 − π0 (A|X)1−α , [︂ ]︂ 1 [︁ π(A|X) ]︁ . V L̂αn (π) ≤ EX∼ν,A∼π(·|X) n π0 (A|X)2α−1 203
Proof. We first bound the bias as [︂ ]︂ B(L̂αn (π)) = E L̂αn (π) − L(π) , ]︃ [︃ n π(Ai |Xi ) 1 ∑︂ − L(π) , EXi ∼ν,Ai ∼π0 (·|Xi ),Ci ∼p(·|Xi ,Ai ) Ci = n i=1 π0 (Ai |Xi )α [︃ ]︃ π(A|X) (i) = E(X,A,C)∼µπ0 C − L(π) , π0 (A|X)α [︄ ]︄ [︄ ]︄ ∑︂ ∑︂ π(a|X) = EX∼ν c(X, a) − EX∼ν c(X, a)π(a|X) , α−1 π 0 (a|X) a∈A a∈A [︄ ]︄ ∑︂ = EX∼ν c(X, a)π(a|X)(π0 (a|X)1−α − 1) , a∈A
[︁ ]︁ = EX∼ν,A∼π(·|X) c(X, A)(π0 (A|X)1−α − 1) , where (i) follows from the i.i.d. assumption. Since π0 (A|X)1−α ≤ 1 for any x ∈ X and a ∈ A, we have that [︁ ]︁ |B(L̂αn (π))| ≤ EX∼ν,A∼π(·|X) |c(X, A)||π0 (A|X)1−α − 1| , [︁ ]︁ ≤ EX∼ν,A∼π(·|X) 1 − π0 (A|X)1−α . The variance is bounded as n [︂ [︂ ]︂ 1 ∑︂ π(Ai |Xi ) ]︂ , VXi ∼ν,Ai ∼π0 (·|Xi ),Ci ∼p(·|Xi ,Ai ) Ci V L̂αn (π) = 2 n i=1 π0 (Ai |Xi )α [︂ π(A|X) ]︂ 1 = V(X,A,C)∼µπ0 C , n π0 (A|X)α [︂ 1 π(A|X)2 ]︂ ≤ E(X,A,C)∼µπ0 C 2 , n π0 (A|X)2α [︂ π(A|X)2 ]︂ 1 , ≤ EX∼ν,A∼π0 (·|X) n π0 (A|X)2α [︂ ∑︂ π(a|X)2 ]︂ 1 = EX∼ν , n π0 (a|X)2α−1 a∈A [︂ π(A|X) ]︂ 1 = EX∼ν,A∼π(·|X) . n π0 (A|X)2α−1
E.2
Proofs for Off-Policy Learning
In this section, we provide the complete proofs for our OPL results in Section 8.3. We start with proving Theorem 4 in Section E.2.1. We then state the extension of Theorem 4 204
along with its proof in Section E.2.2. After that, in Section E.2.3, we provide the proof for Proposition 6. Finally, in Section E.2.4, we discuss in detail and prove our claims regarding the number of samples needed so that the performance of the learned policy is close to that of the optimal policy. Notation: This section uses the risk formulation with L(π) = −V (π) and L̂αn (π) = −V̂ α (π). All bounds translate directly to the value formulation in the main paper. Recall that we assume the costs to be deterministic for simplicity: Ci = c(Xi , Ai ).
E.2.1
Proof of Theorem 4
In this section, we prove Theorem 4.
Proof. First, we decompose the difference L(πQ ) − L̂αn (πQ ) as
L(πQ ) − L̂αn (πQ ) = L(πQ ) − ⏞
n n n 1 ∑︂ 1 ∑︂ α 1 ∑︂ L(πQ |Xi ) + L(πQ |Xi ) − L (πQ |Xi ) n i=1 n i=1 n i=1 ⏟⏟ ⏞ ⏞ ⏟⏟ ⏞ I1
I2
+
n ∑︂
1 Lα (πQ |Xi ) − L̂αn (πQ ) , n i=1 ⏟⏟ ⏞ ⏞ I3
where L(πQ ) = EX∼ν ,A∼πQ (·|X) [c(X, A)] , L(πQ |Xi ) = EA∼πQ (·|Xi ) [c(Xi , A)] , [︃ ]︃ πQ (A|Xi ) α L (πQ |Xi ) = EA∼π0 (·|Xi ) c(Xi , A) , π0 (A|Xi )α n 1 ∑︂ π(Ai |Xi ) α L̂n (π) = Ci . n i=1 π0 (Ai |Xi )α Our goal is to bound |L(πQ ) − L̂αn (πQ )| and thus we need to bound |I1 | + |I2 | + |I3 |. We start with |I1 |, Alquier (2021, Theorem 3.3) yields that following inequality holds with probability at least 1 − δ/2 for any distribution Q on H √︄ |I1 | ≤
√
DKL (Q∥P) + log 4 δ n . 2n 205
(E.1)
Moreover, |I2 | can be bounded by decomposing it as ⃓ n ]︃⃓⃓ [︃ n ⃓ 1 ∑︂ ∑︂ 1 πQ (A|Xi ) ⃓ ⃓ c(Xi , A) ⃓ |I2 | = ⃓ EA∼πQ (·|Xi ) [c(Xi , A)] − EA∼π0 (·|Xi ) α ⃓n ⃓ n i=1 π0 (A|Xi ) ⃓ ⃓ i=1 n ∑︂ ⃓ ⃓ 1 ∑︂ πQ (a|Xi ) ⃓ ⃓ c(Xi , a)⃓ =⃓ πQ (a|Xi )c(Xi , a) − π0 (a|Xi ) α ⃓ ⃓n π0 (a|Xi ) ⃓ i=1 a∈A ⃓ n ∑︂ (︂ ⃓ 1 ∑︂ ⃓ πQ (a|Xi ) )︂ ⃓ ⃓ =⃓ πQ (a|Xi ) − α−1 c(Xi , a)⃓ ⃓n ⃓ π0 (a|Xi ) ⃓ ⃓ i=1 a∈A n ∑︂ (︂ ⃓ ⃓ 1 ∑︂ )︂ ⃓ ⃓ 1−α 1 − π0 (a|Xi ) πQ (a|Xi )c(Xi , a)⃓ , =⃓ ⃓ ⃓n ≤
i=1 a∈A n ∑︂ ∑︂
⃓ ⃓ 1 ⃓1 − π01−α (a|Xi )⃓ πQ (a|Xi ) |c(Xi , a)| . n i=1 a∈A
But 1 − π01−α (a|x) ≥ 0 and |c(x, a)| ≤ 1 for any a ∈ A and x ∈ X . Thus |I2 | ≤
n [︁ ]︁ 1 ∑︂ EA∼πQ (·|Xi ) 1 − π01−α (A|Xi ) . n i=1
(E.2)
Finally, we need to bound the main term |I3 |. To achieve this, we borrow the following technical lemma from Haddouche and Guedj (2022). It is slightly different from the one in Haddouche and Guedj (2022); their result holds for any n ≥ 1 while we state a simpler version where n is fixed in advance. Lemma 11. Let Z be an instance space and let Sn = (zi )i∈[n] be an n-sized dataset for some n ≥ 1. Let (Fi )i∈{0}∪[n] be a filtration adapted to Sn . Also, let H be a hypothesis space and (fi (Si , h))i∈[n] be a martingale difference sequence for any h ∈ H, that is for any i ∈ [n], and ∑︁ h ∈ H , we have that E [fi (Si , h) |Fi−1 ] = 0. Moreover, for any h ∈ H, let Mn (h) = ni=1 fi (Si , h). Then for any fixed prior, P, on H, any λ > 0, the following holds with probability 1 − δ over the sample Sn , simultaneously for any Q, on H DKL (Q∥P) + log(2/δ) λ + (Eh∼Q [⟨M ⟩n (h) + [M ]n (h)]) , λ 2 [︁ ]︁ ∑︁ ∑︁ where ⟨M ⟩n (h) = ni=1 E fi (Si , h)2 |Fi−1 and [M ]n (h) = ni=1 fi (Si , h)2 . |Eh∼Q [Mn (h)]| ≤
To apply Lemma 11, we need to construct an adequate martingale difference sequence (fi (Si , h))i∈[n] for h ∈ H that allows us to retrieve |I3 |. To achieve this, we define Sn = (Ai )i∈[n] as the set of n taken actions. Also, we let (Fi )i∈{0}∪[n] be a filtration adapted to Sn . For h ∈ H, we define fi (Si , h) as [︃ ]︃ I{h(Xi )=A} I{h(Xi )=Ai } c(Xi , A) − c(Xi , Ai ) . fi (Si , h) = fi (Ai , h) = EA∼π0 (·|Xi ) α π0 (A|Xi ) π0 (Ai |Xi )α We stress that fi (Si , h) only depends on the last action in Si , Ai , and the predictor h. For this reason, we denote it by fi (Ai , h). The function fi is indexed by i since it depends 206
on the fixed i-th context, Xi . The context Xi is fixed and thus randomness only comes from Ai ∼ π0 (·|Xi ). It follows that the expectations are under Ai ∼ π0 (·|Xi ). First, we have that E [fi (Ai , h) |Fi−1 ] = 0 for any i ∈ [n] , h ∈ H. This follows from ⃓ [︂ ]︂ ⃓ E [fi (Ai , h) |Fi−1 ] = EAi ∼π0 (·|Xi ) fi (Ai , h) ⃓A1 , . . . , Ai−1 , ]︃ ]︃ [︃ [︃ ⃓ I{h(Xi )=Ai } I{h(Xi )=A} ⃓ c(Xi , A) − c(Xi , Ai )⃓A1 , . . . , Ai−1 , = EAi ∼π0 (·|Xi ) EA∼π0 (·|Xi ) π0 (A|Xi )α π0 (Ai |Xi )α [︃ ]︃ [︃ ]︃ ⃓ I{h(Xi )=A} I{h(Xi )=Ai } (i) ⃓ = EA∼π0 (·|Xi ) c(Xi , A) − EAi ∼π0 (·|Xi ) c(Xi , Ai )⃓A1 , . . . , Ai−1 . π0 (A|Xi )α π0 (Ai |Xi )α [︂ I ]︂ i )=A} In (i) we use the fact that given Xi , EA∼π0 (·|Xi ) π{h(X c(X , A) is deterministic. Now i α 0 (A|Xi ) Ai does not depend on A1 , . . . , Ai−1 since logged data is i.i.d. Hence [︃ EAi ∼π0 (·|Xi )
]︃ [︃ ]︃ ⃓ I{h(Xi )=Ai } I{h(Xi )=Ai } ⃓ c(Xi , Ai )⃓A1 , . . . , Ai−1 = EAi ∼π0 (·|Xi ) c(Xi , Ai ) , π0 (Ai |Xi )α π0 (Ai |Xi )α ]︃ [︃ I{h(Xi )=A} c(Xi , A) . = EA∼π0 (·|Xi ) π0 (A|Xi )α
It follows that E [fi (Ai , h) |Fi−1 ] ]︃ ]︃ [︃ [︃ ⃓ I{h(Xi )=Ai } I{h(Xi )=A} ⃓ c(Xi , A) − EAi ∼π0 (·|Xi ) c(Xi , Ai )⃓A1 , . . . , Ai−1 , = EA∼π0 (·|Xi ) π0 (A|Xi )α π0 (Ai |Xi )α ]︃ ]︃ [︃ [︃ I{h(Xi )=A} I{h(Xi )=A} = EA∼π0 (·|Xi ) c(Xi , A) − EA∼π0 (·|Xi ) c(Xi , A) , π0 (A|Xi )α π0 (A|Xi )α = 0. Therefore, for any h ∈ H, (fi (Ai , h))i∈[n] is a martingale difference sequence. Hence we apply Lemma 11 and obtain that the following inequality holds with probability at least 1 − δ/2 for any Q on H |Eh∼Q [Mn (h)]| ≤
DKL (Q∥P) + log(4/δ) λ + (Eh∼Q [⟨M ⟩n (h) + [M ]n (h)]) , λ 2
where Mn (h) = ⟨M ⟩n (h) =
n ∑︂ i=1 n ∑︂
fi (Ai , h) , [︁ ]︁ E fi (Ai , h)2 |Fi−1 ,
i=1
[M ]n (h) =
n ∑︂
fi (Ai , h)2
i=1
207
(E.3)
Now these terms can be decomposed as Eh∼Q [Mn (h)] =
n ∑︂
Eh∼Q [fi (Ai , h)] ,
i=1 n ∑︂
[︃ [︃ ]︃ ]︃ I{h(Xi )=A} I{h(Xi )=Ai } = Eh∼Q EA∼π0 (·|Xi ) c(Xi , A) − c(Xi , Ai ) , π0 (A|Xi )α π0 (Ai |Xi )α i=1 [︃ [︃ ]︃]︃ [︃ ]︃ n I{h(Xi )=A} I{h(Xi )=Ai } (i) ∑︂ = Eh∼Q EA∼π0 (·|Xi ) c(Xi , A) − Eh∼Q c(Xi , Ai ) , α α π (A|X ) π (A |X ) 0 i 0 i i i=1 ]︄ [︄ [︁ ]︁ [︁ ]︁ n Eh∼Q I{h(Xi )=Ai } Eh∼Q I{h(Xi )=A} (ii) ∑︂ = c(Xi , A) − c(Xi , Ai ) , EA∼π0 (·|Xi ) α α π (A|X ) π (A |X ) 0 i 0 i i i=1 [︃ ]︃ ∑︂ n n πQ (A|Xi ) πQ (Ai |Xi ) (iii) ∑︂ = EA∼π0 (·|Xi ) c(Xi , A) − c(Xi , Ai ) , α π0 (A|Xi ) π (Ai |Xi )α i=1 i=1 0 where we use the linearity of the expectation in both (i) and (ii). In (iii), we use our definition of policies in (8.10). Therefore, we have that Eh∼Q [Mn (h)] =
n ∑︂
[︃ EA∼π0 (·|Xi )
i=1
]︃ ∑︂ n πQ (Ai |Xi ) πQ (A|Xi ) c(X , A) − c(Xi , Ai ) , i α π0 (A|Xi )α π (A |X ) 0 i i i=1
n (i) ∑︂ α = L (πQ |Xi ) − nL̂αn (πQ ) , i=1
(E.4)
= nI3 , where we used the fact that Ci = c(Xi , Ai ) for any i ∈ [n] in (i). Now we focus on the terms ⟨M ⟩n (h) and [M ]n (h). First, we have that
]︃ )︂2 I{h(Xi )=Ai } I{h(Xi )=A} fi (Ai , h) = EA∼π0 (·|Xi ) c(X , A) − c(X , A ) , (E.5) i i i π0 (A|Xi )α π0 (Ai |Xi )α [︃ ]︃2 (︂ )︂2 I{h(Xi )=A} I{h(Xi )=Ai } = EA∼π0 (·|Xi ) c(Xi , A) + c(Xi , Ai ) π0 (A|Xi )α π0 (Ai |Xi )α ]︃ [︃ I{h(Xi )=Ai } I{h(Xi )=A} − 2EA∼π0 (·|Xi ) c(Xi , A) c(Xi , Ai ) , α π0 (A|Xi ) π0 (Ai |Xi )α [︃ ]︃2 I{h(Xi )=A} I{h(Xi )=Ai } = EA∼π0 (·|Xi ) c(X , A) + c(Xi , Ai )2 i π0 (A|Xi )α π0 (Ai |Xi )2α [︃ ]︃ I{h(Xi )=A} I{h(Xi )=Ai } − 2EA∼π0 (·|Xi ) c(X , A) c(Xi , Ai ) . i π0 (A|Xi )α π0 (Ai |Xi )α 2
(︂
[︃
Moreover, fi (Ai , h)2 does not depend on A1 , . . . , Ai−1 . Thus, [︁ ]︁ [︁ ]︁ E fi (Ai , h)2 |Fi−1 = EAi ∼π0 (·|Xi ) fi (Ai , h)2 |Fi−1 , [︁ ]︁ [︁ ]︁ = EAi ∼π0 (·|Xi ) fi (Ai , h)2 = EA∼π0 (·|Xi ) fi (A, h)2 . 208
[︁ ]︁ Computing EA∼π0 (·|Xi ) fi (A, h)2 using the decomposition in (E.5) yields ]︁ [︁ ]︁ [︁ E fi (Ai , h)2 |Fi−1 = EA∼π0 (·|Xi ) fi (A, h)2 , [︃ ]︃2 ]︃ [︃ I{h(Xi )=A} I{h(Xi )=A} 2 = −EA∼π0 (·|Xi ) c(Xi , A) + EA∼π0 (·|Xi ) c(Xi , A) π0 (A|Xi )α π0 (A|Xi )2α
(E.6)
Combining (E.5) and (E.6) leads to ]︃ I{h(Xi )=Ai } I{h(Xi )=A} 2 c(Xi , A) + c(Xi , Ai )2 2α π0 (A|Xi ) π0 (Ai |Xi )2α [︃ ]︃ I{h(Xi )=A} I{h(Xi )=Ai } − 2EA∼π0 (·|Xi ) c(X , A) c(Xi , Ai ) , i π0 (A|Xi )α π0 (Ai |Xi )α ]︃ [︃ (i) I{h(Xi )=Ai } I{h(Xi )=A} 2 ≤ EA∼π0 (·|Xi ) c(Xi , A) + c(Xi , Ai )2 . 2α π0 (A|Xi ) π0 (Ai |Xi )2α (E.7) [︂ I ]︂ I {h(Xi )=Ai } i )=A} The inequality in (i) holds because −2EA∼π0 (·|Xi ) π{h(X c(X , A) c(Xi , Ai ) ≤ i α π0 (Ai |Xi )α 0 (A|Xi ) 0. Therefore, we have that ]︃ [︃ n ∑︂ I{h(Xi )=Ai } I{h(Xi )=A} 2 c(Xi , A) + c(Xi , Ai )2 . ⟨M ⟩n (h) + [M ]n (h) ≤ EA∼π0 (·|Xi ) 2α 2α π (A|X ) π (A |X ) 0 i 0 i i i=1 [︁ ]︁ E fi (Ai , h)2 |Fi−1 + fi (Ai , h)2 = EA∼π0 (·|Xi )
[︃
Finally, by using the linearity of the expectation and the definition of policies in (8.10), we get that Eh∼Q [⟨M ⟩n (h) + [M ]n (h)] (E.8) [︄ ]︄ [︁ ]︁ [︁ ]︁ n ∑︂ Eh∼Q I{h(Xi )=A} Eh∼Q I{h(Xi )=Ai } 2 ≤ EA∼π0 (·|Xi ) c(X , A) c(Xi , Ai )2 , + i 2α 2α π0 (A|Xi ) π0 (Ai |Xi ) i=1 ]︃ [︃ n ∑︂ πQ (Ai |Xi ) πQ (A|Xi ) 2 c(X , A) c(Xi , Ai )2 . (E.9) + = EA∼π0 (·|Xi ) i 2α 2α π (A|X ) π (A |X ) 0 i 0 i i i=1 Combining (E.3) and (E.8) yields n|I3 | = |
n ∑︂
Lα (πQ |Xi ) − nL̂αn (πQ )|
i=1
[︃ ]︃ n DKL (Q∥P) + log(4/δ) λ ∑︂ πQ (A|Xi ) πQ (Ai |Xi ) 2 ≤ + EA∼π0 (·|Xi ) c(Xi , A) + c(Xi , Ai )2 . 2α λ 2 i=1 π0 (A|Xi ) π0 (Ai |Xi )2α (E.10)
This means that the following inequality holds with probability at least 1 − δ/2 for any distribution Q on H ]︃ [︃ n DKL (Q∥P) + log(4/δ) λ ∑︂ πQ (A|Xi ) 2 |I3 | ≤ + EA∼π0 (·|Xi ) c(Xi , A) nλ 2n i=1 π0 (A|Xi )2α n
λ ∑︂ πQ (Ai |Xi ) + c(Xi , Ai )2 . 2n i=1 π0 (Ai |Xi )2α 209
(E.11)
However we know that c(x, a)2 ≤ 1 for any x ∈ X and a ∈ A and that c(Xi , Ai ) = Ci for any i ∈ [n]. Thus the following inequality holds with probability at least 1 − δ/2 for any distribution Q on H ]︃ [︃ n n λ ∑︂ λ ∑︂ πQ (Ai |Xi ) 2 πQ (A|Xi ) DKL (Q∥P) + log(4/δ) + + EA∼π0 (·|Xi ) C . |I3 | ≤ nλ 2n i=1 π0 (A|Xi )2α 2n i=1 π0 (Ai |Xi )2α i (E.12)
The union bound of (E.1) and (E.12) combined with the deterministic result in (E.2) yields that the following inequality holds with probability at least 1 − δ for any distribution Q on H √︄ √ n 4 n [︁ ]︁ D (Q∥P) + log 1 ∑︂ KL δ |L(πQ ) − L̂αn (πQ )| ≤ + EA∼πQ (·|Xi ) 1 − π01−α (A|Xi ) 2n n i=1 ]︃ [︃ n n λ ∑︂ λ ∑︂ πQ (Ai |Xi ) 2 DKL (Q∥P) + log(4/δ) πQ (A|Xi ) + + + EA∼π0 (·|Xi ) C . nλ 2n i=1 π0 (A|Xi )2α 2n i=1 π0 (Ai |Xi )2α i (E.13)
E.2.2
Extensions of Theorem 4
Proposition 15 (Extension of Theorem 4 to hold simultaneously for any λ ∈ (0, 1)). Let n ≥ 1, δ ∈ [0, 1], α ∈ [0, 1], and let P be a fixed prior on H, then with probability at least 1 − δ over draws Dn ∼ µnπ0 , the following holds simultaneously for any posterior Q on H, and for any λ ∈ (0, 1) that √︃ kl′ 1 (πQ , λ) kl′ 2 (πQ , λ) λ α + Bnα (πQ ) + + Varαn (πQ ) . |L(πQ ) − L̂n (πQ )| ≤ 2n nλ 2 where
√ 8 n kl 1 (πQ , λ) = DKL (Q∥P) + log , δλ (︁ 8 )︁ kl′ 2 (πQ , λ) = 2 DKL (Q∥P) + log , δλ n [︁ ]︁ 1 ∑︂ α Bn (πQ ) = 1 − EA∼πQ (·|Xi ) π01−α (A|Xi ) , n i=1 ]︃ [︃ n πQ (Ai |Xi ) 2 1 ∑︂ πQ (A|Xi ) α Varn (πQ ) = EA∼π0 (·|Xi ) + C . 2α n i=1 π0 (A|Xi ) π0 (Ai |Xi )2α i ′
Proof. Let δ ∈ (0, 1). For any i ≥ 1, we define λi = 2−i and let δi = δλi . Then Theorem 4 yields that for any i ≥ 1, the following inequality holds with probability at least 1 − δi for any Q on H √︄ √ 4 n D (Q∥P) + log DKL (Q∥P) + log δ4i KL λi δi |L(πQ ) − L̂αn (πQ )| ≤ + Bnα (πQ ) + + Varαn (πQ ) . 2n nλi 2 210
∑︁ ∑︁∞ Now notice that ∞ i=1 λi = 1, and hence i=1 δi = δ. Therefore, the union bound of the above inequalities over i ≥ 1 yields that with probability at least 1 − δ, the following inequality holds with probability at least 1 − δ for any Q on H and for any i ≥ 1 √︄ |L(πQ ) − L̂αn (πQ )| ≤
√
DKL (Q∥P) + log 4 δin 2n
+ Bnα (πQ ) +
DKL (Q∥P) + log δ4i nλi
+
λi Varαn (πQ ) . 2 (E.14)
Let ⌈·⌉ denote the ceiling function, then we have that for any λ ∈ (0, 1), there exists log λ j = ⌈ −log ⌉ ≥ 1 such that λ/2 ≤ λj ≤ λ. Since (E.14) holds for any i ≥ 1, it holds 2 in particular for j. In addition to this, we have that λ1j ≤ λ2 , that λj ≤ λ and that 1 2 . This yields that the following inequality holds with probability at least = λ1j δ ≤ δλ δj 1 − δ for any Q on H and for any λ ∈ (0, 1) √︄ |L(πQ ) − L̂αn (πQ )| ≤
√
8 DKL (Q∥P) + log 8δλn DKL (Q∥P) + log δλ λ α + Bn (πQ ) + 2 + Varαn (πQ ) . 2n nλ 2 (E.15)
D
(Q∥P)+log 8
δλ appears since we used that λ1j ≤ λ2 . Similarly, the The additional 2 in 2 KL nλ 2 additional λ2 in the logarithmic terms is due to the fact that δ1j ≤ δλ . Finally, setting
√ 8 n kl 1 (πQ , λ) = DKL (Q∥P) + log , δλ (︁ 8 )︁ kl′ 2 (πQ , λ) = 2 DKL (Q∥P) + log , δλ ′
concludes the proof.
Next, we provide a similar proof to extend Theorem 4 to any α ∈ (0, 1]. While we only provide a one-sided inequality, the same covering technique can be used to obtain the other side of the inequality. Proposition 16 (One-sided extension of Theorem 4 to hold simultaneously for any α ∈ (0, 1) ∪ {1} ). Let n ≥ 1, δ ∈ [0, 1], λ > 0, and let P be a fixed prior on H, then with probability at least 1 − δ over draws Dn ∼ µnπ0 , the following holds simultaneously for any posterior Q on H, and for any α ∈ (0, 1] that √︃ L(πQ ) ≤ L̂αn (πQ ) +
kl′′ 1 (πQ , α) kl′′ 2 (πQ , α) λ α + Bn (πQ ) + + Var2α n (πQ ) . 2n nλ 2 211
where √ 8 n kl 1 (πQ , α) = DKL (Q∥P) + log , δα 8 kl′′ 2 (πQ , α) = DKL (Q∥P) + log , δα n ]︁ [︁ 1 ∑︂ Bnα (πQ ) = 1 − EA∼πQ (·|Xi ) π01−α (A|Xi ) , n i=1 [︃ ]︃ n πQ (A|Xi ) πQ (Ai |Xi ) 2 1 ∑︂ α EA∼π0 (·|Xi ) + C . Varn (πQ ) = 2α n i=1 π0 (A|Xi ) π0 (Ai |Xi )2α i ′′
Proof. Let δ ∈ (0, 1). For any i ≥ 0, we define αi = 2−i and let δi = δαi /2. Then Theorem 4 yields that for any i ≥ 0, the following inequality holds with probability at least 1 − δi for any Q on H √︄ √ 4 n D (Q∥P) + log DKL (Q∥P) + log δ4i KL λ δi + Bnαi (πQ ) + + Varnαi (πQ ) . |L(πQ ) − L̂nαi (πQ )| ≤ 2n nλ 2 ∑︁∞ ∑︁ Now notice that ∞ i=0 δi = δ. Therei=0 αi = 2, and hence by definition of δi , we have fore, the union bound of the above inequalities over i ≥ 0 yields that with probability at least 1 − δ, the following inequality holds with probability at least 1 − δ for any Q on H and for any i ≥ 0 √︄ √ 4 n D (Q∥P) + log DKL (Q∥P) + log δ4i KL λ δi αi αi |L(πQ ) − L̂n (πQ )| ≤ + Bn (πQ ) + + Varnαi (πQ ) . 2n nλ 2 (E.16) Let ⌊·⌋ denote the floor function, then we have that for any α ∈ (0, 1], there exists j = log α ⌊ −log ⌋ ≥ 0 such that α ≤ αj ≤ 2α. Since (E.16) holds for any i ≥ 0, it holds in particular 2 for j. In addition, we have that Bnα (πQ ) and L̂αn (πQ ) are decreasing in α while Varαn (πQ ) α α is increasing in α. Therefore, we have that L̂nj (πQ ) ≤ L̂αn (πQ ) , Bn j (πQ ) ≤ Bnα (πQ ) , and 1 2 Varαnj (πQ ) ≤ Var2α n (πQ ). Moreover, we have that δj ≤ δα . This yields that the following inequality holds with probability at least 1 − δ for any Q on H and for any α ∈ (0, 1] √︄ √ 8 n 8 D (Q∥P) + log DKL (Q∥P) + log δα λ KL δα L(πQ ) ≤ L̂αn (πQ ) + + Bnα (πQ ) + + Var2α n (πQ ) . 2n nλ 2 (E.17) Finally, setting √ 8 n kl 1 (πQ , α) = DKL (Q∥P) + log , δα 8 kl′′ 2 (πQ , α) = DKL (Q∥P) + log , δα ′′
concludes the proof. 212
E.2.3
Proof of Proposition 6
Haddouche and Guedj (2022, Theorem 7) provides an application of Lemma 11 to the general PAC-Bayes learning problems in Section 8.3.1. We cannot apply their theorem directly to get Proposition 6 for two reasons. They assume that the loss function is nonnegative and they derive a one-sided generalization bound. In our case, the loss function is negative and we want to derive a two-sided generalization bound. Fortunately, we show with a slight modification of their proof that the result can be extended to two-sided inequalities with negative losses. In fact, the only requirement is that the sign of loss is fixed. We show next how this is achieved. Proof. First, note that Lemma 11 does not make any assumption on the sign of the martingale difference sequence (fi (Si , h))i∈[n] nor on the sign of the terms that decompose it. Now similarly to the proof in Section E.2.1, we define Sn = (Xi , Ai )i∈[n] as the set of n observed contexts and taken actions. Also, we let (Fi )i∈{0}∪[n] be a filtration adapted to Sn . For h ∈ H, we define fi (Si , h) as fi (Si , h) = fi (Xi , Ai , h) = f (Xi , Ai , h) , ]︃ [︃ I{h(Xi )=Ai } I{h(X)=A} c(X, A) − c(Xi , Ai ) . = EX∼ν,A∼π0 (·|X) π0 (A|X)α π0 (Ai |Xi )α Here fi (Si , h) only depends on the last samples Xi , Ai and the predictor h. For this reason, we denote it by fi (Xi , Ai , h). Also, the function fi does not depend on i and this is why we simplify the notation as fi (Xi , Ai , h) = f (Xi , Ai , h). Moreover, the randomness in f (Xi , Ai , h) is only due Xi ∼ ν and Ai ∼ π0 (·|Xi ); all other terms are deterministic. Thus the expectations are under Xi ∼ ν, Ai ∼ π0 (·|Xi ). Now similarly to the proof in Section E.2.1, we have that E [f (Xi , Ai , h) |Fi−1 ] = 0 for any i ∈ [n] , h ∈ H. Therefore, (f (Xi , Ai , h))i∈[n] is a martingale difference sequence for any h ∈ H. Thus we apply Lemma 11 and get that that with probability at least 1 − δ, the following holds simultaneously for any distribution Q on H DKL (Q∥P) + log(2/δ) λ + (Eh∼Q [⟨M ⟩n (h) + [M ]n (h)]) , (E.18) |Eh∼Q [Mn (h)]| ≤ λ 2 where n ∑︂ Mn (h) = f (Xi , Ai , h) , i=1
⟨M ⟩n (h) = [M ]n (h) =
n ∑︂ i=1 n ∑︂
[︁ ]︁ E f (Xi , Ai , h)2 |Fi−1 ,
f (Xi , Ai , h)2 .
i=1
Now we compute Eh∼Q [Mn (h)] as ]︃ [︃ n ∑︂ πQ (Ai |Xi ) πQ (A|X) Eh∼Q [Mn (h)] = EX∼ν,A∼π0 (·|X) c(X, A) − c(Xi , Ai ) , α α π (A|X) π (A |X ) 0 0 i i i=1 [︃ ]︃ ∑︂ n πQ (Ai |Xi ) πQ (A|X) = nEX∼ν,A∼π0 (·|X) c(X, A) − c(Xi , Ai ) , (E.19) α π0 (A|X) π (Ai |Xi )α i=1 0 213
where we used the linearity of the expectation Eh∼Q [·] and the definition of policies in (8.10). Moreover, similarly to the proof in Section E.2.1, we have that n ∑︂ [︁ ]︁ ⟨M ⟩n (h) + [M ]n (h) = E f (Xi , Ai , h)2 |Fi−1 + f (Xi , Ai , h)2 i=1 n ∑︂
[︃
]︃ I{h(X)=A} I{h(Xi )=Ai } 2 = EX∼ν,A∼π0 (·|X) c(X, A) + c(Xi , Ai )2 2α 2α π (A|X) π (A |X ) 0 0 i i i=1 ]︃ [︃ I{h(Xi )=Ai } I{h(X)=A} c(X, A) c(Xi , Ai ) , − 2EX∼ν,A∼π0 (·|X) π0 (A|X)α π0 (Ai |Xi )α [︃ ]︃ ∑︂ n (i) I{h(X)=A} I{h(Xi )=Ai } 2 ≤ nEX∼ν,A∼π0 (·|X) c(X, A) + c(Xi , Ai )2 , 2α 2α π0 (A|X) π (A |X ) i i i=1 0
(E.20) ]︂ I I {h(Xi )=Ai } where (i) holds since −2EX∼ν,A∼π0 (·|X) π{h(X)=A} c(Xi , Ai ) ≤ 0 for any α c(X, A) π0 (Ai |Xi )α 0 (A|X) i ∈ [n]. This is where the non-negative loss assumption is not needed. Our loss Lα (h, x, a, c) = I{h(X)=A} c is negative since c ∈ [−1, 0]. However, we only need the product between the π0 (a|x)α loss and its expectation to be non-negative. This holds in particular when the loss has a fixed sign. In that case, the expectation of the loss and the loss itself will have the same sign and thus their product will be non-negative. In our case, the loss has a fixed negative sign and this is all we needed. Now notice that ]︃ [︃ πQ (A|X) c(X, A) = nRα (πQ ) , nEX∼ν,A∼π0 (·|X) π0 (A|X)α n ∑︂ πQ (Ai |Xi ) c(Xi , Ai ) = nL̂αn (πQ ) , α π (A |X ) i i i=1 0 [︂
where we used that c(Xi , Ai ) = Ci for any i ∈ [n] in the second equality. Using these two equalities and plugging (E.19) and (E.20) in (E.18) yields that with probability at least 1 − δ, the following holds simultaneously for any distribution Q on H [︃ ]︃ ⃓ ⃓ D (Q∥P) + log(2/δ) λ (︂ πQ (A|X) ⃓ α ⃓ KL α 2 n ⃓L (πQ ) − L̂n (πQ )⃓ ≤ + nEX∼ν,A∼π0 (·|X) c(X, A) λ 2 π0 (A|X)2α n )︂ ∑︂ πQ (Ai |Xi ) 2 + c(X , A ) . i i 2α π (A |X ) 0 i i i=1 (E.21)
Again we used the linearity of the expectation Eh∼Q [·] and the definition of policies in (8.10). Finally, we have that c(Xi , Ai ) = Ci for any i ∈ [n]. Thus with probability at least 1 − δ the following inequality holds for any distribution Q on H [︃ ]︃ ⃓ ⃓ D (Q∥P) + log(2/δ) λ πQ (A|X) ⃓ α ⃓ KL 2 α c(X, A) + EX∼ν,A∼π0 (·|X) ⃓L (πQ ) − L̂n (πQ )⃓ ≤ nλ 2 π0 (A|X)2α n λ ∑︂ πQ (Ai |Xi ) 2 + C . 2n i=1 π0 (Ai |Xi )2α i
(E.22)
This concludes the proof. 214
E.2.4
Sample Complexity
Notation reminder: Minimizing risk L(π) is equivalent to maximizing value V (π) = −L(π). Achieving L(π̂) ≤ L(π∗ ) + ϵ is equivalent to V (π̂) ≥ V (π∗ ) − ϵ. Proposition 17. Let M1 (H) be the set of probability distributions on the hypothesis space H, and let λ > 0, n ≥ 1, δ ∈ [0, 1], α ∈ [0, 1], and let P be a fixed prior on H, then with probability at least 1 − δ over draws Dn ∼ µnπ0 , we have √︃
kl2 (πQ∗ ) kl1 (πQ∗ ) + 2Bnα (πQ∗ ) + 2 + λ Varαn (πQ∗ ) . 2n nλ √︂ kl1 (πQ ) α where πQ̂n is the learned policy with Q̂n = argminQ∈M1 (H) L̂n (πQ ) + + Bnα (πQ ) + 2n L(πQ̂n ) ≤ L(πQ∗ ) + 2
kl2 (πQ ) + λ2 Varαn (πQ ) , Q∗ = argminQ∈M1 (H) L(πQ ), and nλ
√ 4 n 4 kl1 (πQ ) = DKL (Q∥P) + log , kl2 (πQ ) = DKL (Q∥P) + log , δ δ n ∑︂ [︁ ]︁ 1 Bnα (πQ ) = 1 − EA∼πQ (·|Xi ) π01−α (A|Xi ) , n i=1 [︃ ]︃ n πQ (A|Xi ) πQ (Ai |Xi )Ci2 1 ∑︂ α EA∼π0 (·|Xi ) . Varn (πQ ) = + n i=1 π0 (A|Xi )2α π0 (Ai |Xi )2α Proof. First, Theorem 4 holds for any potentially data dependent distribution Q on H. In particular, we have that with probability at least 1 − δ the following inequalities hold simultaneously for Q̂n and Q∗ √︄ kl2 (πQ̂n ) λ kl1 (πQ̂n ) + Bnα (πQ̂n ) + + Varαn (πQ̂n ) , |L(πQ̂n ) − L̂αn (πQ̂n )| ≤ 2n nλ 2 √︃ kl2 (πQ∗ ) λ kl1 (πQ∗ ) |L(πQ∗ ) − L̂αn (πQ∗ )| ≤ + Bnα (πQ∗ ) + + Varαn (πQ∗ ) . 2n nλ 2 Taking only one side of these inequalities yields that with probability at least 1 − δ the following inequalities hold simultaneously for Q̂n and Q∗ √︄ kl1 (πQ̂n ) kl2 (πQ̂n ) λ L(πQ̂n ) ≤ L̂αn (πQ̂n ) + + Bnα (πQ̂n ) + + Varαn (πQ̂n ) , 2n nλ 2 ⏞ ⏟⏟ ⏞ (I)
√︃ L̂αn (πQ∗ ) ≤ L(πQ∗ ) +
kl1 (πQ∗ ) kl2 (πQ∗ ) λ + Bnα (πQ∗ ) + + Varαn (πQ∗ ) . 2n nλ 2
Now using the definition of πQ̂n , we know that √︃ I ≤ L̂αn (πQ∗ ) +
kl1 (πQ∗ ) kl2 (πQ∗ ) λ + Bnα (πQ∗ ) + + Varαn (πQ∗ ) . 2n nλ 2 215
This yields that with probability at least 1 − δ the following inequalities hold simultaneously for Q̂n and Q∗ √︃ kl2 (πQ∗ ) λ kl1 (πQ∗ ) L(πQ̂n ) ≤ L̂αn (πQ∗ ) + + Bnα (πQ∗ ) + + Varαn (πQ∗ ) , 2n nλ 2 √︃ kl1 (πQ∗ ) kl2 (πQ∗ ) λ L̂αn (πQ∗ ) ≤ L(πQ∗ ) + + Bnα (πQ∗ ) + + Varαn (πQ∗ ) . 2n nλ 2 Computing the sum of these two inequalities concludes the proof. {︁ }︁ Corollary 2 (Special case of Proposition 17). Let H = hθ ; θ ∈ RdK of mappings hθ (x) = argmaxa∈A ϕ(x)⊤ θa for any x ∈ X . Let n ≥ 1, δ ∈ [0, 1], α ∈ [0, 1], and let P = N (µ0 , IdK ) be a fixed prior on H, then with probability at least 1 − δ over draws Dn ∼ µnπ0 , we have that √︂ √ ∥µ∗ − µ0 ∥2 + 2 log 4 δ n √ L(πQ̂n ) ≤ L(πQ∗ ) + n ∥µ∗ − µ0 ∥2 + 2 log 4δ K 2α−1 + K 2α √ √ + 2(1 − K α−1 ) + + . n n √︂ kl1 (πQ ) α where πQ̂n is the learned policy with Q̂n = argminQ=N (µ,IdK ) L̂n (πQ )+ +Bnα (πQ )+ 2n kl2 (πQ ) + λ2 Varαn (πQ ) , Q∗ = argminQ=N (µ,IdK ) L(πQ ). nλ
Proof. This result follows from the general Proposition 17 by simply setting P = N (µ0 , IdK ) and Q∗ = N (µ∗ , IdK ). First, since the covariance matrices of both distributions are IdK , their KL divergence is DKL (Q∥P) = ∥µ∗ − µ0 ∥2 /2. Moreover, since the logging policy is uniform then B√nα (πQ ) = (1 − K α−1 ) and Varαn (πQ ) ≤ K 2α−1 + K 2α . Using these quantities, setting λ = 1/ n and applying Proposition 17 yields that with probability at least 1 − δ over draws Dn ∼ µnπ0 , we have that √︂ √ ∥µ∗ − µ0 ∥2 + 2 log 4 δ n √ + 2(1 − K α−1 ) L(πQ̂n ) ≤ L(πQ∗ ) + n ∥µ∗ − µ0 ∥2 + 2 log 4δ K 2α−1 + K 2α √ √ + . + n n This concludes the proof. The above corollary allows us to give insights into the sample complexity of our procedure. That is, the number of samples needed so that the performance of the learned policy πQ̂n is close to that of the optimal one. Let ϵ > 2(1 − K α−1 ) for α ∈ [1 − log 2/ log K, 1]. This condition on α ensures that ϵ ∈ [0, 1] and it is mild as α is often close to 1. Let δ, then the following implication holds √︂ √ ∥µ∗ − µ0 ∥2 + 2 log 4 δ n ∥µ∗ − µ0 ∥2 + 2 log 4δ K 2α−1 + K 2α √ √ √ ϵ≥ + 2(1 − K α−1 ) + + n n n =⇒ P(L(πQ̂n ) ≤ L(πQ∗ ) + ϵ) ≥ 1 − δ . (E.23) 216
√︂ √ √ 4 n 2 ∥µ∗ − µ0 ∥ + 2 log δ ≤ ∥µ∗ −µ0 ∥+ 2 log 4 δ n . Moreover we bound 2α
√︂
First, we use that K 2α−1 + K 2α ≤ 2K . Then the implication in (E.23) becomes ∥µ∗ − µ0 ∥ + ∥µ∗ − µ0 ∥ + 2 log 4δ + √ n≥ ϵ − 2(1 − K α−1 ) 2
√︂ √ 2 log 4 δ n + 2K 2α
=⇒ P(L(πQ̂n ) ≤ L(πQ∗ ) + ϵ) ≥ 1 − δ . (E.24)
We only provide intuition on the sample complexity and aim at having easy-to-interpret terms. Thus we omit the logarithmic terms in (E.24) and assume that ∥µ∗ − µ0 ∥2 ≥ ∥µ∗ −µ0 ∥. This leads to the claim made in Section 8.4.1. Of course, a more √︂precise√sample √ complexity analysis can be made by studying the function h(x) = x − 2 log 4 δ x /(ϵ − 2(1 − K α−1 )) and finding x such that f (x) ≥
E.3
Experiments
E.3.1
Setup
∥µ∗ −µ0 ∥+∥µ∗ −µ0 ∥2 +2 log 4δ +2K 2α . ϵ−2(1−K α−1 )
We consider the standard supervised-to-bandit conversion (Agarwal et al., 2014). Precisely, let Sntr and Sntsts be the training and testing set of a classification dataset, respectively. First, we transform the training set Sntr to a logged bandit data Dn as described in Algorithm 3. The resulting logged data Dn is then used to train our policies. After that, the learned policies are tested on Sntsts as described in Algorithm 4. We consider that the resulting reward in Algorithm 4 is a good proxy for the unknown true reward of the learned policies. This will be our performance metric, the higher the better. In our experiments, we use the following image classification datasets MNIST (LeCun et al., 1998), FashionMNIST (Xiao et al., 2017), EMNIST (Cohen et al., 2017) and CIFAR100 (Krizhevsky et al., 2009). We provide a summary of the statistics of these datasets in Table E.1. Algorithm 3 takes as input a logging policy π0 which we define as π0 (a|x) = ∑︁
exp(η0 · ϕ(x)⊤ µ0,a ) , ⊤ a′ ∈A exp(η0 · ϕ(x) µ0,a′ )
∀(x, a) ∈ X × A .
(E.25)
Here ϕ(x) ∈ Rd is the feature transformation function that outputs a d-dimensional vector, µ0 = (µ0,a )a∈A ∈ RdK are learnable parameters and η0 is an inverse-temperature parameter for the softmax in (E.25). We explain next how these quantities are derived in detail. The feature transformation function ϕ(x) ∈ Rd : for all the datasets, except CIFAR100, x the feature transformation function ϕ(·) is defined as ϕ(x) = ∥x∥ for any x ∈ X . That is, we simply normalize the features x ∈ X by their L2 norm ∥x∥. In contrast, CIFAR100 is a more challenging problem. Thus we use transfer learning to extract features ϕ(x) expressive enough so that a linear softmax model would enjoy a reasonable performance. Precisely, we retrieve the last hidden layer of a ResNet-50 network, pre-trained on the ImageNet dataset, to output 2048-dimensional features. Finally, the obtained features 217
Table E.1: Statistics of the datasets used in our experiments. Data set
Nbr. train samples n
Nbr. test samples nts
Nbr. actions K
Dimension d
MNIST
60000
10000
10
FashionMNIST
60000
10000
10
784 784
EMNIST
112800
18800
47
784
CIFAR100
50000
10000
100
2048
x are normalized as ∥x∥ and this whole process (ResNet-50 + normalization) corresponds to ϕ(·) for CIFAR100.
The parameters µ0 = (µ0,a )a∈A ∈ RdK : we learn the parameters µ0 using 5% of the training set Sntr . Precisely, we use the cross-entropy loss with an L2 regularization of 10−6 to prevent the logging policy π0 from being degenerate. This ensures that the learning policies are absolutely continuous with respect to the logging policy π0 , a condition under which standard IPS is unbiased. In optimization, we use Adam (Kingma and Ba, 2014) with a learning rate of 0.1 for 10 epochs. In all the experiments, we set the prior P = N (η0 µ0 , IdK ) for the Gaussian policies in (8.13) and we set it as P = N (η0 µ0 , IdK ) × G(0, 1)K for the mixed-logit policies in (8.12). Our theory requires that the prior does not depend on data. Given that µ0 is learned on the 5% portion of data, we only train our learning policies on the remaining 95% portion of the data to match our theoretical requirements. The inverse-temperature parameter η0 ∈ R: this controls the performance of the logging policy. A high positive value of η0 leads to a well-performing logging policy, while a negative one leads to a low-performing logging policy. When η0 = 0, π0 is identical to the uniform policy. In our experiments η0 varies between 0 and 1. Algorithm 3 Supervised-to-bandit: creating logged data Input: training classification set Sntr = {(Xi , yi )}ni=1 , logging policy π0 . Output: logged bandit data Dn = (Xi , Ai , Ci )i∈[n] . Initialize Dn = {} for i = 1, . . . , n do Ai ∼ π0 (·|Xi ) Ci = −I{Ai =yi } Dn ← Dn ∪ {(Xi , Ai , Ci )} . Algorithm 4 Supervised-to-bandit: testing policies ts Input: image classification dataset Sntsts = {(Xi , yi )}ni=1 , learned policy π̂n . Output: reward r. for i = 1, . . . , nts do Ai ∼ π̂n (·|Xi ) Ri = I{Ai =yi } ∑︁ ts r = n1ts ni=1 Ri . Now it remains to explain the learning policies πQ and the corresponding closed-form bounds using either our results or those in existing works (London and Sandler, 2019; 218
Sakhi et al., 2022).
E.3.2
Policies
Here we present the two families of policies that we use in our experiments, Gaussian and mixed-logit policies. Mixed-Logit }︁ {︁ Let H = hθ,γ ; θ ∈ RdK , γ ∈ RK be a hypothesis space of mappings hθ,γ (x) = argmaxa∈A ϕ(x)⊤ θa + γa for any x ∈ X . Here ϕ(x) outputs a d-dimensional representation of context x ∈ X . Now assume that for any a ∈ A, γa is a standard Gumbel perturbation, γa ∼ G(0, 1), then we have that exp(ϕ(x)⊤ θa ) sof ∑︁ πθ (a|x) = , ⊤ a′ ∈A exp(ϕ(x) θa′ ) [︁ ]︁ = Eγ∼G(0,1)K I{hθ,γ (x)=a} .
(E.26)
In addition, we randomize θ such as θ ∼ N (µ, σ 2 IdK ) where µ ∈ RdK and σ > 0. It follows that the posterior Q is a multivariate Gaussian N (µ, σ 2 IdK ) over the parameters mixL θ with standard Gumbel perturbations γ ∼ G(0, 1)K . We denote such policies by πµ,σ and they are defined as ]︃ [︃ exp(ϕ(x)⊤ θa ) mixL , πµ,σ (a|x) = Eθ∼N (µ,σ2 IdK ) ∑︁ ⊤ a′ ∈A exp(ϕ(x) θa′ ) = Eθ∼N (µ,σ2 IdK ) [πθsof (a|x)] , [︁ ]︁ = Eθ∼N (µ,σ2 IdK ) ,γ∼G(0,1)K I{hθ,γ (x)=a} .
(E.27)
mixL , we first sample θ ∼ N (µ, σ 2 IdK ) and To sample from the mixed-logit policies πµ,σ K γ ∼ G(0, 1) and then set the sampled action as a ← hθ,γ (x). Now we also need to compute the gradient of the expectation in (E.27). This needs additional care since the distribution under which we take the expectation depends on the parameters µ, σ. Fortunately, the reparameterization trick can be used in this case. Roughly speaking, it allows us to express a gradient of the expectation in (E.27) as an expectation of a gradient. In our case, we use the local reparameterizaton trick (Kingma et al., 2015) which is known for reducing the variance of stochastic gradients. Precisely, we rewrite (E.27) as
]︃ exp(ϕ(x)⊤ µa + σϵa ) ∑︁ . ⊤ a′ ∈A exp(ϕ(x) µa′ + σϵa′ ) [︃ ]︃ exp(ϕ(x)⊤ µa + σϵa ) , = Eϵ∼N (0,IK ) ∑︁ ⊤ a′ ∈A exp(ϕ(x) µa′ + σϵa′ ) [︃
mixL πµ,σ (a|x) = Eϵ∼N (0,∥ϕ(x)∥2 IK )
where we used that ∥ϕ(x)∥2 = 1 since features are normalized. It follows that gradients read [︃ ]︃ exp(ϕ(x)⊤ µa + σϵa ) mixL ∇µ,σ πµ,σ (a|x) = Eϵ∼N (0,IK ) ∇µ,σ ∑︁ . ⊤ a′ ∈A exp(ϕ(x) µa′ + σϵa′ ) 219
Moreover, the propensities are approximated as mixL πµ,σ (a|x) ≈
exp(ϕ(x)⊤ µa + σϵi,a ) 1 ∑︂ ∑︁ , ⊤ S a′ ∈A exp(ϕ(x) µa′ + σϵi,a′ )
ϵi ∼ N (0, IK ) , ∀i ∈ [S] .
(E.28)
i∈[S]
In all our experiments, we set S = 32. Gaussian }︁ {︁ We define the hypothesis space H = hθ ; θ ∈ RdK of mappings hθ (x) = argmaxa∈A ϕ(x)⊤ θa gaus read for any x ∈ X . It follows that the learning policies πQ = πµ,σ [︁ ]︁ gaus πµ,σ (a|x) = Eθ∼N (µ,σ2 IdK ) I{hθ (x)=a} . (E.29) To see why this can be beneficial (Sakhi et al., 2022), let π∗ be the optimal policy. Given x ∈ X , π∗ (·|x) should be deterministic; it chooses the best action for context x with probability 1. That is, there exists µ∗ ∈ RdK such that π∗ = I{hµ∗ (x)=a} . When µ → µ∗ and σ → 0, the Gaussian policy in (E.29) approaches π∗ . In contrast, the mixed-logit policy in (E.27) approaches πµsof . However, πµsof is not deterministic due to the additional ∗ ∗ ⊤ randomness in γ and is equal to π∗ only if ϕ(x) µ∗,a∗ (x) → ∞. This explains the choice of removing the Gumbel noise. First, Sakhi et al. (2022) showed that (E.29) can be written as [︂ ∏︂ (︁ ϕ(x)⊤ (µa − µa′ ) )︁]︂ gaus πµ,σ , (a|x) = Eϵ∼N (0,1) Φ ϵ+ σ∥ϕ(x)∥ a′ ̸=a where Φ is the cumulative distribution function of a standard normal variable. But ∥ϕ(x)∥ = 1 in all our experiments. Thus gaus πµ,σ (a|x) = Eϵ∼N (0,1)
[︂ ∏︂ (︁ ϕ(x)⊤ (µa − µa′ ) )︁]︂ . Φ ϵ+ σ ′ a ̸=a
Then similarly to mixed-logit policies, the gradient reads gaus ∇µ,σ πµ,σ (a|x) = Eϵ∼N (0,1)
[︂ ∏︂ (︁ ϕ(x)⊤ (µa − µa′ ) )︁]︂ ∇µ,σ Φ ϵ+ . σ a′ ̸=a
Moreover, the propensities are approximated as gaus πµ,σ (a|x) ≈
1 ∑︂ ∏︂ (︁ ϕ(x)⊤ (µa − µa′ ) )︁ Φ ϵi + , S σ ′ a ̸=a
ϵi ∼ N (0, 1) , ∀i ∈ [S] .
(E.30)
i∈[S]
In all our experiments, we set S = 32.
E.3.3
Baselines
Here we present all the methods that we use in our experiments. For each method, we state the result that holds for any learning policy π. After that, we derive the corresponding 220
closed-form bounds for Gaussian and mixed-logit policies that we presented previously. All the baselines require computing the KL divergence between the prior P and the posterior Q. Thus before presenting them, we state the following lemma that allows bounding the KL divergence between the prior P and the posterior Q in the cases of mixed-logit or Gaussian policies. Lemma 12 (KL divergence for Gaussian distributions with Gumbel noise). For distributions P = N (µ0 , σ02 IdK ) × G(0, 1)K and Q = N (µ, σ 2 IdK ) × G(0, 1)K , with µ0 , µ ∈ RdK and 0 < σ 2 ≤ σ02 < ∞, ∥µ − µ0 ∥2 dK σ02 DKL (Q∥P) ≤ + log . 2σ02 2 σ2 Moreover, this result holds when the Gumbel noise is removed. That is when P = N (µ0 , σ02 IdK ) and Q = N (µ, σ 2 IdK ). We borrow this lemma from London and Sandler (2019). In particular, Lemma 12 shows that the KL terms for both policies can be bounded by the same quantity. As a result, the corresponding bounds will be the same; the only difference is the space of learning policies on which we optimize. For completeness, however, we write these bounds for both types of policies although they are similar. Since existing approaches are not named, we name them as (Author, Policy) where Author ∈ {Ours, London et al., Sakhi et al. 1, Sakhi et al. 2} and Policy ∈ {Gaussian, Mixed-Logit} . Here Ours, London et al., Sakhi et al. 1 and Sakhi et al. 2 correspond to Theorem 4, London and Sandler (2019, Theorem 1), Sakhi et al. (2022, Proposition 1), Sakhi et al. (2022, Proposition 3), respectively. For example, London and Sandler (2019, Theorem 1) leads to two baselines (London et al., Gaussian) and (London et al., Mixed-Logit). In all our experiments, the learning policies are trained using Adam (Kingma and Ba, 2014) with a learning rate of 0.1 for 20 epochs. Ours, Theorem 4 (Ours, Gaussian) Here we use the Gaussian policies in (E.29). Thus we only replace the term, DKL (Q∥P), with its closed-form bound in Lemma 12. This leads to the following objective. √︄ √ ∥µ−µ0 ∥2 dK (︂ (︁ 2 + log 4 n )︁ − log σ gaus α gaus 2 2 δ + Bnα (πµ,σ ) min L̂n πµ,σ + dK 2n µ∈R ,σ>0 ∥µ−µ0 ∥2 )︂ − dK log σ 2 + log 4δ λ gaus 2 2 + + Varαn (πµ,σ ) , nλ 2 where we used that σ√︃ 0 = 1 since our prior is P = N (η0 µ0 , IdK ) for Gaussian policies. Moreover, we set λ =
2
∥µ−µ0 ∥2 − dK log σ 2 +log 4δ 2 2 α gaus ) n Varn (πµ,σ
.
(Ours, Mixed-Logit) Here we use the mixed-logit policies in (E.27). Thus we only replace the terms, DKL (Q∥P), with their closed-form bound in Lemma 12. This leads to 221
the following objective. √︄ min
µ∈RdK ,σ>0
(︂
(︁ mixL )︁ + L̂αn πµ,σ
√ ∥µ−µ0 ∥2 dK 2 + log 4 n − log σ 2 2 δ
2n +
∥µ−µ0 ∥2 − dK log σ 2 + log 4δ 2 2
nλ
+
mixL ) + Bnα (πµ,σ
)︂ λ mixL Varαn (πµ,σ ) , 2
where we used that σ0 = 1 since√︃our prior is P = N (η0 µ0 , IdK ) × G(0, 1)K for mixed-logit policies. Moreover, we set λ =
2
∥µ−µ0 ∥2 − dK log σ 2 +log 4δ 2 2 α mixL ) n Varn (πµ,σ
.
London and Sandler (2019, Theorem 1) Proposition 18. Let τ ∈ (0, 1), n ≥ 1, δ ∈ (0, 1) and let P be a fixed prior on H, then with probability at least 1 − δ over draws Dn ∼ µnπ0 , the following holds simultaneously for all posteriors, Q, on H that ⌜ (︂ )︂ (︁ ⃓ )︁ )︁ (︁ ⃓ 2 L̂τ (πQ ) + 1 DKL (Q∥P) + log n n 2 DKL (Q∥P) + log nδ ⎷ τ δ τ R (πQ ) ≤ L̂n (πQ ) + + . τ (n − 1) τ (n − 1) (E.31) Baseline 1: (London et al., Gaussian) Here we use the Gaussian policies in (E.29). Thus we only replace the terms, DKL (Q∥P), with their closed-form bound in Lemma 12. This leads to the following objective. ⌜ (︂ )︂ (︂ )︂ ⃓ (︁ )︁ ∥µ−µ0 ∥2 dK 2 + log n ⃓ 2 L̂τ π gaus + 1 (︂ (︁ − log σ )︁ ⎷ n µ,σ τ 2 2 δ gaus min L̂τn πµ,σ + τ (n − 1) µ∈RdK ,σ>0 (︂ )︂ ∥µ−µ0 ∥2 dK n )︂ 2 2 − 2 log σ + log δ 2 + , τ (n − 1) where we used that σ0 = 1 since our prior is P = N (η0 µ0 , IdK ) for Gaussian policies. Baseline 2: (London et al., Mixed-Logit) Here we consider the mixed-logit policies in (E.27). Since the additional Gumbel noise does not affect the KL divergence (Lemma 12), we have the same objective as in the Gaussian case. That is ⌜ (︂ )︂ (︂ )︂ ⃓ (︁ )︁ ∥µ−µ0 ∥2 dK 2 + log n ⃓ 2 L̂τ π mixL + 1 (︂ (︁ − log σ )︁ ⎷ n µ,σ τ 2 2 δ mixL min L̂τn πµ,σ + τ (n − 1) µ∈RdK ,σ>0 (︂ )︂ ∥µ−µ0 ∥2 dK n )︂ 2 2 − 2 log σ + log δ 2 + , τ (n − 1) where we used that σ0 = 1 since our prior is P = N (η0 µ0 , IdK ) × G(0, 1)K for mixed-logit policies. 222
Sakhi et al. (2022, Proposition 1) Proposition 19. Let τ ∈ (0, 1), n ≥ 1, δ ∈ (0, 1) and let P be a fixed prior on H, then with probability at least 1 − δ over draws Dn ∼ µnπ0 , the following holds simultaneously for all posteriors, Q, on H that √
2 n )︂ (︂ D (Q∥P)+log 1 δ −τ λL̂τn (πQ )+ KL n R (πQ ) ≤ min 1−e . λ λ>0 τ (e − 1)
(E.32)
Baseline 3: (Sakhi et al. 1, Gaussian) Here we use the Gaussian policies in (E.29). min
(︂
µ∈RdK ,σ>0,λ>0
(︂ 1 τ gaus 1 − e−τ λL̂n (πµ,σ )+ λ τ (e − 1)
√ ∥µ−µ0 ∥2 dK 2 n − 2 log σ 2 +log 2 δ n
)︂)︂
(E.33)
,
where we used that σ0 = 1 since our prior is P = N (η0 µ0 , IdK ) for Gaussian policies. Baseline 4: (Sakhi et al. 1, Mixed-Logit) Here we consider the mixed-logit policies in (E.27). min
µ∈RdK ,σ>0,λ>0
(︂
(︂ 1 mixL + −τ λL̂τn (πµ,σ ) 1 − e λ τ (e − 1)
√ ∥µ−µ0 ∥2 dK 2 n − 2 log σ 2 +log 2 δ n
)︂)︂
.
(E.34)
where we used that σ0 = 1 since our prior is P = N (η0 µ0 , IdK ) × G(0, 1)K for mixed-logit policies. Sakhi et al. (2022, Proposition 3) Proposition 20. Let τ ∈ (0, 1), n ≥ 1, δ ∈ (0, 1), let P be a fixed prior on H, and let Λ = {λi }i∈[nλ ] a set of nλ positive scalars. Then with probability at least 1 − δ over draws Dn ∼ µnπ0 , the following holds simultaneously for all posteriors, Q, on H and any λi ∈ Λ, √︄ √ (︃ )︃ DKL (Q∥P) + log 4 δ n DKL (Q∥P) + log 2nδ λ λ λ τ L (πQ ) ≤ L̂n (πQ ) + + + g Vnτ (πQ ) , 2n λ n τn (E.35) [︂ ]︂ ∑︁n π0 (A|Xi ) 1 τ and V (π ) = E . where g : u → exp(u)−1−u 2 Q 2 A∼π (·|X ) n i Q i=1 u n max(τ,π (A|X )) 0
i
Baseline 5: (Sakhi et al. 2, Gaussian) Here we consider the Gaussian policies in (E.29).
√︄ min
µ∈RdK ,σ>0,λ∈Λ
(︂
(︁ gaus )︁ L̂τn πµ,σ +
√ ∥µ−µ0 ∥2 dK 2 + log 4 n − log σ 2 2 δ
2n
+
∥µ−µ0 ∥2 − dK log σ 2 + log 2nδ λ 2 2
λ + g n
(︃
λ τn
λ )︃
(︁ gaus )︁ )︂ Vnτ πµ,σ , (E.36)
223
where we used that σ0 = 1 since our prior is P = N (η0 µ0 , IdK ) for Gaussian policies. Baseline 6: (Sakhi et al. 2, Mixed-Logit) Here we consider the mixed-logit policies in (E.27). √︄ min
(︂
µ∈RdK ,σ>0,λ∈Λ
(︁ mixL )︁ L̂τn πµ,σ +
√ ∥µ−µ0 ∥2 dK 2 + log 4 n − log σ 2 2 δ
2n
+
∥µ−µ0 ∥2 − dK log σ 2 + log 2nδ λ 2 2
λ + g n
(︃
λ τn
λ )︃
(︁ mixL )︁ )︂ Vnτ πµ,σ , (E.37)
where we used that σ0 = 1 since our prior is P = N (η0 µ0 , IdK ) × G(0, 1)K for mixed-logit policies.
E.3.4
Additional Results and Discussion
In Figure E.1, we report the reward of the learned policy using one of the considered methods. We make the following observations: √ • Choice of τ and α: in Figure E.1, we set τ = 1/ 4 n ≈ 0.06 and α = 1 − √4 1/ n ≈ 0.94 so that when n is large enough, both L̂τn (π) and L̂αn (π) approach L̂ips n (π) (Ionides, 2008). This is because standard IPS should be preferred when n → ∞. For completeness, we also show in Figure E.2 that the choice of α and τ does not affect the conclusions that we make here. We also include in Figure E.2 the results with an adaptive and data-dependent α obtained using (8.15) in Section 8.3.4. The results in Figure E.2 will be discussed in detail after we finish analyzing the results in Figure E.1. • Overall performance: our method outperforms the baselines for any class of learning policies (Gaussian or mixed-logit) and any choice of logging policies. The only exception is when the logging policy is uniform. • Effect of the class of learning policies: the class of policies, Gaussian or mixedlogit, affects the performance of all the baselines. In general, Gaussian policies behave better than mixed-logit policies. However, this is less significant for our method; the performance of both Gaussian and mixed-logit policies are comparable, and in both cases, our method outperforms the baselines with Gaussian policies. Therefore, in general, Gaussian policies should be preferred over mixed-logit policies. But in case engineering constraints impose the choice of mixed-logit or softmax policies, then the performance of our method is robust to this choice. • Effect of the logging policy: our method reaches the maximum reward even when the logging policy is not performing well. In contrast, the baselines only reach their best reward when the logging policy is already well-performing (η0 ≈ 1), in which case minor to no improvements are made. Note that the baselines have a better reward than ours when the logging policy is uniform. But our method has better reward when the logging policy is not uniform, that is when η0 > 0. This is more common in practice since the logging policy is deployed in production and thus it is expected to perform better than the uniform policy. 224
In Figure E.2, we compare our method to (Sakhi et al. 2) with Gaussian policies since this was the best-performing baseline in our experiments in Figure E.1. Note that we did not include CIFAR100 in Figure E.1 as it was computationally heavy to run these experiments with varying η0 , α and τ for a very high-dimensional dataset such as CIFAR100. We consider 20 varying values of τ and α evenly spaced in (0, 1). We also include the results using the adaptive tuning procedure of α described in Section 8.3.4 (green curve). We make the following observations:
• Adaptive and data-dependent α: This procedure is reliable since the performance with an adaptive α (green curve) is comparable with the best possible choice of α. This is consistent for the three datasets. • Effect of the choice α: as we observed before, the only case where the choice of α may lead to bad-performing policies is when the logging policy is uniform. When the logging policy is not uniform, our method outperforms the best baseline with the best τ for a wide range of values of α. Also, note that there is no very bad choice of α, in contrast with τ ≈ 0 that led to a very bad performing policy that slightly improved upon the logging policy. This attests to the robustness of our method to the choice of α. Moreover, our bound regularizes better α; it contains a bias-variance trade-off term for α. Also, the bound of (Sakhi et al. 2) has a 1/τ making it vacuous for small values of τ . • Best choice of α: To see the effect of α for varying problems, we consider the following experiment. We split the logging policies into two groups. The first is modest logging which corresponds to logging policies whose η0 is between 0 and 0.5. This includes uniform logging policies and other average-performing logging policies. The second is good logging which corresponds to logging policies whose η0 is between 0.5 and 1. After that, for each α, we compute the average reward of the learned policy across either the group of modest or good logging policies. For each dataset, this leads to the two red and green curves in the second row of Figure E.2. Overall, we observe that α ≈ 0.7 leads to the best performance for the modest logging group. Thus when the performance of the logging policy is average, regularizing the importance weights can be critical. In contrast, when the performance of the logging policy is already good, regularization is less needed and we can set α ≈ 1. Fortunately, one of the main strengths of this work is that our bound also holds for standard IPS recovered for α = 1. The bounds in all prior works cannot provide good performance for standard IPS due to their dependency on 1/τ . 225
MNIST, K=10, d=784
0.5
0.6
0.6
0.4
0.5
0.5 0.4 0.3 0.2
0.3
0.4
Logging Ours Sakhi et al. 2 Ours, Adaptive ®
0.2
0.3
0.1
0.2 0.1
0.0
inverse-temperature parameter ´0
inverse-temperature parameter ´0
MNIST, K=10, d=784
FashionMNIST, K=10, d=784
0.80 reward of the learned policy
0.90 0.85 0.80 0.75 0.70
Modest Logging Good Logging
0.65 0.60 0.0
0.2
0.4
0.6
0.8
0.70 0.65 0.60 Modest Logging Good Logging
0.55 0.50 0.0
1.0
0.2
0.4
0.6
0.8
inverse-temperature parameter ´0 EMNIST, K=10, d=784
0.60
0.75
smoothing parameter ®
EMNIST, K=47, d=784
0.6
0.7
0.7
0.1
reward of the learned policy
FashionMNIST, K=10, d=784
0.8
0.8
reward of the learned policy
reward of the learned policy
0.9
0.55 0.50 0.45 0.40 0.35 0.30 Modest Logging Good Logging
0.25 0.20 0.0
1.0
smoothing parameter ®
0.2
0.4
0.6
0.8
1.0
smoothing parameter ®
reward of the learned policy
Figure E.2: In the first row, we report the reward of the learned policy with 20 evenly space values of τ ∈ (0, 1) and α ∈ (0, 1) and varying η0 ∈ [0, 1], and for an adaptive and data-dependent α obtained using (8.15) in Section 8.3.4. The blue-to-cyan colors correspond to different values of τ . The lighter the color, the higher the value of τ . For instance, the cyan lines correspond to high values of τ while the blue ones correspond to very small values of τ . Similarly, the red-to-yellow colors correspond to different values α. The lighter the color, the higher the value of α. For instance, the yellow lines correspond to high values of α while the red ones correspond to very small values of α. Finally, the green curve corresponds to the reward of the learned policy using an adaptive and data-dependent α described in (8.15) (Section 8.3.4). In the second row, we report the average reward of the learned policies using our method across the modest logging group (η0 ∈ [0, 0.5] in red) and the good logging group (η0 ∈ [0.5, 1] in green). MNIST, K=10, d=784
0.9
FashionMNIST, K=10, d=784
0.8
0.8
0.7
0.7
0.6
0.6
0.5
0.5
0.4
0.4
EMNIST, K=47, d=784
0.6 0.5
0.25
0.4
0.20
0.3
0.15
0.2
0.10 0.05
0.3
0.3
0.2
0.2
0.1
0.1 0.0
0.1 0.0
0.0 0.0
0.2
0.4
0.6
0.8
1.0
inverse-temperature parameter ´0 Ours, Gaussian Ours, Mixed-Logit
0.2
0.4
0.6
0.8
1.0
inverse-temperature parameter ´0 London et al., Gaussian London et al., Mixed-Logit
0.2
0.4
0.6
0.8
1.0
inverse-temperature parameter ´0
Sakhi et al. 1, Gaussian Sakhi et al. 1, Mixed-Logit
CIFAR, K=100, d=2048
0.30
0.00 0.0
0.2
0.4
0.6
0.8
1.0
inverse-temperature parameter ´0
Sakhi et al. 2, Gaussian Sakhi et al. 2, Mixed-Logit
Logging
Figure E.1: The reward of the learned policy for four datasets with varying quality of the logging policy η0 ∈ [0, 1].
E.3.5
Learning Principles
Here we compare our bound in Theorem 4 and our learning principle in (8.17) to the one in London and Sandler (2019). We do not include the learning principle in Swaminathan and Joachims (2015a) since the one in London and Sandler (2019) enjoys similar performance and is far more scalable. The learning principle of London and Sandler (2019) is defined 226
as min L̂τn (πµ ) + λ∥µ − µ0 ∥2 . µ
(E.38)
where λ is a tunable hyper-parameters, πµ is the softmax policy defined in (E.26) and µ ∈ RdK is its parameter vector. This learning principle is referred to as (London et al., LP). In contrast, our learning principle is defined as L̂αn (πµ ) + λ1 ∥µ − µ0 ∥2 + λ2 Varαn (πµ ) + λ3 Bnα (πµ ) ,
(E.39)
reward of the learned policy
where λ1 , λ2 and λ3 are tunable hyper-parameters and πµ is the Gaussian policy in (8.13) with a fixed σ = 1. Our learning principle is referred to as (Ours, LP). Finally, our bound in Theorem 4 with Gaussian policies is√referred to as (Ours, Bound). Similarly √4 4 to the previous experiments, we set τ = 1/ n ≈ 0.06 and α = 1 − 1/ n ≈ 0.94 so that when n is large enough, both L̂τn (π) and L̂αn (π) approach L̂ips n (π) (Ionides, 2008). For the learning principles, we tried multiple values of hyper-parameters λ, λ1 , λ2 and λ3 , all between 10−5 and 10−1 . For instance, we found that the best hyper-parameter for London and Sandler (2019) is λ = 10−5 which matches the value they found in their FashionMNIST experiments. For our learning principle, the best hyper-parameters were λ1 = 10−5 , λ2 = 10−5 and λ3 = 10−5 . In contrast, our bound does not require hyperparameter tuning. We report in Figure E.3 the reward of the learned policy on the FashionMNIST for all these methods with varying values of hyper-parameters. To reduce clutter, we only report the reward for good choices of hyper-parameters λ, λ1 , λ2 and λ3 . We observe that for a wide range of hyper-parameters, our learning principle outperforms the one in London and Sandler (2019). However, both learning principles are sensitive to the choice of hyper-parameters. In contrast, our bound does not require the tuning of any additional hyper-parameter and it achieves the best performance except for the uniform logging policy. In addition to being more theoretically grounded, this approach also enjoys favorable empirical performance without additional hyper-parameter tuning, an important practical consideration. FashionMNIST, K=10, d=784
0.8 0.7 0.6 0.5
Logging Ours, LP London et al., LP Ours, Bound
0.4 0.3 0.2 0.1 0.0
0.2
0.4
0.6
0.8
1.0
inverse-temperature parameter ´0
Figure E.3: The reward of the learned policy using either our bound in Theorem 4 (referred to as (Ours, Bound) in green), our learning principle in (8.17) (referred to as (Ours, LP) in red for multiple values of hyper-parameters) or the learning principle in London and Sandler (2019) (referred to as (London et al., LP) in blue) for multiple values of hyper-parameters). 227
E.3.6
Other Importance Weight Corrections
Su et al. (2020); Metelli et al. (2021) also proposed corrections that are different from hard clipping (a detailed comparison is given in Section 8.2). However, they were not included in our main experiments since they do not provide generalization guarantees; they focus on OPE and only propose a heuristic for OPL in their Appendix B.2 and Section 6.1.2, respectively. Those heuristics are not based on theory, in contrast with ours which is directly derived from our generalization bound. However, for completeness, we also compare our regularization of importance weights to theirs. To make such a comparison, we use the hyper-parameters and tuning procedures provided in Section 6 and Appendix B.2 for Metelli et al. (2021) and Sections 5 and 6.1.2 for Su et al. (2020). Overall, we observe in Figure E.4 that our method outperforms these baselines in OPL and the gap is more significant when the logging policy is not performing well.
reward of the learned policy
The reward of the learned policy using one of the baselines with varying quality of the logging policy ´0 2 [0; 1]. MNIST, K=10, d=784
0.9
FashionMNIST, K=10, d=784
0.8
0.8
0.7
0.7
0.6
0.6
0.5 0.4
0.5
0.5
0.3
0.4
0.4
0.2
0.3
0.3
0.2
0.2
0.1 0.0
0.1 0.0
0.2
0.4
0.6
0.8
1.0
inverse-temperature parameter ´0
Ours
EMNIST, K=47, d=784
0.6
0.1 0.2
0.4
0.6
0.8
1.0
inverse-temperature parameter ´0
Meteli et al.(2021)
Su et al.(2020)
0.0 0.0
0.2
0.4
0.6
0.8
1.0
inverse-temperature parameter ´0
Logging
Figure E.4: The reward of the learned policy with varying quality of the logging policy η0 ∈ [0, 1] using either our regularization (α-IPS) or the ones in Su et al. (2020); Metelli et al. (2021).
228
Bibliography
Y. Abbasi-Yadkori, D. Pál, and C. Szepesvári. Improved algorithms for linear stochastic bandits. Advances in neural information processing systems, 24, 2011. M. Abeille and A. Lazaric. Linear Thompson sampling revisited. In Proceedings of the 20th International Conference on Artificial Intelligence and Statistics, 2017. A. Agarwal, D. Hsu, S. Kale, J. Langford, L. Li, and R. Schapire. Taming the monster: A fast and simple algorithm for contextual bandits. In International Conference on Machine Learning, pages 1638–1646. PMLR, 2014. S. Agrawal and N. Goyal. Thompson sampling for contextual bandits with linear payoffs. In Proceedings of the 30th International Conference on Machine Learning, pages 127– 135, 2013a. S. Agrawal and N. Goyal. Further optimal regret bounds for thompson sampling. In Proceedings of the Sixteenth International Conference on Artificial Intelligence and Statistics, pages 99–107, 2013b. S. Agrawal and N. Goyal. Near-optimal regret bounds for thompson sampling. Journal of the ACM (JACM), 64(5):1–24, 2017. P. Alquier. User-friendly introduction to pac-bayes bounds. arXiv:2110.11216, 2021.
arXiv preprint
I. Aouali. Linear diffusion models meet contextual bandits with large action spaces. In NeurIPS 2023 Foundation Models for Decision Making Workshop, 2023. I. Aouali. Diffusion models meet contextual bandits. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. I. Aouali and O. Sakhi. Off-policy learning in large action spaces: Optimization matters more than estimation. arXiv preprint arXiv:2509.03456, 2025. I. Aouali, S. Ivanov, M. Gartrell, D. Rohde, F. Vasile, V. Zaytsev, and D. Legrand. Combining reward and rank signals for slate recommendation. arXiv preprint arXiv:2107.12455, 2021. 229
I. Aouali, A. Benhalloum, M. Bompaire, A. Ait Sidi Hammou, S. Ivanov, B. Heymann, D. Rohde, O. Sakhi, F. Vasile, and M. Vono. Reward optimizing recommendation using deep learning and fast maximum inner product search. In proceedings of the 28th ACM SIGKDD conference on knowledge discovery and data mining, pages 4772–4773, 2022a. I. Aouali, A. Benhalloum, M. Bompaire, B. Heymann, O. Jeunen, D. Rohde, O. Sakhi, and F. Vasile. Offline evaluation of reward-optimizing recommender systems: The case of simulation. arXiv preprint arXiv:2209.08642, 2022b. I. Aouali, A. A. S. Hammou, O. Sakhi, D. Rohde, and F. Vasile. Probabilistic rank and reward: A scalable model for slate recommendation. arXiv preprint arXiv:2208.06263, 2022c. I. Aouali, V.-E. Brunel, D. Rohde, and A. Korba. Exponential smoothing for off-policy learning. In Proceedings of the 40th International Conference on Machine Learning, pages 984–1017. PMLR, 2023a. I. Aouali, B. Kveton, and S. Katariya. Mixed-effect thompson sampling. In International Conference on Artificial Intelligence and Statistics, pages 2087–2115. PMLR, 2023b. I. Aouali, V.-E. Brunel, D. Rohde, and A. Korba. Unified pac-bayesian study of pessimism for offline policy learning with regularized importance sampling. In Uncertainty in Artificial Intelligence, pages 88–109. PMLR, 2024. I. Aouali, V.-E. Brunel, D. Rohde, and A. Korba. Bayesian off-policy evaluation and learning for large action spaces. In International Conference on Artificial Intelligence and Statistics, pages 136–144. PMLR, 2025. P. Auer. Using confidence bounds for exploitation-exploration trade-offs. Journal of Machine Learning Research, 3:397–422, 2002. P. Auer, N. Cesa-Bianchi, and P. Fischer. Finite-time analysis of the multiarmed bandit problem. Machine Learning, 47:235–256, 2002. M. G. Azar, A. Lazaric, and E. Brunskill. Sequential transfer in multi-armed bandit with finite set of models. In Advances in Neural Information Processing Systems 26, pages 2220–2228, 2013. H. Bang and J. M. Robins. Doubly robust estimation in missing data and causal inference models. Biometrics, 61(4):962–973, 2005. H. Bastani, D. Simchi-Levi, and R. Zhu. Meta dynamic pricing: Transfer learning across experiments. CoRR, abs/1902.10918, 2019. URL https://arxiv.org/abs/ 1902.10918. S. Basu, B. Kveton, M. Zaheer, and C. Szepesvari. No regrets for learning the prior in bandits. In Advances in Neural Information Processing Systems 34, 2021. B. Bercu and A. Touati. Exponential inequalities for self-normalized martingales with applications. 2008. 230
C. M. Bishop. Pattern Recognition and Machine Learning, volume 4 of Information science and statistics. Springer, 2006. L. Bottou, J. Peters, J. Quiñonero-Candela, D. X. Charles, D. M. Chickering, E. Portugaly, D. Ray, P. Simard, and E. Snelson. Counterfactual reasoning and learning systems: The example of computational advertising. Journal of Machine Learning Research, 14(11), 2013. S. Bubeck, N. Cesa-Bianchi, et al. Regret analysis of stochastic and nonstochastic multiarmed bandit problems. Foundations and Trends® in Machine Learning, 5(1):1–122, 2012. O. Catoni. Pac-bayesian supervised classification: the thermodynamics of statistical learning. arXiv preprint arXiv:0712.0248, 2007. L. Cella, A. Lazaric, and M. Pontil. Meta-learning with stochastic linear bandits. In Proceedings of the 37th International Conference on Machine Learning, 2020. L. Cella, K. Lounici, and M. Pontil. Multi-task representation learning with stochastic linear bandits. arXiv preprint arXiv:2202.10066, 2022. O. Chapelle and L. Li. An empirical evaluation of Thompson sampling. In Advances in Neural Information Processing Systems 24, pages 2249–2257, 2012. M. Chen, R. Gummadi, C. Harris, and D. Schuurmans. Surrogate objectives for batch policy optimization in one-step decision making. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. W. Chu, L. Li, L. Reyzin, and R. Schapire. Contextual bandits with linear payoff functions. In Proceedings of the 14th International Conference on Artificial Intelligence and Statistics, pages 208–214, 2011. H. Chung, J. Kim, M. T. Mccann, M. L. Klasky, and J. C. Ye. Diffusion posterior sampling for general noisy inverse problems. arXiv preprint arXiv:2209.14687, 2022. M. Cief, J. Golebiowski, P. Schmidt, Z. Abedjan, and A. Bekasov. Learning action embeddings for off-policy evaluation. In European Conference on Information Retrieval, pages 108–122. Springer, 2024. P. Clavier, T. Huix, and A. Durmus. Vits: Variational inference thomson sampling for contextual bandits. arXiv preprint arXiv:2307.10167, 2023. G. Cohen, S. Afshar, J. Tapson, and A. Van Schaik. Emnist: Extending mnist to handwritten letters. In 2017 international joint conference on neural networks (IJCNN), pages 2921–2926. IEEE, 2017. V. Dani, T. P. Hayes, and S. M. Kakade. Stochastic linear optimization under bandit feedback. In COLT, volume 2, page 3, 2008. A. A. Deshmukh, U. Dogan, and C. Scott. Multi-task learning for contextual bandits. In Advances in Neural Information Processing Systems 30, pages 4848–4856, 2017. 231
P. Dhariwal and A. Nichol. Diffusion models beat gans on image synthesis. Advances in neural information processing systems, 34:8780–8794, 2021. M. Dimakopoulou, N. Vlassis, and T. Jebara. Marginal posterior sampling for slate bandits. In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, IJCAI-19, pages 2223–2229. International Joint Conferences on Artificial Intelligence Organization, 2019. M. Dudík, J. Langford, and L. Li. Doubly robust policy evaluation and learning. International Conference on Machine Learning, 2011. M. Dudík, D. Erhan, J. Langford, and L. Li. Sample-efficient nonstationary policy evaluation for contextual bandits. In Proceedings of the Twenty-Eighth Conference on Uncertainty in Artificial Intelligence, UAI’12, page 247–254, Arlington, Virginia, USA, 2012. AUAI Press. M. Dudik, D. Erhan, J. Langford, and L. Li. Doubly robust policy evaluation and optimization. Statistical Science, 29(4):485–511, 2014. A. Durand, C. Achilleos, D. Iacovides, K. Strati, G. D. Mitsis, and J. Pineau. Contextual bandits for adapting treatment in a mouse model of de novo carcinogenesis. In Proceedings of the 3rd Machine Learning for Healthcare Conference, volume 85, pages 67–82, 2018. M. Farajtabar, Y. Chow, and M. Ghavamzadeh. More robust doubly robust off-policy evaluation. In International Conference on Machine Learning, pages 1447–1456. PMLR, 2018. A. Farid and A. Majumdar. Generalization bounds for meta-learning via pac-bayes and uniform stability. Advances in Neural Information Processing Systems, 34:2173–2186, 2021. S. Filippi, O. Cappe, A. Garivier, and C. Szepesvari. Parametric bandits: The generalized linear case. In Advances in Neural Information Processing Systems 23, pages 586–594, 2010. H. Flynn, D. Reeb, M. Kandemir, and J. Peters. Pac-bayes bounds for bandit problems: A survey and experimental comparison. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(12):15308–15327, 2023. D. J. Foster, C. Gentile, M. Mohri, and J. Zimmert. Adapting to misspecification in contextual bandits. Advances in Neural Information Processing Systems, 33:11478– 11489, 2020. G. Gabbianelli, G. Neu, and M. Papini. Importance-weighted offline learning done right. In International Conference on Algorithmic Learning Theory, pages 614–634. PMLR, 2024. G. Garrigos and R. M. Gower. Handbook of convergence theorems for (stochastic) gradient methods. arXiv preprint arXiv:2301.11235, 2023. 232
C. Gentile, S. Li, and G. Zappella. Online clustering of bandits. In Proceedings of the 31st International Conference on Machine Learning, pages 757–765, 2014. A. Gilotte, C. Calauzènes, T. Nedelec, A. Abraham, and S. Dollé. Offline a/b testing for recommender systems. In Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining, pages 198–206, 2018. A. Gilotte, O. Sakhi, I. Aouali, and B. Heymann. Offline contextual bandit with counterfactual sample identification. arXiv preprint arXiv:2509.10520, 2025. A. Gopalan, S. Mannor, and Y. Mansour. Thompson sampling for complex online problems. In Proceedings of the 31st International Conference on Machine Learning, pages 100–108, 2014. B. Guedj. A primer on pac-bayesian learning. arXiv preprint arXiv:1901.05353, 2019. S. Gupta, S. Chaudhari, S. Mukherjee, G. Joshi, and O. Yagan. A unified approach to translate classical bandit algorithms to the structured bandit setting. CoRR, abs/1810.08164, 2018. URL https://arxiv.org/abs/1810.08164. M. Haddouche and B. Guedj. Pac-bayes with unbounded losses through supermartingales. arXiv preprint arXiv:2210.00928, 2022. J. Ho, A. Jain, and P. Abbeel. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020. J. Hong, B. Kveton, M. Zaheer, Y. Chow, A. Ahmed, and C. Boutilier. Latent bandits revisited. In Advances in Neural Information Processing Systems 33, 2020. J. Hong, B. Kveton, S. Katariya, M. Zaheer, and M. Ghavamzadeh. Deep hierarchy in bandits. In Proceedings of the 39th International Conference on Machine Learning, 2022a. J. Hong, B. Kveton, M. Zaheer, and M. Ghavamzadeh. Hierarchical Bayesian bandits. In Proceedings of the 25th International Conference on Artificial Intelligence and Statistics, 2022b. J. Hong, B. Kveton, M. Zaheer, S. Katariya, and M. Ghavamzadeh. Multi-task off-policy learning from bandit feedback. In International Conference on Machine Learning, pages 13157–13173. PMLR, 2023. D. G. Horvitz and D. J. Thompson. A generalization of sampling without replacement from a finite universe. Journal of the American statistical Association, 47(260):663–685, 1952. J. Hu, X. Chen, C. Jin, L. Li, and L. Wang. Near-optimal representation learning for linear bandits and linear rl. In International Conference on Machine Learning, pages 4349–4358. PMLR, 2021. E. L. Ionides. Truncated importance sampling. Journal of Computational and Graphical Statistics, 17(2):295–311, 2008. 233
O. Jeunen and B. Goethals. Pessimistic reward models for off-policy learning in recommendation. In Fifteenth ACM Conference on Recommender Systems, pages 63–74, 2021. Y. Jin, Z. Yang, and Z. Wang. Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pages 5084–5096. PMLR, 2021. E. Kaufmann, N. Korda, and R. Munos. Thompson sampling: An asymptotically optimal finite-time analysis. In International conference on algorithmic learning theory, pages 199–213. Springer, 2012. D. P. Kingma and J. Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. D. P. Kingma, T. Salimans, and M. Welling. Variational dropout and the local reparameterization trick. Advances in neural information processing systems, 28, 2015. D. Koller and N. Friedman. Probabilistic graphical models: principles and techniques. MIT press, 2009. A. Korba and F. Portier. Adaptive importance sampling meets mirror descent: a biasvariance tradeoff. In International Conference on Artificial Intelligence and Statistics, pages 11503–11527. PMLR, 2022. N. Korda, E. Kaufmann, and R. Munos. Thompson sampling for 1-dimensional exponential family bandits. Advances in neural information processing systems, 26, 2013. A. Krizhevsky, G. Hinton, et al. Learning multiple layers of features from tiny images. 2009. I. Kuzborskij and C. Szepesvári. Efron-stein pac-bayesian inequalities. arXiv preprint arXiv:1909.01931, 2019. I. Kuzborskij, C. Vernade, A. Gyorgy, and C. Szepesvári. Confident off-policy evaluation and selection through self-normalized importance weighting. In International Conference on Artificial Intelligence and Statistics, pages 640–648. PMLR, 2021. B. Kveton, M. Zaheer, C. Szepesvari, L. Li, M. Ghavamzadeh, and C. Boutilier. Randomized exploration in generalized linear bandits. In International Conference on Artificial Intelligence and Statistics, pages 2066–2076. PMLR, 2020. B. Kveton, M. Konobeev, M. Zaheer, C.-W. Hsu, M. Mladenov, C. Boutilier, and C. Szepesvari. Meta-Thompson sampling. In Proceedings of the 38th International Conference on Machine Learning, 2021. S. Lam and J. Herlocker. MovieLens Dataset. http://grouplens.org/datasets/movielens/, 2016. T. Lattimore and R. Munos. Bounded regret for finite-armed structured bandits. In Advances in Neural Information Processing Systems 27, pages 550–558, 2014. 234
T. Lattimore and C. Szepesvari. Bandit Algorithms. Cambridge University Press, 2019. B. Laurent and P. Massart. Adaptive estimation of a quadratic functional by model selection. Annals of Statistics, pages 1302–1338, 2000. Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998. L. Li, W. Chu, J. Langford, and R. Schapire. A contextual-bandit approach to personalized news article recommendation. In Proceedings of the 19th International Conference on World Wide Web, 2010. L. Li, W. Chu, J. Langford, and X. Wang. Unbiased offline evaluation of contextualbandit-based news article recommendation algorithms. In Proceedings of the fourth ACM international conference on Web search and data mining, pages 297–306, 2011. L. Li, Y. Lu, and D. Zhou. Provably optimal algorithms for generalized linear contextual bandits. In Proceedings of the 34th International Conference on Machine Learning, pages 2071–2080, 2017. D. Liang and N. Vlassis. Local policy improvement for recommender systems. arXiv preprint arXiv:2212.11431, 2022. D. Lindley and A. Smith. Bayes estimates for the linear model. Journal of the Royal Statistical Society: Series B (Methodological), 34(1):1–18, 1972. B. London and T. Sandler. Bayesian counterfactual risk minimization. In International Conference on Machine Learning, pages 4125–4133. PMLR, 2019. X. Lu and B. Van Roy. Information-theoretic confidence bounds for reinforcement learning. In Advances in Neural Information Processing Systems 32, 2019. R. D. Luce. Individual choice behavior: A theoretical analysis. Courier Corporation, 2012. C. J. Maddison, D. Tarlow, and T. Minka. A* sampling. Advances in neural information processing systems, 27, 2014. O.-A. Maillard and S. Mannor. Latent bandits. In Proceedings of the 31st International Conference on Machine Learning, pages 136–144, 2014. A. Maurer and M. Pontil. Empirical bernstein bounds and sample variance penalization. arXiv preprint arXiv:0907.3740, 2009. D. A. McAllester. Some pac-bayesian theorems. In Proceedings of the eleventh annual conference on Computational learning theory, pages 230–234, 1998. P. McCullagh and J. A. Nelder. Generalized Linear Models. Chapman & Hall, 1989. J. Mei, C. Xiao, B. Dai, L. Li, C. Szepesvari, and D. Schuurmans. Escaping the gravitational pull of softmax. In Advances in Neural Information Processing Systems, volume 33, pages 21130–21140. Curran Associates, Inc., 2020a. 235
J. Mei, C. Xiao, C. Szepesvári, and D. Schuurmans. On the global convergence rates of softmax policy gradient methods. In Proceedings of the 37th International Conference on Machine Learning, ICML’20. JMLR.org, 2020b. A. M. Metelli, A. Russo, and M. Restelli. Subgaussian and differentiable importance sampling for off-policy evaluation and learning. Advances in Neural Information Processing Systems, 34:8119–8132, 2021. A. Nair, A. Gupta, M. Dalal, and S. Levine. Awac: Accelerating online reinforcement learning with offline datasets. arXiv preprint arXiv:2006.09359, 2020. N. Nguyen, I. Aouali, A. György, and C. Vernade. Prior-dependent allocations for bayesian fixed-budget best-arm identification in structured bandits. In International Conference on Artificial Intelligence and Statistics, pages 379–387. PMLR, 2025. M. Papini, A. M. Metelli, L. Lupo, and M. Restelli. Optimistic policy optimization via multiple importance sampling. In International Conference on Machine Learning, pages 4989–4999. PMLR, 2019. A. Peleg, N. Pearl, and R. Meirr. Metalearning linear bandits by prior update. In Proceedings of the 25th International Conference on Artificial Intelligence and Statistics, 2022. J. Peng, H. Zou, J. Liu, S. Li, Y. Jiang, J. Pei, and P. Cui. Offline policy evaluation in large action spaces via outcome-oriented action grouping. In Proceedings of the ACM Web Conference 2023, pages 1220–1230, 2023. X. B. Peng, A. Kumar, G. Zhang, and S. Levine. Advantage-weighted regression: Simple and scalable off-policy reinforcement learning. arXiv preprint arXiv:1910.00177, 2019. J. Peters. Reinforcement learning by reward-weighted regression. In NIPS 2006 Workshop: Towards a New Reinforcement Learning?, 2006. J. Rappaz, J. McAuley, and K. Aberer. Recommendation on Live-Streaming Platforms: Dynamic Availability and Repeat Consumption, page 390–399. Association for Computing Machinery, 2021. I. Rejwan and Y. Mansour. Top-k combinatorial bandits with full-bandit feedback. In ALT, pages 752–776, 2020. S. Rendle, C. Freudenthaler, Z. Gantner, and L. Schmidt-Thieme. Bpr: Bayesian personalized ranking from implicit feedback. arXiv preprint arXiv:1205.2618, 2012. D. A. Reynolds et al. Gaussian mixture models. Encyclopedia of biometrics, 741(659-663), 2009. C. Riquelme, G. Tucker, and J. Snoek. Deep bayesian bandits showdown: An empirical comparison of bayesian deep networks for thompson sampling. arXiv preprint arXiv:1802.09127, 2018. 236
H. Robbins and S. Monro. A stochastic approximation method. The annals of mathematical statistics, pages 400–407, 1951. J. M. Robins and A. Rotnitzky. Semiparametric efficiency in multivariate regression models with missing data. Journal of the American Statistical Association, 90(429): 122–129, 1995. R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. D. Russo and B. Van Roy. Learning to optimize via posterior sampling. Mathematics of Operations Research, 39(4):1221–1243, 2014. D. Russo and B. Van Roy. An information-theoretic analysis of Thompson sampling. Journal of Machine Learning Research, 17(68):1–30, 2016. D. J. Russo, B. Van Roy, A. Kazerouni, I. Osband, Z. Wen, et al. A tutorial on thompson sampling. Foundations and Trends® in Machine Learning, 11(1):1–96, 2018. N. Sachdeva, Y. Su, and T. Joachims. Off-policy bandits with deficient support. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 965–975, 2020. N. Sachdeva, L. Wang, D. Liang, N. Kallus, and J. McAuley. Off-policy evaluation for large action spaces via policy convolution. In Proceedings of the ACM Web Conference 2024, pages 3576–3585, 2024. Y. Saito and T. Joachims. Off-policy evaluation for large action spaces via embeddings. arXiv preprint arXiv:2202.06317, 2022. Y. Saito, Q. Ren, and T. Joachims. Off-policy evaluation for large action spaces via conjunct effect modeling. In international conference on Machine learning, pages 29734– 29759. PMLR, 2023. Y. Saito, J. Yao, and T. Joachims. POTEC: Off-policy contextual bandits for large action spaces via policy decomposition. In The Thirteenth International Conference on Learning Representations, 2025. O. Sakhi, S. Bonner, D. Rohde, and F. Vasile. Blob: A probabilistic model for recommendation that combines organic and bandit signals. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 783– 793, 2020. O. Sakhi, N. Chopin, and P. Alquier. Pac-bayesian offline contextual bandits with guarantees. arXiv preprint arXiv:2210.13132, 2022. O. Sakhi, D. Rohde, and N. Chopin. Fast slate policy optimization: Going beyond plackett-luce. Transactions on Machine Learning Research, 2023. ISSN 2835-8856. URL https://openreview.net/forum?id=f7a8XCRtUu. 237
O. Sakhi, I. Aouali, P. Alquier, and N. Chopin. Logarithmic smoothing for pessimistic offpolicy evaluation, selection and learning. Advances in Neural Information Processing Systems, 37:80706–80755, 2024. S. Scott. A modern bayesian look at the multi-armed bandit. Applied Stochastic Models in Business and Industry, 26:639 – 658, 2010. A. Shrivastava and P. Li. Asymmetric lsh (alsh) for sublinear time maximum inner product search (mips). In Advances in Neural Information Processing Systems, volume 27. Curran Associates, Inc., 2014. M. Simchowitz, C. Tosh, A. Krishnamurthy, D. Hsu, T. Lykouris, M. Dudik, and R. Schapire. Bayesian decision-making under misspecified priors with applications to meta-learning. In Advances in Neural Information Processing Systems 34, 2021. A. Slivkins. Introduction to multi-armed bandits. Foundations and Trends® in Machine Learning, 12(1-2):1–286, 2019. J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International conference on machine learning, pages 2256–2265. PMLR, 2015. Y. Su, M. Dimakopoulou, A. Krishnamurthy, and M. Dudík. Doubly robust off-policy evaluation with shrinkage. In International Conference on Machine Learning, pages 9167–9176. PMLR, 2020. A. Swaminathan and T. Joachims. Batch learning from logged bandit feedback through counterfactual risk minimization. The Journal of Machine Learning Research, 16(1): 1731–1755, 2015a. A. Swaminathan and T. Joachims. The self-normalized estimator for counterfactual learning. advances in neural information processing systems, 28, 2015b. A. Swaminathan, A. Krishnamurthy, A. Agarwal, M. Dudik, J. Langford, D. Jose, and I. Zitouni. Off-policy evaluation for slate recommendation. Advances in Neural Information Processing Systems, 30, 2017. M. F. Taufiq, A. Doucet, R. Cornish, and J.-F. Ton. Marginal density ratio for off-policy evaluation in contextual bandits. Advances in Neural Information Processing Systems, 36, 2024. W. R. Thompson. On the likelihood that one unknown probability exceeds another in view of the evidence of two samples. Biometrika, 25(3-4):285–294, 1933. S. Tomkins, P. Liao, P. Klasnja, and S. Murphy. Intelligentpooling: Practical thompson sampling for mhealth. Machine learning, 110(9):2685–2727, 2021. N. Tripuraneni, C. Jin, and M. Jordan. Provable meta-learning of linear representations. In International Conference on Machine Learning, pages 10434–10443. PMLR, 2021. 238
J. A. Tropp. User-friendly tail bounds for sums of random matrices. Foundations of computational mathematics, 12(4):389–434, 2012. I. Urteaga and C. Wiggins. Variational inference for the multi-armed contextual bandit. In Proceedings of the 21st International Conference on Artificial Intelligence and Statistics, pages 698–706, 2018. T. Van Erven and P. Harremos. Rényi divergence and kullback-leibler divergence. IEEE Transactions on Information Theory, 60(7):3797–3820, 2014. M. Wan, R. Misra, N. Nakashole, and J. J. McAuley. Fine-grained spoiler detection from large-scale review corpora. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pages 2605–2610. Association for Computational Linguistics, 2019. R. Wan, L. Ge, and R. Song. Metadata-based multi-task bandits with Bayesian hierarchical models. In Advances in Neural Information Processing Systems 34, 2021. R. Wan, L. Ge, and R. Song. Towards scalable and robust structured bandits: A metalearning framework. CoRR, abs/2202.13227, 2022. URL https://arxiv.org/abs/ 2202.13227. L. Wang, A. Krishnamurthy, and A. Slivkins. Oracle-efficient pessimism: Offline policy optimization in contextual bandits. arXiv preprint arXiv:2306.07923, 2023. Y.-X. Wang, A. Agarwal, and M. Dudık. Optimal and adaptive off-policy evaluation in contextual bandits. In International Conference on Machine Learning, pages 3589– 3597. PMLR, 2017. Z. Wang, A. Novikov, K. Zolna, J. S. Merel, J. T. Springenberg, S. E. Reed, B. Shahriari, N. Siegel, C. Gulcehre, N. Heess, et al. Critic regularized regression. Advances in Neural Information Processing Systems, 33:7768–7778, 2020. N. Weiss. A Course in Probability. Addison-Wesley, 2005. H. Xiao, K. Rasul, and R. Vollgraf. Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms. arXiv preprint arXiv:1708.07747, 2017. Y. Xu and A. Zeevi. Upper counterfactual confidence bounds: a new optimism principle for contextual bandits. arXiv preprint arXiv:2007.07876, 2020. J. Yang, W. Hu, J. D. Lee, and S. S. Du. Impact of representation learning in linear bandits. arXiv preprint arXiv:2010.06531, 2020. T. Yu, B. Kveton, Z. Wen, R. Zhang, and O. Mengshoel. Graphical models meet bandits: A variational Thompson sampling approach. In Proceedings of the 37th International Conference on Machine Learning, 2020. W. Zhu. Classification of mnist handwritten digit database using neural network. Proceedings of the research school of computer science. Australian National University, Acton, ACT, 2601, 2018. 239
Y. Zhu, D. J. Foster, J. Langford, and P. Mineiro. Contextual bandits with large action spaces: Made practical. In International Conference on Machine Learning, pages 27428–27453. PMLR, 2022.
240
Titre : Apprentissage on-policy et off-policy pour les grands espaces d’actions Mots clés : apprentissage, apprentissage par renforcement, systèmes interactifs, approximation Résumé : Cette thèse étudie l’apprentissage de politiques dans les systèmes interactifs où un agent observe un contexte, choisit une action parmi un très grand ensemble, puis reçoit un retour partiel. Le cadre principal est celui des bandits contextuels, avec deux paradigmes : l’apprentissage en ligne, où l’agent interagit séquentiellement avec l’environnement et minimise le regret, et l’apprentissage hors politique, où il apprend à partir de données journalisées par une politique de logging. Dans les grands espaces d’actions, ces deux cadres soulèvent des difficultés majeures : exploration coûteuse, faible couverture des données, forte variance des poids d’importance, biais d’extrapolation et objectifs difficiles à optimiser. La première partie propose des méthodes bayésiennes structurées pour l’apprentissage en ligne. Nous introduisons meTS, une extension de Thompson sam-
pling fondée sur des effets mixtes, puis dTS, qui exploite des priors inspirés des modèles de diffusion. Ces méthodes partagent l’information entre actions et obtiennent des garanties de regret dépendant d’un nombre effectif d’actions. La seconde partie traite l’apprentissage hors politique. Nous proposons sDM, une méthode directe structurée fondée sur des variables latentes, montrons que l’erreur d’optimisation peut dominer l’erreur d’estimation dans les grands espaces d’actions, et introduisons des objectifs de vraisemblance pondérée par la politique, concaves et efficaces à optimiser. Enfin, nous développons des méthodes pessimistes différentiables fondées sur le lissage exponentiel et des bornes PAC-bayésiennes pour contrôler le compromis biais-variance des estimateurs par importance sampling.
Title : On-Policy and Off-Policy Learning for Large Action Spaces Keywords : learning, reinforcement learning, interactive systems, approximation Abstract : This thesis studies policy learning in interactive systems where an agent observes a context, selects an action from a very large set, and receives partial feedback. The main framework is contextual bandits, with two paradigms: on-policy learning, where the agent interacts sequentially with the environment and minimizes regret, and off-policy learning, where it learns from logged data collected by a logging policy. In large action spaces, both settings face major challenges: inefficient exploration, sparse data coverage, high-variance importance weights, extrapolation bias, and difficult optimization landscapes. The first part develops structured Bayesian methods for on-policy learning. We introduce meTS, a mixed-effect extension of Thompson sampling, and dTS, which le-
Institut Polytechnique de Paris 91120 Palaiseau, France
verages diffusion-inspired priors to model dependencies between actions. These methods share information across actions and yield regret guarantees depending on an effective number of actions. The second part addresses off-policy learning. We propose sDM, a structured direct method based on latent variables, show that optimization error can dominate estimation error in large action spaces, and introduce policy-weighted log-likelihood objectives that are concave and efficiently optimizable. Finally, we develop differentiable pessimistic methods based on exponential smoothing and PAC-Bayesian bounds to control the bias-variance trade-off of regularized importance-sampling estimators.