ConceptioArchivearXiv CS
arXiv CSopen access

Misspecified Explore-then-Exploit Leads to Supra-Competitive Prices

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

Misspecified Explore-then-Exploit Leads to Supra-Competitive Prices

arXiv:2605.16064v1 [cs.GT] 15 May 2026

Jackie Baek∗

Vivek F. Farias†

Farrell Wu‡

Abstract We study whether simple algorithmic pricing systems can systematically produce collusivelike prices in multi-firm markets. We consider firms using an explore-then-exploit pipeline: they randomize prices during an initial exploration phase, then estimate demand from their own historical data and set prices myopically thereafter. The estimation step relies on a misspecified, monopoly-style model that omits competitors’ prices. We characterize when this pipeline converges to supra-competitive prices above the Nash equilibrium, via a fluid-limit ordinary differential equation analysis. We show that supra-competitive prices arise when firms explore within similar price ranges on the same side of the Nash price. Moreover, prices can be substantially above the Nash price; we show that prices can reach monopoly levels under symmetric exploration. Simulations calibrated to a real multifamily rental market confirm that supracompetitive outcomes arise robustly beyond our theoretical assumptions, including under finite horizons, heterogeneous products, and nonlinear logit demand.

1

Introduction

Businesses increasingly rely on algorithms to automate pricing decisions (Brown and MacKay, 2023). This trend has raised a fundamental question: can algorithmic pricing produce collusive-like outcomes? A growing literature suggests that the answer can be yes for certain learning algorithms. For example, the seminal work of Calvano et al. (2020) demonstrates that Q-learning agents can reach and sustain supra-competitive prices through reward-and-punishment dynamics. These results are highly sensitive to the explicit learning dynamics assumed1 . However, it remains unclear whether these findings apply to the pricing systems most commonly deployed in practice: classic reinforcement learning approaches such as Q-learning are complex to implement, require long and costly training, and they are outperformed by simpler pricing alternatives (den Boer et al., 2024). In practice, many firms use far simpler, model-based approaches to algorithmic pricing: estimate a demand curve from their historical data of past prices and realized sales, then apply a myopic ∗

Stern School of Business, New York University, [email protected] Massachusetts Institute of Technology, [email protected] ‡ Massachusetts Institute of Technology, [email protected] 1 For instance, in the case of Calvano et al. (2020), the definition of agent state, discount factor and exploration probability must all be carefully calibrated to get sustained collusion. †

1

“estimate-then-optimize” rule to update prices, often without explicitly modeling competitors’ prices or strategic responses. This approach has been widely studied in the single-agent learning literature under various names such as certainty-equivalence in adaptive control (Simon, 1956), and greedy or exploitation policies in reinforcement learning (Sutton and Barto, 1998). It is also a performant baseline in the revenue management and inventory control literature (Lariviere and Porteus, 1999; Aviv and Pazgal, 2002; Farias and Van Roy, 2010). It is unclear whether these more realistic dynamics systematically produce supra-competitive outcomes. To this end, Cooper et al. (2015), building on earlier work by Kirman (1975, 1986, 1995), investigate sellers that use monopoly (i.e., misspecified) demand models under estimatethen-optimize dynamics. Their finding is discouraging: the limiting price depends entirely on the initial conditions, where any outcome (supra-competitive, sub-competitive, or Nash) is achievable. Furthermore, the limiting price is highly sensitive to the initial conditions, making the outcome practically unpredictable and thus leaving no basis for determining whether collusive-like outcomes will emerge. We argue this unpredictability is an artifact of treating the firms’ initial prices as arbitrary, rather than integrating idiosyncratic variation in the initial prices. We consider a simple adjustment to this setup: an initial exploration phase where firms experiment with (independent) random prices before settling into the estimate-then-optimize exploitation dynamic above. This initial idiosyncratic price variation, in addition to being realistic, allows for a more substantive statement on limiting prices: we prove that when firms initially explore within a well-defined region of similar prices, the limiting prices of all firms are supra-competitive. Crucially, the relevant region of exploration covers a nontrivial fraction of the exploration space and arises naturally when prices are clustered. The mechanism for supra-competitive prices does not require punishment, communication, or explicit coordination; the misspecified pricing dynamic suffices. Computational evidence on both synthetic and real-world-calibrated markets further shows that supra-competitive outcomes persist beyond the conditions of our analytical results and in realistic rental-market calibrations.

1.1

Summary of model and results

Model setup. We study a competitive market with N symmetric firms over T time periods under a linear demand model. Each firm uses an explore-then-exploit policy to choose prices. In the exploration phase, firm i’s price is drawn independently over time from a distribution with mean µi . Then, throughout the exploitation phase, each firm estimates a monopoly-style demand curve in each period using only its own past prices and realized sales. This demand curve omits competitor prices and is thus misspecified. This misspecification reflects a common real-world constraint where firms have limited visibility into competitors’ real-time pricing or choose a simpler model for statistical or computational efficiency. We analyze a “fluid” scaling regime where both the exploration and exploitation horizons grow large. Under this asymptotic scaling, we show that terminal prices converge to a deterministic limit described by a price-moments ordinary differential equation (ODE). This ODE keeps track of the 2

running price means and covariances for all firms, and a careful analysis of this ODE yields the results we describe below. Main theoretical result. Our primary result (Theorem 3.2) states that limiting prices are supracompetitive when firms’ exploration means µi lie on the same side of the Nash price within what we call the best-response cones (left plot of Figure 1). The upper cone yields supra-competitive prices immediately, while the lower cone yields supra-competitive prices after a finite waiting period. Informally, the cone condition is that all firms explore with “similar prices” on the same side of the Nash price, with the allowable price dispersion across firms increasing as prices move further from Nash. More formally, these cones are regions where every firm’s best-response price, if it were computed against the current exploration means, would move all firms’ prices in the same direction. Crucially, under our dynamic, firms do not actually best-respond, nor do they model competitor prices. Rather, firms simply update their prices to maximize revenue as predicted by a misspecified demand model fit to their own historical prices and sales. The core innovation in our analysis lies precisely in analyzing such a dynamic when exploration occurs within a best-response cone. Final price

µ2

Monopoly θ Nash

Nash

45◦

Nash

µ1

Nash Monopoly

s

Figure 1: Left: the shaded regions depict the best-response cones in the (µ1 , µ2 ) plane, where µi is firm i’s average exploration price. We show that the terminal prices are supra-competitive whenever (µ1 , µ2 ) lies in the shaded region. The angle θ depends on the demand parameters but is always at least 45◦ , so the cones cover more than one quarter of the feasible exploration-mean space. Right: Final price under symmetric exploration, in the regime of vanishing exploration variance and an infinite exploitation horizon, where s is the average price during exploration. In a duopoly, the best-response cones cover at least one quarter of the feasible price space for any demand parameters. We also show that in a duopoly, the cone arises endogenously under our dynamics: under any initial conditions, the limit price falls in the best-response cones or the Nash price (Proposition 3.1). This has an iterative implication: even when initial exploration falls outside a cone, the resulting terminal prices land inside one; so if firms re-explore around those prices, supracompetitive outcomes are guaranteed in the next round. For general N , we show that naturally clustered exploration profiles land in a best-response cone with nontrivial probability. In particular, under a random interval prior in which firms’ exploration means are drawn from a common random 3

price band, the probability of landing in a best-response cone is at least 1/4, and this probability increases with N (Proposition 3.2). Quantifying the price gap. Beyond establishing supra-competitive prices, we quantify the size of the price gap above Nash under symmetric exploration. We consider a setting where all firms have the same average exploration price s, where we prove a sharp analytical characterization of the final price as exploration noise vanishes and the exploitation horizon grows (right plot of Figure 1, Theorem 3.3). We show that prices converge to the monopoly price if s is either below the Nash price or above the monopoly price, and when s is between the Nash and monopoly prices, prices converge to s. This reveals that limiting prices can be substantially above the Nash price, potentially reaching the monopoly benchmark. We confirm that the price gap above Nash is substantial beyond the theorem’s assumptions, using numerical ODE evaluation (Section 4.1) and stochastic simulations (Section 4.2). Key mechanism: common-sign movement relative to trailing means. We establish the following key invariant with respect to estimate-then-optimize dynamics: after “sufficient” exploration at prices µi that lie within a best-response cone, estimate-then-optimize dynamics yield, for each firm, price trajectories that move with a common sign relative to trailing means. More precisely, in the lower (resp., upper) cone, the price posted by a firm at each point in time is weakly above (resp., weakly below) the trailing mean of the prices posted by that firm up to that point in time. The exploration phase is crucial because it initializes this common-sign pattern, and we show that once established, this pattern persists throughout the exploitation phase. This invariant is a key ingredient to show that the price changes across firms remain positively correlated (in an appropriate fluid-regime sense), which in turn, leads to supra-competitive limiting prices. Simulations calibrated to a real rental market. In assessing the robustness of our theory there are several issues that merit consideration. In no particular order these include: (1) real-world time scales vs. the fluid regime; (2) nonlinear demand and firm asymmetry; (3) exploration outside of the best-response cones; and (4) a richer supply side with product-specific costs. To address these considerations, we run empirically calibrated simulations to the multifamily rental market in the Greater Boston area (Section 5). Specifically, using demand estimates and market structure from Calder-Wang and Kim (2024), we simulate firms deploying the same explore–then–exploit pricing pipeline in a realistic environment with heterogeneous products, heterogeneous customers, nonlinear logit demand, and product-specific shadow costs calibrated so that observed rents form the Nash benchmark. This analysis shows that the qualitative predictions of our theory persist: supra-competitive terminal prices arise across a broad range of exploration parameters, and the effect is strongest when firms experiment with similar prices on the same side of the Nash benchmark. The finitehorizon simulations further show that these effects appear quickly, not only in the asymptotic fluid regime. These results confirm that supra-competitive outcomes arise robustly across a range 4

of problem parameters and model features that extend beyond the assumptions required for our analytical results, including parameters calibrated to a real-world rental market.

1.2

Related Work

There is a burgeoning literature studying how pricing algorithms can give rise to algorithmic collusion. A commonly studied class of algorithms is reinforcement learning, especially Q-learning, which has been shown to converge to supra-competitive outcomes in repeated oligopoly pricing environments (Calvano et al., 2020; Klein, 2021; Hettich, 2021; Abada et al., 2024; Asker et al., 2022). In many of these models, collusive behavior emerges through reward-and-punishment dynamics: agents learn to coordinate on cooperative price levels and to implement credible punishments following deviations. Such mechanisms typically require extensive training, careful state representation, and the capacity to learn complex, history-dependent strategies. In contrast, we show that much simpler pricing pipelines can systematically produce supra-competitive prices without any punishment logic or learned threats. Closer to our work, a line of research studies learning dynamics under misspecified demand models that ignore the impact of competition (Cooper et al., 2015; Hansen et al., 2021; Banchio and Mantegazza, 2022; Lin and Sarıtaç, 2025; Douglas et al., 2024). The most direct comparison to our work is Cooper et al. (2015), who study a pure exploitation rule in a similar misspecified-demand environment. In their most comparable case (both intercept and slope unknown), long-run prices can be highly path dependent: the system admits infinitely many limit points that depend on initial conditions and may lie above, equal to, or below the Nash equilibrium price. This contrasts with our main result, which provides a sharp characterization of when supra-competitive prices arise. The key difference relative to Cooper et al. (2015) is the explicit exploration phase in our pipeline. In Cooper et al. (2015), there is no exploration or noise; the system is fully deterministic and reaches a steady state in just three time steps. This extremely short trajectory and the lack of price variation severely limit the scope for learning, which is why their long-run outcomes remain so unpredictable and dependent on initial settings. By contrast, our model allows for systematic learning that leads to robustly supra-competitive results. A recent work by Lin and Sarıtaç (2025) studies ordinary least squares dynamics with a misspecified linear demand model, similar to our paper. Their main result derives a stability condition that guarantees convergence to the competitive equilibrium; this is in contrast to our paper’s main result, which is to identify conditions that lead to supra-competitive prices. They also show that supra-competitive prices can be caused by correlation in exploration across firms, or asymmetry in the response time of algorithmic price updates. Under the misspecified demand model, Hansen et al. (2021), Banchio and Mantegazza (2022) and Douglas et al. (2024) identify supra-competitive outcomes under specific learning algorithms in a stylized environment with two actions. Hansen et al. (2021) analyze a UCB bandit policy in a noiseless setting; because both firms run the same deterministic algorithm, their internal learning states coincide after exactly two rounds, so their subsequent actions become fully synchronized 5

in the third round. Douglas et al. (2024) study a similar noiseless setting and generalize this result from UCB to all symmetric deterministic bandit algorithms. Banchio and Mantegazza (2022) study reinforcement learners (e.g., ϵ-greedy Q-learning) in a two-action prisoner’s dilemma and show that when learning is sufficiently slow, payoff estimates can become coupled across firms, sustaining collusive outcomes, although this coupling can disappear under alternative learning-rate specifications. In contrast to these works, we analyze a natural and widely used estimate–then– optimize rule in a continuous action space. Our results do not rely on exact synchronization of algorithms across firms, nor on specific algorithmic parameters. A complementary line of work studies algorithmic collusion in settings where algorithms are explicitly designed to sustain collusive outcomes (Meylahn and den Boer, 2022; Loots and den Boer, 2023; Aouad and den Boer, 2021). Yang et al. (2023) study how fairness regulations on prices interact with collusion in a repeated-game setting. Arunachaleswaran et al. (2024) view deploying an algorithm as a commitment and study the environment as a leader-follower game. They show that if the leader uses an arbitrary no-regret algorithm, then supra-competitive prices arise if the follower approximately optimizes within this environment. Recent papers also evaluate whether large language models can exhibit collusive behavior in pricing environments (Fish et al., 2024; Keppo et al., 2025). On the empirical side, there is growing evidence that pricing algorithms can lead to elevated prices in practice (Chen et al., 2016; Abada and Lambin, 2023; Calder-Wang and Kim, 2024; Assad et al., 2024). We build on the empirical analysis of Calder-Wang and Kim (2024) of the U.S. multifamily rental market by using their calibrated demand model as a more realistic environment than our stylized framework, allowing us to assess the robustness of our results.

2

Model

There are N symmetric firms competing over several periods under a linear demand model with idiosyncratic shocks. Firms follow an explore–then–exploit pricing policy. During the first K periods, firms explore by posting random prices and using the resulting sales data to estimate demand. Thereafter, for the remaining T periods, they exploit by using the estimated demand model to choose prices myopically. Importantly, firms continue to update their demand estimates throughout the exploitation phase as new price and quantity observations arrive. A key feature of the model is that firms estimate a misspecified, monopoly-style demand curve using only their own price and sales history. Thus, although true demand depends on all firms’ prices, each firm treats competitors’ prices as unobserved variation when fitting its own demand curve. We now make this setup precise.

2.1

True demand model

Let Pi,t ∈ R+ denote the price posted by firm i ∈ [N ] in period t ∈ [K + T ]. Firm i realizes quantity Qi,t ∈ R, which depends on the full vector of current-period prices (Pj,t )j∈[N ] through the linear 6

demand system Qi,t = a − bPi,t +

c X Pj,t + εi,t , N −1

i ∈ [N ].

(1)

j̸=i

Here a, b, c > 0. The parameter b captures the own-price effect, while c captures the cross-price effect from competitors’ prices. We assume b > c, so that own-price effects dominate cross-price effects. The demand shocks form a martingale difference sequence with respect to the filtration {Ft }t≥0 generated by the history of prices and quantities:   2 E ε2i,t+1 | Ft ≤ σenv ,

E[εi,t+1 | Ft ] = 0,

for all i ∈ [N ] and t ∈ {0, . . . , K + T − 1}. Thus the shocks may be history-dependent, but they are conditionally mean zero and have uniformly bounded conditional second moments. Constraints on price and quantity. Prices are restricted to the interval [Pmin , Pmax ], where a Pmin > 0. We assume 2b−c < Pmax ≤ ab . The lower bound ensures that the price cap does not force

firms to remain below the competitive Nash price, which we define in Section 2.3. The upper bound ensures that conditional expected demand is nonnegative over the feasible price range. Since Qi,t represents realized sales, we also assume Qi,t ≥ 0 almost surely for all i ∈ [N ] and t ∈ [K + T ].

2.2

Explore-then-exploit dynamics

We next describe how firms choose prices, recalling that the first K periods correspond to the exploration phase, and the remaining T periods correspond to the exploitation phase. Exploration (t ≤ K). During exploration, firms randomize prices according to a fixed distribution Dexp on RN : i.i.d.

(P1,t , . . . , PN,t ) ∼ Dexp ,

t = 1, . . . , K.

We assume prices across firms are independent. Let µ ∈ [Pmin , Pmax ]N denote the mean of 2 2 2 Dexp , and let Σexp := diag(σ1,exp , . . . , σN,exp ) denote its covariance matrix, with σi,exp > 0 for all

i ∈ [N ]. We assume Dexp is supported on [Pmin , Pmax ]N . The exploration phase generates the initial price–quantity histories from which firms begin estimating demand in the exploitation phase. Exploitation (t > K). During exploitation, each firm follows a myopic estimate–then–optimize rule. At each time t ≥ K, firm i uses its own history {(Pi,s , Qi,s )}ts=1 to fit the linear demand curve q(p) = α + βp.

(2)

This model is misspecified because the true demand model (1) also depends on competitors’ prices. Equivalently, firm i treats the variation in demand caused by competitors’ prices as unobserved noise. Let (α̂i,t , β̂i,t ) denote the ordinary least squares estimates of (α, β) based on firm i’s history

7

through period t. The fitted demand curve is q̂i,t (p) := α̂i,t + β̂i,t p. Firm i then chooses the next period’s price by maximizing predicted profit over the feasible price interval: Pi,t+1 :=

arg max

 p α̂i,t + β̂i,t p ,

(3)

t = K, . . . , K + T − 1.

p∈[Pmin ,Pmax ]

This problem has a unique maximizer; see Lemma A.1. To summarize the feedback loop: prices generate data, the data update the fitted demand curve, and the updated fitted demand curve determines the next price.

2.3

Price benchmarks

Lastly, we define two natural price benchmarks: the competitive Nash price and the monopoly price. Competitive benchmark. For firm i and price vector p = (p1 , . . . , pN ), let p̄−i := N 1−1

P

j̸=i pj

denote the average price of its competitors. Suppose firms compete under the correct demand model (1) and observe each other’s posted prices. Then firm i’s profit-maximization problem depends on its competitors’ prices only through p̄−i . The best-response price for firm i is BR(p̄−i ) =

a + c p̄−i . 2b

(4)

The symmetric Nash equilibrium is the fixed point at which every firm charges its best response to the common price charged by its competitors. Thus the unique Nash price, denoted pNE , is pNE =

a . 2b − c

(5)

We also refer to pNE as the competitive price, thus labelling prices above pNE as supra-competitive. Monopoly benchmark. The monopoly price, denoted pMNP , is the symmetric price that would be chosen if firms jointly maximized total industry profits under the same demand system. It is given by pMNP =

a . 2(b − c)

(6)

The Nash price pNE and the monopoly price pMNP therefore form natural benchmarks for contextualizing the prices generated by the learning dynamics. We assume pNE ≥ Pmin , but do not necessarily assume pMNP ≤ Pmax . When c ≤ b/2, the monopoly price satisfies pMNP ≤ a/b.

8

3

Convergence to Supra-Competitive Prices: Analytical Results

We now present the main analytical results. The first step is a convergence result for the stochastic system: in a fluid scaling where the exploration and exploitation horizons grow proportionally, terminal prices converge to a deterministic limit characterized by an ODE that tracks running price means and accumulated covariances (Section 3.1). In Section 3.2, we use this ODE to establish two main results. First, we give a sufficient condition for supra-competitive limiting prices, expressed through best-response cones in the space of exploration means (Theorem 3.2). Second, under symmetric exploration, we characterize the limiting price as a function of the exploration mean, quantifying the gap above Nash (Theorem 3.3). Finally, we discuss the geometry and scope of these cones, showing that they have a simple duopoly interpretation and arise with nontrivial probability under clustered exploration profiles (Section 3.3).

3.1

Terminal convergence and the price-moments ODE

For each pair (K, T ), the explore-then-exploit dynamics in Section 2 induce a random terminal price vector (P1,K+T , . . . , PN,K+T ). We show that when K, T → ∞ with (K + T )/K → α, this random vector converges to the solution of a deterministic price-moments ODE, defined as follows. Definition 3.1 (Price-moments ODE). Fix an exploration-mean vector µ = (µ1 , . . . , µN ) and covariance matrix Σexp . The price-moments ODE is the system for (U (t), V (t)) ∈ RN × RN ×N , t ≥ 1, with U (1) = µ, V (1) = Σexp , and U̇ =

P −U , t

V̇ = (P − U )(P − U )⊤ ,

where P (t) = P (t; µ, Σexp ) ∈ RN is the vector of posted prices, defined componentwise. Writing P P Ū−i := N 1−1 j̸=i Uj and V̄i,−i := N 1−1 j̸=i Vij , (a + c Ū−i )Vii − c Ui V̄i,−i Pei := , 2(bVii − cV̄i,−i )

 [Pe ] i [Pmin ,Pmax ] , Pi := P , max

if − bVii + cV̄i,−i < 0, otherwise.

(7)

The vector U (t) tracks running price means, and the matrix V (t) accumulates price covariances, where Vii captures the accumulated variance of firm i’s own prices and Vij the accumulated covariance between firms i and j. Each Pi (t) is firm i’s posted price at time t; Eq. (7) is firm i’s misspecified OLS price, expressed in terms of the moments (U, V ) (see Appendix A.1 for the derivation). This formula (7) also makes the source of misspecification bias transparent. When V̄i,−i = 0, the posted price reduces to the true best-response Pei = BR(Ū−i ) = (a + cŪ−i )/(2b). When V̄i,−i > 0 (representing positive correlation in firm i’s historical prices with its competitors’), this creates an upward omitted-variable bias, pushing Pei above the true best-response. The ODE uses a scaled horizon t, where t = 1 marks the end of exploration and t corresponds to 9

(K + T )/K in the stochastic model. We abbreviate P ODE (t; µ, Σexp ) := P (t; µ, Σexp ) (or P ODE (t) when µ and Σexp are fixed) for the ODE-implied terminal price at scaled horizon t. The following theorem states that this deterministic limit is the fluid limit of the original stochastic pricing process (proof in Appendix A). Theorem 3.1 (Convergence to the price-moments ODE). Let {(Km , Tm )}m∈N satisfy Km → ∞, Tm → ∞, and (Km + Tm )/Km → α ∈ [1, ∞). Then, (P1,Km +Tm , . . . , PN,Km +Tm ) − → P ODE (α; µ, Σexp ) P

3.2

as m → ∞.

Supra-competitive limiting prices

The previous subsection reduces terminal prices to the deterministic limit P ODE (α; µ, Σexp ). We now state two results of this limit. First, we give a sufficient condition on the exploration-mean vector µ under which the limiting prices are supra-competitive (Theorem 3.2). Second, under symmetric exploration, we characterize the limiting price as a function of the exploration mean (Theorem 3.3). We begin by defining the regions of exploration mean µ that enter the sufficient condition. Definition 3.2 (Best-response cones). For an exploration-mean vector µ ∈ [Pmin , Pmax ]N , write P µ̄−i := N 1−1 j̸=i µj . The upper and lower best-response cones are  C + := µ ∈ [Pmin , Pmax ]N : µi > BR(µ̄−i ) ∀i ∈ [N ] ,  C − := µ ∈ [Pmin , Pmax ]N : µi < BR(µ̄−i ) ∀i ∈ [N ] .

(8) (9)

Membership in C + (resp., C − ) means that every firm explores above (resp., below) its best response to its competitors’ average exploration mean. The left side of Figure 1 in the introduction illustrates the cones for N = 2. We defer a detailed discussion and interpretation of these cones to Section 3.3. We now state our main result, whose proof is in Section 6. Theorem 3.2 (Supra-competitive limiting prices). Under the same assumptions as Theorem 3.1, (a) if µ ∈ C + , then for every α ∈ [1, ∞), P ODE (α; µ, Σexp ) > pNE 1; (b) if µ ∈ C − , then there exists α0 ∈ [1, ∞) such that for all α > α0 , P ODE (α; µ, Σexp ) > pNE 1. Theorem 3.2 identifies simple sufficient conditions on the mean exploration prices µ under which the long-run outcome is supra-competitive. The two cones have slightly different timing implications. If µ ∈ C − , then prices begin below the Nash price, so the exploitation dynamics must first push prices upward; the threshold α0 is the waiting period required for the deterministic dynamics to enter the supra-competitive region. In contrast, if µ ∈ C + , the supra-competitive conclusion holds for every α ∈ [1, ∞). While Theorem 3.2 establishes that limiting prices can exceed the Nash price, it does not specify 2 I , we obtain a sharp by how much. Under symmetric exploration, where µ = s1 and Σexp = σexp N

characterization in a limit where the exploitation horizon grows and the exploration noise vanishes. 10

Theorem 3.3 (Symmetric exploration). Suppose µ = s1 for some s ∈ [Pmin , Pmax ] and Σexp = 2 I . Let p̄MNP := min{pMNP , P σexp max }. Then N

lim

lim P

ODE

σexp →0 α→∞

(α; µ, Σexp ) =

 s 1,

s ∈ [pNE , p̄MNP ],

 MNP p̄ 1,

s < pNE or s > p̄MNP .

The proof is in Appendix B.2. Theorem 3.3 reveals a stark threshold structure, illustrated in the right plot of Figure 1. When s ∈ [pNE , p̄MNP ], the limiting price locks in at the exploration mean s itself. Outside this range, either below Nash or above the capped monopoly price, prices converge all the way to p̄MNP . Thus the price gap above Nash can be substantial, potentially reaching the monopoly benchmark when pMNP ≤ Pmax .2

3.3

Coverage of best-response cones

The sets C + and C − consist of exploration profiles where, if each firm were to best-respond to its competitors’ average exploration price, all firms would adjust their prices in the same direction (a single round of fictitious play). Concretely, µ ∈ C + means each firm i explores at a price µi exceeding its best-response to the average exploration mean of its competitors; µ ∈ C − is the analogous region where every firm explores below that best-response. This is only a characterization of the cones, not what firms actually do: under the misspecified demand model, firms do not compute bestresponses. The actual pricing rule in Definition 3.1 can be decomposed into a best-response term plus a covariance-driven bias. We interpret the cones in two parts. First, in the duopoly case, they admit a simple geometric description and also arise endogenously from the dynamics. Second, for general N , we introduce a random interval prior that models clustered exploration profiles and show that the probability of landing in a cone is at least 1/4 and often much higher, for any N and any demand parameters. 3.3.1

Cone coverage for N = 2

For N = 2, the cones in Definition 3.2 simplify because µ̄−1 = µ2 and µ̄−2 = µ1 :  C + = (µ1 , µ2 ) ∈ [Pmin , Pmax ]2 : µ1 > BR(µ2 ) and µ2 > BR(µ1 ) ,  C − = (µ1 , µ2 ) ∈ [Pmin , Pmax ]2 : µ1 < BR(µ2 ) and µ2 < BR(µ1 ) . Figure 2 illustrates these regions in the (µ1 , µ2 ) plane. The two best-response lines µ2 = BR(µ1 ) and µ1 = BR(µ2 ) intersect at (pNE , pNE ), and C + and C − correspond to the upper-right and lower-left wedges. Informally, the cones capture settings where both firms explore at similar prices and where the first best-response adjustment points in the same direction for both firms. 2

We show that under symmetric histories, the pairwise price correlation ρ also admits a conduct-parameter interpretation, mapping ρ = 0 to the Nash price and ρ = 1 to the capped monopoly price; see Appendix B.1 for details.

11

µ2 µ1 = BR(µ2 ) C+ pNE , pNE

 µ2 = BR(µ1 )

C− µ1 Figure 2: Best-response cones in the duopoly case, shown in the (µ1 , µ2 ) plane for a = b = 1 and c = 1/2. The two best-response lines intersect at (pNE , pNE ); the upper-right and lower-left wedges are C + and C − , respectively. Beyond this geometric interpretation, the cones also arise endogenously from the dynamic: from any exploration price vector µ, the limit price must lie in C + ∪ C − , where C ± denotes the closure of C ± (allowing equality in the defining inequalities (8)-(9)). Proposition 3.1 (Duopoly limit points lie in the closed best-response cones). Take N = 2. Fix a diagonal covariance matrix with positive diagonal entries and an exploration-mean vector µ. If the limit P ∞ (µ, Σexp ) := lim P ODE (α; µ, Σexp ) exists, then P ∞ (µ, Σexp ) ∈ C + ∪ C − . α→∞

The proof is in Appendix C.1. To see why this matters, suppose firms’ exploration prices are themselves the output of an earlier round of estimate-then-optimize pricing, rather than arbitrary points in the feasible set; this is natural if real-world prices inherit structure from past pricing rather than being drawn uniformly. Proposition 3.1 guarantees that such “recycled” exploration prices lie in C + ∪ C − , so by Theorem 3.2, restarting the dynamic from these prices yields supra-competitive outcomes (except at the knife-edge best-response boundaries, which include the Nash point). In this sense, the cones are not an unusual corner of the parameter space but a typical destination of the dynamic. 3.3.2

Cone coverage for general N

We now ask, for general N , how often a randomly drawn exploration profile µ ∈ [Pmin , Pmax ]N lands in a best-response cone. In practice, exploratory prices tend to cluster within a common local band, since firms typically tie their experiments to recently observed market prices or industry-typical price levels. We introduce the following random interval prior to formalize this clustering. Definition 3.3 (Random interval prior). Fix an outer band [P , P ] containing pNE . Two firms draw exploration means uniformly from this band, iid

x, y ∼ Unif[P , P ],

ℓ := min{x, y}, 12

u := max{x, y},

and we set µ1 = x and µ2 = y. The remaining N − 2 firms draw exploration means uniformly from the random interval [ℓ, u]: iid

µi | (ℓ, u) ∼ Unif[ℓ, u],

i = 3, . . . , N.

Two firms act as anchors by drawing from the full outer band, forming a random subinterval [ℓ, u]; the remaining N − 2 firms then cluster within that subinterval. The only input to this construction is the outer band [P , P ]. Letting r := c/b denote the ratio of the cross-price effect to the own-price effect, the following result gives a universal lower bound on cone membership (proof in Appendix C.2). Proposition 3.2 (Cone membership under the random interval prior). Under the random interval prior, for every N ≥ 2 and r ∈ (0, 1),  Pr µ ∈ C + ∪ C − ≥

(N − 1)(2 − r) 1 ≥ . 4(N − 1) − r(N − 2) 4

Under the random interval prior, the probability of landing in a cone is at least 1/4 for any number of firms N and any demand parameters; Table 1 reports the bound for representative values, and the bound converges to (2 − r)/(4 − r) as N → ∞. Perhaps surprisingly, the bound increases in N for each fixed r. This runs counter to a naive volume intuition: under a uniform prior on the hypercube [Pmin , Pmax ]N , the best-response cones shrink exponentially as N grows. Under the random interval prior, however, the inner interval [ℓ, u] is the same for all firms, so additional firms only reinforce the alignment that places µ inside a cone. N 2 3 5 10 ∞

r = 0.25 0.4375 0.4518 0.4591 0.4633 0.4667

r = 0.50 0.3750 0.4006 0.4143 0.4222 0.4286

r = 0.75 0.3125 0.3461 0.3647 0.3756 0.3846

r → 1.00 0.2500 0.2877 0.3094 0.3225 0.3333

Table 1: Lower bounds on the probability of landing in C + ∪ C − under the random interval prior, for representative values of N and r = c/b. The column r → 1.00 is the limiting boundary case.

4

Computational Results via the ODE

The previous section shows that terminal prices are characterized by the deterministic price-moments ODE and gives sufficient conditions for supra-competitive limits. We now use the ODE to computationally evaluate prices beyond the analytical cases, where we focus on the duopoly case (N = 2) for ease of illustration. In this section, we compute the ODE terminal prices over a grid of exploration means, illustrating how the supra-competitive region changes with µ, α and σexp (Section 4.1). 13

Then, we compare these ODE values to the original finite-sample stochastic system from Section 2, showing how the deterministic heatmaps emerge as the exploration sample size grows (Section 4.2).

4.1

ODE-based duopoly simulations

We numerically evaluate the price-moments ODE for N = 2. We fix a = b = 1 and c = 1/2, so that 2 I pNE = 23 and pMNP = 1. We set the covariance during exploration to be Σexp = σexp 2×2 . Figure 3

plots the ODE-implied terminal price P1ODE (α; µ, Σexp ) over a grid of exploration means (µ1 , µ2 ), varying both the scaled horizon α and the exploration noise σexp .

2

2

= 2.0, exp = 0.05

= 5.0, exp = 0.05

= 20.0, exp = 0.05

0.85 0.80 0.75 0.70 0.65 0.60 0.55 0.500.50 0.55 0.60 0.65 0.70 0.75 0.80 0.85 0.50 0.55 0.60 0.65 0.70 0.75 0.80 0.85 0.50 0.55 0.60 0.65 0.70 0.75 0.80 0.85 1

1

1

= 2.0, exp = 0.15

= 5.0, exp = 0.15

= 20.0, exp = 0.15

0.85 0.80 0.75 0.70 0.65 0.60 0.55 0.500.50 0.55 0.60 0.65 0.70 0.75 0.80 0.85 0.50 0.55 0.60 0.65 0.70 0.75 0.80 0.85 0.50 0.55 0.60 0.65 0.70 0.75 0.80 0.85 1

1

1.0

0.9

0.8

0.7 pNE 0.6

0.5

0.4

1

2 I Figure 3: ODE-implied terminal price P1ODE (α; µ, σexp 2×2 ) in the duopoly case. Each panel fixes (α, Σexp ) and varies the exploration means (µ1 , µ2 ). White marks the Nash price pNE = 2/3, while red and blue indicate terminal prices above and below Nash. Thin lines mark the best-response boundaries defining the two cones.

Horizons sharpen cone-like regions. As the horizon α increases (left to right), the heatmaps show convergence that is qualitatively similar to the best-response cones defined in Definition 3.2, though the supra-competitive region is in fact larger than the cones. At larger σexp (bottom plots), this convergence is slower, so the cone-like regions are less sharply separated at the plotted horizons. The diagonal line (µ1 = µ2 ) corresponds to symmetric exploration, and the heatmap shows values

14

consistent with the symmetric exploration result of Theorem 3.3. Specifically, the terminal price is close to the monopoly price when µ1 < pNE , while terminal prices increase with µ1 when µ1 ≥ pNE . Exploration noise pulls prices toward Nash. When the exploration noise σexp is higher (bottom plots), prices are closer to Nash. Because supra-competitive prices arise from correlation in prices across firms, more exploration noise weakens this effect: it gives each firm more uncorrelated own-price variation, diluting the cross-firm correlation that drives the bias. Beyond the duopoly case, Appendix D evaluates the ODE for general N and shows that supra-competitive prices arise robustly under clustered exploration profiles.

4.2

Comparing the ODE to stochastic simulations

We compare the ODE predictions (assuming an asymptotic scaling) to the original discrete-time stochastic system from Section 2 for small, finite horizons. We use the same parameters as Section 4.1 for the demand (a = b = 1 and c = 1/2, so that pNE = 32 and pMNP = 1). We evaluate the stochastic system for (K, T ) ∈ {(10, 50), (100, 500)}, then run the ODE for the corresponding α = (K+T )/K = 6. We fix σexp = 0.05 and set the standard deviation of the random demand shock to 0.05. For each exploration-mean pair (µ1 , µ2 ) ∈ [0.5, 0.85]2 , we run 2500 independent simulation pipelines. Figure 4 compares the mean terminal price of firm 1 to the ODE-implied value P1ODE (α; µ, Σexp ).

Figure 4: ODE and stochastic mean terminal-price heatmaps. The left panel is the deterministic ODE prediction, while the middle and right panels average terminal prices across 2500 stochastic simulation runs. For small K, finite-sample exploration noise blurs the sharp ODE transition regions. As K increases from 10 to 100, the stochastic heatmap becomes visibly closer to the deterministic ODE profile.

Finite samples blur transitions. The stochastic heatmaps are blurred versions of the ODE heatmap, with the blurring diminishing as K grows. This is consistent with Theorem 3.1, which predicts convergence to the ODE profile as K → ∞ with α fixed. When K is small, the exploration phase provides a noisy estimate of the moments that initialize exploitation, so nearby exploration

15

profiles can lead to different dynamic regimes. As K increases, these moments concentrate around their population values and the sharp transitions in the ODE map reappear. Boundary points create mixtures. As the heatmap only presents the mean price over 2500 runs, we also examine the distribution of prices across runs. In Figure 5, we pick two exploration means (µ1 , µ2 ) = (0.66, 0.66) and (µ1 , µ2 ) = (0.75, 0.85), and we plot the distribution of terminal prices across runs. The first point (0.66, 0.66) corresponds to a sharp boundary between Nash-like and monopoly-like outcomes in the ODE heatmap. Finite-sample noise therefore pushes different runs to different sides of the boundary, producing a bimodal distribution: roughly one quarter of runs remain near Nash, while the rest move close to pMNP . By contrast, the second point (0.75, 0.85) lies in a smoother region of the heatmap, where the stochastic noise does not lead to regime switching.

Figure 5: ODE map and terminal-price histograms from stochastic simulations at α = 100, σexp = 0.05, and demand-shock standard deviation 0.05 (250 runs; K = 100, K + T = 10000). The two marked profiles have similar mean terminal prices, but different distributions: (0.66, 0.66) lies near a sharp ODE transition and produces a mixture of Nash-like and monopoly-like outcomes, while (0.75, 0.85) lies in a smoother region and produces a concentrated distribution.

5

Empirically calibrated simulations: Boston multifamily rentals

We complement our analysis with simulations calibrated to a large multifamily rental market, adapting the heterogeneous-logit demand system in Calder-Wang and Kim (2024). We test whether an explore-then-exploit pipeline can generate systematically elevated prices, and evaluate whether the patterns from our theory persist in this richer setting. Relative to our stylized model, this environment features differentiated firms, heterogeneous customers, and nonlinear logit demand.

5.1

Market setup and calibration

Data. We instantiate a market calibrated to the Boston Core multifamily rentals (PUMAs 08001– 08003) using ACS microdata. We exclude rentals below the within-bedroom-count 5th percentile 16

of rent and restrict to adult households with income ≥ $30k. After filtering, our sample comprises N = 801 representative rentals and H = 920 representative households, each weighted by its ACS survey weight. We denote by p0 the vector of observed rents and by s0 the corresponding observed shares (constructed from normalized survey weights). We model each of the N representative rentals as an independent firm offering one product, and thus refer to “rentals” and “firms” interchangeably. Demand-side model. In each period t = 1, . . . , K + T , each household h ∈ {1, . . . , H} chooses among rentals j ∈ {1, . . . , N } (plus an outside option). Given the price vector Pt = (P1,t , . . . , PN,t ) ∈ RN , the total share (demand) for rental j at time t is   exp αh Pj,t + x′j βh + ξj sj,t (Pt ; ξ) = w̃h , P ′ 1+ N k=1 exp αh Pk,t + xk βh + ξk h=1 H X

wh w̃h := PH

ℓ=1 wℓ

,

where xj are observed rental characteristics, ξj is a rental vertical-differentiation fixed-effect, and (αh , βh ) encode renter-specific price sensitivity and characteristic preferences. We adopt the parameterization of (αh , βh ) from Calder-Wang and Kim (2024) and calibrate {ξj }N j=1 to match observed shares s0 at the observed prices p0 (sj (p0 ; ξ) = sj0 ); see Appendix E.1 for details. Supply-side model. We assume that each rental j has a product-specific shadow cost λj > 0 capturing the intertemporal opportunity cost of renting today, so that its period-t objective is  πj,t (Pt , ξ) = Pj,t − λj sj,t (Pt , ξ). We choose λ = (λj )N j=1 so that the observed rent vector p0 is the Nash equilibrium of the calibrated static game. This pins down each shadow cost uniquely via the firm’s first-order condition; the explicit formula is in Appendix E.2. The fitted values are reasonable: the 25th and 75th percentiles of λj /pj0 are 0.817 and 0.877, respectively. Nash and monopoly benchmarks. The Nash benchmark is the observed rent vector p0 by construction, since λ was calibrated to make p0 the Nash equilibrium. The monopoly benchmark pM is defined by joint profit maximization across all rentals; in our calibration, the monopoly markup is relatively uniform, with the 25th and 75th percentiles of pM j /pj0 at 1.236 and 1.263, respectively.

5.2

Explore-then-exploit pricing pipeline

We apply the same explore-then-exploit dynamic introduced in Section 2 at the rental level. During i.i.d.

exploration (t = 1, . . . , K), each rental j posts a perturbed price Pj,t = µj (1 + vj,t ) with vj,t ∼

N (0, σ 2 ) around an exploration mean µj , with price clipped to [0, P̄ ], where P̄ is a large exogenous upper bound. During exploitation (t ≥ K), each rental fits a naive binary logit on its own priceshare history (treating competitors’ prices as unobserved) and sets its next-period rent by myopically 17

maximizing predicted profit. Full equations are in Appendix E.3. Specifying exploration means µ. To mirror the comparative statics in the stylized model, we vary two features of exploration means: their overall level relative to the Nash benchmark and the dispersion of prices across rentals. Specifically, for parameters m, σν ≥ 0, we set µj = m pj0 (1 + νj ),

√ √ i.i.d. νj ∼ Unif(− 3σν , 3σν ).

Thus m shifts the average exploration price vs. Nash, while σν controls cross-rental dispersion.

5.3

Results and Discussion

We report two experiments. First, we vary the exploration-price parameters m, σν , and σ, which respectively control the average exploration level relative to Nash, the cross-rental dispersion in exploration means, and the within-rental exploration noise (Subsection 5.3.1). Second, we vary the time parameters: the exploitation horizon T and the exploration length K (Subsection 5.3.2). For rental j, define the terminal percentage change relative to Nash by  ∆j,T := 100

 Pj,K+T −1 . pj0

The figures plot the 10th, 50th, and 90th percentiles of ∆j,T across rentals. 5.3.1

Supra-competitive prices across exploration-price designs

We fix K = 50 and T = 450 (α = 10), placing the experiment close to the converged regime. We vary the exploration mean multiplier m ∈ {0.80, 0.82, . . . , 1.30} across four combinations of within-rental exploration noise σ and cross-rental dispersion σν , each taking values in {0.02, 0.10}. Figure 6 plots the 10th, 50th, and 90th percentiles of ∆j,T across rentals as a function of m, with the calibrated average monopoly price as a benchmark. Main pattern is consistent with theory. The main finding in Figure 6 is that terminal rents exceed Nash across a wide range of m values and all four panels: supra-competitive outcomes are the norm, not the exception. The median rent change is positive for most of the plotted range, and the 10th percentile lies above Nash except near m = 1. The pattern of when prices are most elevated is consistent with the analytical results. Terminal rents are closest to Nash near m = 1 and rise when exploration is clustered away from Nash on either side, forming the U-shape predicted by Theorem 3.3. The low-m side is especially steep, while the high-m side rises more gradually. Variance drives attenuation to Nash. Cross-rental dispersion σν controls how clustered exploration is across rentals. When σν is small (top panels), exploration means are close together, and the terminal price profile as a function of m closely matches the characterization in Theorem 3.3: 18

% change vs Nash % change vs Nash

50 45 40 35 30 25 20 15 10 5 0 5 10 50 45 40 35 30 25 20 15 10 5 0 5 10

= 0.02, v = 0.02

0.8

0.9

1.0

m

1.1

= 0.10, v = 0.02

90th Percentile 50th Percentile 10th Percentile Avg Monopoly Price

1.2

1.3

0.8

0.9

= 0.02, v = 0.10

0.8

0.9

1.0

m

1.1

1.0

m

1.1

1.2

1.3

1.2

1.3

= 0.10, v = 0.10

1.2

1.3

0.8

0.9

1.0

m

1.1

Figure 6: Terminal rent changes relative to Nash as the exploration mean multiplier m varies. Each panel fixes the within-rental exploration noise σ and cross-rental dispersion σν , with K = 50 and T = 450. The percentile curves summarize the cross-section of rentals, while the horizontal benchmark marks the average monopoly price. The main pattern is the same U-shape predicted by the stylized model: prices are elevated when exploration is clustered away from Nash. prices are near Nash when m ≈ 1 and rise toward the monopoly benchmark as m moves away from Nash in either direction. When σν is larger (bottom panels), this correspondence weakens as exploration means spread across rentals, though supra-competitive outcomes remain over a broad range of m. Within-rental noise σ attenuates supra-competitiveness, especially when m > 1. Larger σ induces more price variation for each rental, reducing the role of endogenous correlation during exploitation. This attenuation is weaker for m < 1, where prices move more during exploitation and therefore generate more correlated price variation after exploration. 5.3.2

Effect of time horizon

We now examine how the time horizon affects terminal prices. We fix σ = σν = 0.05 and vary the exploitation horizon T ∈ [5, 200], comparing two exploration levels (m ∈ {0.80, 1.20}, placing exploration below and above Nash) and two exploration lengths (K ∈ {10, 50}). Figure 7 plots the percentile curves of ∆j,T as a function of T .

19

K = 50, m = 0.80

K = 10, m = 0.80

% change vs Nash

40 20 5

10

5

10

% change vs Nash

15

20

50

100 200

5

10

50

100 200

5

10

T K = 10, m = 1.20

20

50

90th Percentile 50th Percentile 10th Percentile

T K = 50, m = 1.20

100 200

10 5

20

T

20

T

50

100 200

Figure 7: Finite-time dynamics of terminal rent changes relative to Nash. We fix σ = σν = 0.05 and vary the exploitation horizon T on a log scale, comparing K ∈ {10, 50} and m ∈ {0.80, 1.20}. Results show fast convergence. Figure 7 shows that supra-competitive rents emerge quickly and do not require long exploitation horizons. The percentile curves are already above Nash at T = 5 in all four panels, and the curves flatten by moderate horizons such as T = 50. The strongest case is K = 10, m = 0.80: even at the smallest plotted horizons, nearly all rentals are already substantially supra-competitive, with the 10th percentile well above Nash. Across all panels, convergence to a supra-competitive steady state is rapid: most of the price increase relative to Nash occurs within the first T = 50 exploitation periods, and further exploitation yields diminishing additional movement.3 Taken together, these simulations confirm that supra-competitive outcomes arise robustly in a realistic rental market, persisting across a wide range of exploration parameters, finite time horizons, heterogeneous products, and nonlinear demand.

6

Proof of Convergence to Supra-competitive Prices (Theorem 3.2)

We now prove Theorem 3.2, which establishes that limiting prices are strictly supra-competitive, by analyzing the price-moments ODE from Definition 3.1. Throughout, let (U (t), V (t)) denote the ODE state and let P (t) = P (U (t), V (t)) denote the corresponding posted-price vector. Thus the ODE-implied terminal price vector in Theorem 3.2 is P ODE (α; µ, Σexp ) = P (α; µ, Σexp ) in the notation of this proof. 3 The comparison across K should be interpreted through the normalized clock α = (K + T )/K. Holding T fixed, the K = 10 process has a larger α than the K = 50 process, so it is farther along the ODE trajectory. This is why the K = 10 curves tend to be closer to their apparent long-run levels at the same value of T , while the K = 50, m = 0.80 panel starts lower but catches up quickly.

20

For concreteness, we present the argument for the lower cone µ ∈ C − . The proof proceeds in three steps. 1. Key observation: prices remain above historical averages. The central step is to show that under the price-moments ODE, the gap Pi (t) − Ui (t) ≥ 0 for every firm i and all t ≥ 1 (Lemma 6.1). That is, each firm’s price always remains above its historical average. We establish this via forward-invariance of the set {P − U ⪰ 0}: at any boundary point where Pi = Ui for some i, the dynamics push the system inward, so the set is never exited. 2. Correlation drives prices above Nash. Next, we show that the property from Step 1 induces a persistent positive correlation in prices across firms. Specifically, for each firm, the P average cross-covariance V̄i,−i := N 1−1 j̸=i Vij remains positive for all t > 1 (Lemma 6.2). We show that such positive correlation increases prices (Lemma 6.3), and hence we can show that the limiting running mean is weakly above Nash: U ∞ ⪰ pNE 1 (Lemma 6.4). 3. Excluding convergence to exactly Nash. The first two steps leave open the possibility that prices converge exactly to pNE . We rule this out by contradiction: if U (t) → pNE 1, the accumulated covariances stabilize at a strictly positive level, which induces a persistent upward bias that is incompatible with convergence to Nash (Lemmas 6.8–6.9).

6.1

Step 1: Prices remain above historical averages

Our goal is to show that the gap P (t) − U (t) is non-negative for t ≥ 1. Initialization. We first consider t = 1. Since each firm prices independently during exploration, it is easy to show that Pi (1) = BR(µ̄−i ); that is, the prices at the beginning of exploitation are exactly the corresponding best-response prices. Since µ ∈ C − , we have µi < BR(µ̄−i ) for all i, so P (1) − U (1) ≻ 0: all prices jump upward at the start of exploitation. The next lemma shows that this sign pattern is forward invariant, so Pi (t) ≥ Ui (t) for all firms i and all t ≥ 1. Lemma 6.1 (Forward invariance of non-negative gap). Fix t0 ≥ 1. If P − U ⪰ 0 at time t0 , then P − U ⪰ 0 for all t ≥ t0 . The lemma is proved using a forward-invariance argument. We introduce an “energy” quantity Ei (U, V ) whose sign agrees with that of Ui − Pi . To show invariance, consider the first time t1 a coordinate hits the boundary {Ei = 0}. At such a point where Pi − Ui = 0, a direct differentiation of Ei along the ODE gives Ėi (t1 ) ≤ 0. Thus, the trajectory cannot cross into {Ei > 0}, implying that P − U ⪰ 0 persists for all future time.

6.2

Step 2: Correlation drives prices above Nash

Next, we show that Step 1 implies persistent positive correlation: V̄i,−i (t) ≥ 0 for all i and t. Since the off-diagonal entries evolve as V̇ij = (Pi − Ui )(Pj − Uj ), maintaining Pi − Ui ≥ 0 for all firms ensures that cross-covariances remain non-negative—this is formalized in the following lemma. 21

Lemma 6.2. Assume Vij (1) = 0 for all i ̸= j (as under diagonal Σexp ) and P (t) − U (t) ⪰ 0 for all t ≥ 1. If, moreover, P (t0 ) − U (t0 ) ≻ 0 for some t0 ≥ 1 (in particular for t0 = 1 on C − ), then V̄i,−i (t) > 0 for all i and all t > t0 . The following lemma formalizes how positive cross-covariance translates into an upward bias in the OLS price Pei (U, V ), the unclipped OLS price from Definition 3.1. Lemma 6.3 (Correlation bias). Fix i ∈ [N ] and hold U and Vii fixed, with bVii − c V̄i,−i > 0. Then: (1) if V̄i,−i = 0, then Pei (U, V ) = BR(Ū−i ); (2) Pei (U, V ) is strictly increasing in V̄i,−i . Lemma 6.3 implies that when the cross-covariance V̄i,−i is positive, the firm’s algebraic OLS price is strictly higher than the best response, Pei (U, V ) > BR(Ū−i ). A positive cross-covariance therefore biases the posted price upward relative to the best response to the average historical competitor prices. The lemma is proved in Appendix F.3. We show that if the cross-covariance term V̄i,−i (t) remains non-negative, the Nash price serves as a lower bound on the limiting price. Lemma 6.4 (Nash lower bound from nonnegative correlation bias). Along any solution of the pricemoments ODE, assume (i) V̄i,−i (t) ≥ 0 for all i and t ≥ 1, and (ii) U is componentwise monotone in t. Then the limit U ∞ := limt→∞ U (t) exists and satisfies U ∞ ⪰ pNE 1. Remark. In the proof we use an ε-argument: if m∞ := mini Ui∞ < pNE , then for small ε > 0 and all large t, one has Uk (t) ≤ m∞ + ε and Ū−k (t) ≥ m∞ − ε, which yields a uniform lower bound on Pk (t) − Uk (t) and contradicts convergence.

6.3

Step 3: Excluding the Nash boundary

It remains to show U ∞ ≻ pNE 1 (strict inequality) for C − . We first show that if the Nash price is reached by a single firm, then it must be that all firms are at the Nash price. Lemma 6.5 (Reduction to boundary). If U ∞ ⪰ pNE 1 and Uj∞ = pNE for some j, then U ∞ = pNE 1. Therefore, it suffices to rule out U (t) → pNE 1, which we do by contradiction. We start by proving an additional forward-invariance property: once price updates move monotonically upward, they continue to do so. Lemma 6.6 (Forward invariance of the sign of Ṗ on C − ). Assume the hypotheses of Lemma 6.1, so that P − U ⪰ 0 for all t ≥ t0 . Fix t0 ≥ 1. If Ṗ ⪰ 0 at time t0 , then Ṗ ⪰ 0 for all t ≥ t0 . It remains only to initialize the hypothesis of Lemma 6.6. Lemma 6.7 (Initialization of Ṗ on C − ). Suppose µ ∈ C − and U (1) = µ, V (1) = Σexp , where Σexp is positive diagonal. Then P is right-differentiable at t = 1 and Ṗ (1+ ) ≻ 0. 22

By Step 1, P − U ⪰ 0 for all t ≥ 1, and Lemma 6.7 gives Ṗ (1+ ) ≻ 0. Lemma 6.6 therefore gives Ṗ (t) ⪰ 0 for all t ≥ 1. The next result shows that if prices converge to Nash, then P − U decays to zero, and the covariance matrix V (t) stabilizes: Lemma 6.8 (Covariance convergence on the Nash boundary). Given U → pNE 1, and for all t, P − U ⪰ 0 and Ṗ ⪰ 0, (1) P → pNE 1 and P ⪯ pNE 1 for all sufficiently large t. (2) Vii converges to a finite limit Vii∞ and each average cross term V̄i,−i = N 1−1

P

j̸=i Vij converges

∞ > 0. to a strictly positive limit V̄i,−i ∞ > 0 induces a persistent upward bias in the Next, we show that the strictly positive limit V̄i,−i

pricing map via Lemma 6.3(2), which is incompatible with U (t) → pNE 1. Lemma 6.9 (Limiting bias excludes U → pNE 1 for C − ). Suppose µ ∈ C − , Vii → Vii∞ , and V̄i,−i → ∞ > 0. Then, U ∞ ≻ pNE 1. V̄i,−i ∞ > 0. Plugging (U, V ) ≈ (pNE 1, V ∞ ) Proof sketch. By Lemma 6.8, the covariances stabilize with V̄i,−i

into the pricing map and using Lemma 6.3, we obtain a uniform lower bound Pi (t) ≥ pNE + η for some η > 0 and all sufficiently large t. But on C − , where P − U ≥ 0 and Ṗ ≥ 0, this persistent gap prevents U (t) from converging to pNE —a contradiction. Combining Steps 1–3 completes the proof of Theorem 3.2 for µ ∈ C − . Equivalently, since P (t) − U (t) ⪰ 0 on C − , there exists α0 < ∞ such that P ODE (α; µ, Σexp ) ≻ pNE 1

for all α > α0 ,

which proves Theorem 3.2(b). The arguments in Steps 1 and 2 can easily be adapted for the C + case; the details are in Sections F.10 and F.11, respectively. Step 3 is unnecessary for C + , which is explained more in Section F.11. This proves Theorem 3.2(a).

7

Conclusion

This paper demonstrates that simple explore-then-exploit pricing algorithms, relying on misspecified demand models, can lead to supra-competitive prices. We prove that under a scaling limit, the pricing dynamics converge to supra-competitive prices for a broad range of exploration profiles, and the price gap above Nash can be substantial. Simulations empirically calibrated to the Boston rental market confirm that these effects persist under nonlinear demand, product heterogeneity, and realistic market calibration; in fact, they appear quickly even at finite exploration and exploitation 23

horizons. Future work could extend the paper in several directions. One direction is to broaden the theory to relax the current assumptions such as linear demand or symmetric firms. Another is to consider dynamics with persistent demand shocks or ongoing experimentation. Lastly, characterizing when analogous biases arise in other supermodular games would show how broadly the supra-competitive channel extends beyond multi-firm pricing dynamics.

References I. Abada and X. Lambin. Artificial intelligence: Can seemingly collusive outcomes be avoided? Management Science, 69(9):5042–5065, 2023. I. Abada, X. Lambin, and N. Tchakarov. Collusion by mistake: Does algorithmic sophistication drive supra-competitive profits? European Journal of Operational Research, 318(3):927–953, 2024. A. Aouad and A. V. den Boer. Algorithmic collusion in assortment games. Available at SSRN 3930364, 2021. E. R. Arunachaleswaran, N. Collina, S. Kannan, A. Roth, and J. Ziani. Algorithmic collusion without threats. arXiv preprint arXiv:2409.03956, 2024. J. Asker, C. Fershtman, and A. Pakes. Artificial intelligence, algorithm design, and pricing. AEA Papers and Proceedings, 112:452–456, 2022. S. Assad, R. Clark, D. Ershov, and L. Xu. Algorithmic pricing and competition: Empirical evidence from the German retail gasoline market. Journal of Political Economy, 132(3):723–771, 2024. Y. Aviv and A. Pazgal. Pricing of short life-cycle products through active learning. Working paper, Washington University in St. Louis, 2002. M. Banchio and G. Mantegazza. Artificial intelligence and spontaneous collusion. arXiv preprint arXiv:2202.05946, 2022. Z. Y. Brown and A. MacKay. Competition in pricing algorithms. American Economic Journal: Microeconomics, 15(2):109–156, 2023. S. Calder-Wang and G. H. Kim. Algorithmic pricing in multifamily rentals: Efficiency gains or price coordination? Available at SSRN 4403058, 2024. E. Calvano, G. Calzolari, V. Denicolò, and S. Pastorello. Artificial intelligence, algorithmic pricing, and collusion. American Economic Review, 110(10):3267–3297, 2020. L. Chen, A. Mislove, and C. Wilson. An empirical analysis of algorithmic pricing on Amazon Marketplace. In Proceedings of the 25th International Conference on World Wide Web, pages 1339–1349, 2016.

24

W. L. Cooper, T. Homem-de-Mello, and A. J. Kleywegt. Learning and pricing with models that do not explicitly incorporate competition. Operations Research, 63(1):86–103, 2015. A. V. den Boer, J. M. Meylahn, and M. P. Schinkel. Artificial collusion: Examining supracompetitive pricing by Q-learning algorithms. Amsterdam Law School Research Paper No. 2022-25; Amsterdam Center for Law & Economics Working Paper No. 2022-06, 2024. C. Douglas, F. Provost, and A. Sundararajan. Naive algorithmic collusion: When do bandit learners cooperate and when do they compete? arXiv preprint arXiv:2411.16574, 2024. V. F. Farias and B. Van Roy. Dynamic pricing with a prior on market response. Operations Research, 58(1):16–29, 2010. S. Fish, Y. A. Gonczarowski, and R. I. Shorrer. Algorithmic collusion by large language models. arXiv preprint arXiv:2404.00806, 2024. K. T. Hansen, K. Misra, and M. M. Pai. Frontiers: Algorithmic collusion: Supra-competitive prices via independent algorithms. Marketing Science, 40(1):1–12, 2021. M. Hettich. Algorithmic collusion: Insights from deep learning. Available at SSRN 3785966, 2021. J. Keppo, Y. Li, G. Tsoukalas, and N. Yuan. A.I. pricing, agent heterogeneity, and collusion. Available at SSRN 5386338, 2025. A. P. Kirman. Learning by firms about demand conditions. In R. H. Day and T. Groves, editors, Adaptive Economic Models, pages 137–156. Academic Press, New York, 1975. A. P. Kirman. On mistaken beliefs and resultant equilibria. In R. Frydman and E. S. Phelps, editors, Individual Forecasting and Aggregate Outcomes: “Rational Expectations” Examined, pages 147– 168. Cambridge University Press, Cambridge, 1986. A. P. Kirman. Learning in oligopoly: Theory, simulation, and experimental evidence. In A. P. Kirman and M. Salmon, editors, Learning and Rationality in Economics, pages 127–178. Basil Blackwell, Cambridge, MA, 1995. T. Klein. Autonomous algorithmic collusion: Q-learning under sequential pricing. The RAND Journal of Economics, 52(3):538–558, 2021. M. A. Lariviere and E. L. Porteus. Stalking information: Bayesian inventory management with unobserved lost sales. Management Science, 45(3):346–363, 1999. M. Lin and Ö. Sarıtaç. Competition in pricing algorithms: Stability, exploration, and supracompetitive outcomes. Manuscript submitted for review, December 2025. T. Loots and A. V. den Boer. Data-driven collusion and competition in a pricing duopoly with multinomial logit demand. Production and Operations Management, 32(4):1169–1186, 2023. 25

J. M. Meylahn and A. V. den Boer. Learning to collude in a pricing duopoly. Manufacturing & Service Operations Management, 24(5):2577–2594, 2022. H. A. Simon. Dynamic programming under uncertainty with a quadratic criterion function. Econometrica, 24(1):74–81, 1956. R. S. Sutton and A. G. Barto. Reinforcement Learning: An Introduction. MIT Press, Cambridge, MA, 1998. Z. Yang, X. Lei, and P. Gao. Regulating discriminatory pricing in the presence of tacit collusion. Available at SSRN 4633784, 2023.

26

Appendix Table of Contents A Proof of Theorem 3.1: convergence to the price-moments ODE (Section 3.1)

27

B Proofs for the Symmetric-History Results (Section 3.2)

33

C Proofs for the Best-Response Cone Characterizations (Section 3.3)

37

D Additional general-N computational simulations (Section 4.1)

41

E Details for the Empirically Calibrated Simulations (Section 5)

44

F Omitted Lemmas in Proof of Theorem 3.2 (Section 6)

45

G Omitted Lemmas in Proof of Theorem 3.1 (Appendix A)

56

The appendix is organized as follows. 1. Main analytical proofs. Appendix A proves the ODE convergence result supporting Section 3.1; Appendix B proves the symmetric-history results supporting Section 3.2; Appendix C proves the cone-geometry results supporting Section 3.3. 2. Computational and empirically calibrated simulation details. Appendix D gives the general-N ODE simulation details referenced in Section 4.1; Appendix E gives the calibration and implementation details behind Section 5. 3. Deferred technical lemmas. Appendix F supplies the lemmas used in Section 6; Appendix G supplies the stochastic-approximation lemmas used in Appendix A.

A

Proof of Theorem 3.1: convergence to the price-moments ODE (Section 3.1)

This appendix proves Theorem 3.1. The proof is organized in three steps. First, we write the empirical price–quantity moments as a stochastic approximation recursion. Second, we show that this recursion tracks a deterministic mean-field ODE on the fluid time scale. Third, we identify the posted-price coordinate of that ODE with the price-moments ODE in Definition 3.1. The argument is written for the diagonal exploration covariance matrix Σexp from the model section, with positive diagonal entries. Proofs of the auxiliary lemmas stated below are collected in Appendix G. The main point is that the stochastic process and the ODE are written in two different coordinate systems. We first prove convergence in a raw empirical-moment state, where the stochastic

27

approximation is transparent, and only afterward change variables to the centered price moments used in the main text.

A.1

Empirical moments and the OLS price map

We begin by defining a state that contains exactly the information used by each firm’s own-price OLS regression and by the conditional mean of the next observation. For each period t, let Pt = (P1,t , . . . , PN,t )⊤ ,

Qt = (Q1,t , . . . , QN,t )⊤ ,

and define the one-period moment vector  Xt := Pt , Qt , Pt Pt⊤ , Pt ⊙ Qt ∈ RN × RN × RN ×N × RN , where ⊙ denotes componentwise multiplication. Let t

X̄t :=

 1X Xs = P̄t , Q̄t , P P t , P Qt . t s=1

Thus P̄t and Q̄t are running price and quantity means, P P t is the raw price second-moment matrix, and P Qt is the vector of own price–quantity moments. We keep raw second moments in this state because they update linearly as running averages. The centered variances and covariances needed for OLS can then be recovered by subtracting products of first moments. For a generic state x = (xP , xQ , xP P , xP Q ) ∈ RN × RN × RN ×N × RN , define the local own-price variance and own price–quantity covariance terms Si (x) := xPi Q − xPi xQ i .

Ri (x) := xPii P − (xPi )2 ,

On states with Ri (x) > 0, the OLS slope and intercept in firm i’s own-price regression are β̂i (x) =

Si (x) , Ri (x)

P α̂i (x) = xQ i − β̂i (x)xi .

Substituting these coefficients into the myopic one-dimensional pricing problem gives a deterministic price as a function of the empirical moments. This is the key Markovian reduction: the full past history enters future prices only through X̄t .

28

The induced OLS pricing map is # " P xP Q − xP P xQ x  ii i   i Pi Q , if Si (x) < 0, Q P 2 xi − x i xi πi (x) := [Pmin ,Pmax ]    Pmax , if Si (x) ≥ 0. Let π(x) := (π1 (x), . . . , πN (x))⊤ . Lemma A.1 (OLS price from empirical moments). For every exploitation period t ≥ K, Pt+1 = π(X̄t ). Moreover, on the finite-horizon neighborhoods used in the ODE-tracking argument below, the map π is continuous and Lipschitz. After this lemma, proving terminal convergence is reduced to proving convergence of the empirical moment vector X̄t and then applying the continuous price map π.

A.2

Stochastic approximation recursion

We next compute the conditional drift of the moment state during exploitation. Conditional on the current history, the next posted price is already fixed by Lemma A.1; only the demand shocks remain random. For a posted-price vector p ∈ RN , define expected demand by di (p) := a − bpi +

c X pj , N −1

d(p) := (d1 (p), . . . , dN (p))⊤ .

j̸=i

Given the empirical moment state x, the conditional mean of the next observation is  f (x) := π(x), d(π(x)), π(x)π(x)⊤ , π(x) ⊙ d(π(x)) . Define the drift h(x) := f (x) − x.

(10)

Thus f (x) is the population one-period moment vector induced by the price chosen from state x, and h(x) is the gap between this new target observation and the current running average. Lemma A.2 (Conditional mean update). For every exploitation period t ≥ K, E[Xt+1 | Ft ] = f (X̄t ). The running-average identity then gives the stochastic approximation recursion. The factor 1/(t + 1) below is the usual running-average step size. It is this slowly vanishing step size that produces the multiplicative, or fluid, time scale used in the limiting ODE.

29

Lemma A.3 (SA recursion and start condition). Let ξt+1 := Xt+1 − f (X̄t ). Then, for all t ≥ K, X̄t+1 = X̄t +

 1 h(X̄t ) + ξt+1 , t+1

E[ξt+1 | Ft ] = 0.

Moreover, there is a constant σξ < ∞ such that   E ∥ξt+1 ∥22 | Ft ≤ σξ2 . Under exploration, letting µX := E[X1 ] and ΣX := Var(X1 ), E[X̄K ] = µX ,

  tr(ΣX ) E ∥X̄K − µX ∥22 = . K

This lemma gives the two ingredients needed for the fluid limit: the noise is a controlled martingale difference, and the exploration phase initializes the recursion near the deterministic point µX . The exploration mean µX has the following explicit coordinates. Define µ̄−i :=

1 X µj , N −1

Σ̄exp,i,−i :=

j̸=i

1 X (Σexp )ij . N −1 j̸=i

Writing  µX = µP , µQ , µP P , µP Q , we have µQ i = a − bµi + cµ̄−i ,

µPi = µi , µP P = µµ⊤ + Σexp ,

µPi Q = µi µQ i − b(Σexp )ii + cΣ̄exp,i,−i .

These formulas make explicit how the exploration distribution enters the limiting initial condition. In particular, the covariance matrix Σexp affects the initial price–quantity moments and therefore the initial OLS slopes.

A.3

Mean-field ODE and ODE tracking

Dropping the martingale term from the stochastic approximation recursion gives the deterministic mean-field dynamics. The time variable is normalized so that t = 1 corresponds to the end of exploration. The deterministic ODE associated with the recursion in Lemma A.3 is 1 Ṁ (t) = h(M (t)), t

M (1) = µX , 30

t ≥ 1.

(11)

Write  M (t) = M P (t), M Q (t), M P P (t), M P Q (t) , where M P P (t) is an N × N matrix. If P (t) := π(M (t)),

q(t) := d(P (t)),

then (11) has the coordinate form

ṀijP P (t) =

ṀiP (t) =

Pi (t) − MiP (t) , t

ṀiQ (t) =

qi (t) − MiQ (t) , t

Pi (t)Pj (t) − MijP P (t) , t

ṀiP Q (t) =

i, j ∈ [N ],

Pi (t)qi (t) − MiP Q (t) . t

Each coordinate has the same running-average interpretation: the current mean-field moment is pulled toward the population moment generated by the current deterministic price P (t). The next lemma is the fluid-limit step. Lemma A.4 (Mean-field ODE convergence). Let M (·) solve (11). Let {Km }m≥1 satisfy Km → ∞, and let {nm }m≥1 satisfy nm ≥ Km and nm /Km → τ ∈ [1, ∞). Then P

X̄nm − → M (τ ). The proof of Lemma A.4 uses a localized Lipschitz argument. Although the OLS pricing rule contains ratios of empirical moments and is not globally Lipschitz, the mean-field path stays in a region where expected demand is bounded away from zero and own-price variance is bounded below. This gives a fixed tube around the path on which the drift is regular. Thus the possible singularities in the OLS ratios are avoided on every fixed fluid-time window, and the martingale error in Lemma A.3 vanishes in the limit.

A.4

Identification with the price-moments ODE

The convergence lemma is stated in the raw moment coordinates M . The ODE in Definition 3.1, however, is written only in terms of running price means and centered price co-movements, so we now translate between the two descriptions. It remains to rewrite the mean-field ODE in centered price-moment coordinates. Define U (t) := M P (t),

 V (t) := t M P P (t) − U (t)U (t)⊤ .

31

Thus U (t) is the running price mean and V (t) is the scaled centered price second-moment matrix. The scaling by t records accumulated price variation rather than the centered second moment itself. This normalization is what gives V the simple evolution V̇ = (P − U )(P − U )⊤ . Lemma A.5 (Mean-field ODE in price-moments coordinates). Let M (·) solve (11), and define (U, V ) as above. Then (U, V ) solves U (1) = µ,

V (1) = Σexp ,

U̇ =

P −U , t

V̇ = (P − U )(P − U )⊤ ,

where P (t) = π(M (t)) is given componentwise as follows. For each i ∈ [N ], let Ū−i :=

1 X Uj , N −1

V̄i,−i :=

j̸=i

Then Pi =

 [Pe ] i [P

min ,Pmax ]

j̸=i

if − bVii + cV̄i,−i < 0,

,

P , max with

1 X Vij . N −1

otherwise,

(a + cŪ−i )Vii − cUi V̄i,−i . Pei = 2(bVii − cV̄i,−i )

Equivalently, the mean-field coordinates are recovered from (U, V ) by MiP (t) = Ui (t),

MiQ (t) = a − bUi (t) + cŪ−i (t),

1 M P P (t) = U (t)U (t)⊤ + V (t), t and MiP Q (t) = Ui (t)MiQ (t) +

 1 −bVii (t) + cV̄i,−i (t) . t

The recovery formulas show that no information relevant for the posted-price coordinate is lost: once (U, V ) is known, the quantity moments appearing in the raw OLS state are pinned down by the linear demand model. These dynamics and the posted-price coordinate P (t) coincide with the price-moments ODE in Definition 3.1. This completes the deterministic identification step: the stochastic approximation is tracked in the full moment state, but its limiting posted prices can be read from the price-moments coordinates of Definition 3.1.

A.5

Terminal convergence

Proof of Theorem 3.1. We now combine the preceding pieces. The empirical moment state converges to the mean-field state, the terminal price is a continuous function of that state, and the resulting deterministic price is the posted-price coordinate of the price-moments ODE.

32

Let {(Km , Tm )}m≥1 satisfy Km → ∞,

Km + T m → α ∈ [1, ∞). Km

Tm → ∞,

Set nm := Km + Tm − 1. Since nm Km + Tm 1 = − → α, Km Km Km Lemma A.4 gives P

→ M (α). X̄nm − This is convergence of the full empirical state through the last period whose data are used to compute the terminal exploitation price. The terminal exploitation price is computed from the empirical moments through period Km + Tm − 1. Therefore, by Lemma A.1, PKm +Tm = π(X̄nm ). Because π is continuous in a neighborhood of the limiting path, the continuous mapping theorem implies P

PKm +Tm = π(X̄nm ) − → π(M (α)). The remaining task is only notational: the limit π(M (α)) is still expressed in the raw moment coordinates, while the theorem states the limit in the price-moments notation. By Lemma A.5, the vector π(M (t)) is exactly P (t), the posted-price coordinate of the pricemoments ODE initialized at (U (1), V (1)) = (µ, Σexp ). In the notation of Definition 3.1, this gives π(M (α)) = P ODE (α; µ, Σexp ). Therefore (P1,Km +Tm , . . . , PN,Km +Tm ) − → P ODE (α; µ, Σexp ), P

as claimed. This proves the desired terminal convergence and completes the reduction from the finite-sample explore–then–exploit process to the deterministic price-moments ODE.

B

Proofs for the Symmetric-History Results (Section 3.2)

This appendix contains two symmetric-history calculations. First, we formalize the historicalcorrelation interpretation from Section 3.2: in a symmetric history, the common pairwise price 33

correlation ρ generates the same symmetric price as an auxiliary conduct game in which each firm internalizes rivals’ profits with weight ρ. Second, we prove the symmetric-exploration limit in Theorem 3.3. Throughout, BR(x) = (a + cx)/(2b) and p̄MNP := min{pMNP , Pmax }.

B.1

Historical correlation as an implied conduct parameter

Consider a symmetric price history in the price-moments state. The OLS demand regression is misspecified because each firm regresses demand on its own price while omitting rivals’ prices, even though rivals’ prices affect demand. The relevant moments are (U, V ): Ui is firm i’s historical average price, Vii is its accumulated own-price variance, and Vij is the accumulated covariance between firms i and j. A symmetric equicorrelated history has Ui = u, Vii = v > 0, and Vij = ρv for every pair i ̸= j. Then

Vij p = ρ, Vii Vjj

so ρ is the common pairwise correlation of firms’ historical price paths. We show that this same ρ also has a conduct interpretation: the self-consistent OLS price is the same as the symmetric price in an auxiliary game where each firm acts as if it internalizes a fraction ρ of rivals’ profits. Thus ρ = 0 corresponds to Nash pricing, while ρ = 1 corresponds to full joint-profit internalization. For the calculation, recall the OLS pricing rule 

 (a + cŪ−i )Vii − cUi V̄i,−i Pi = , 2(bVii − cV̄i,−i ) [Pmin ,Pmax ]

Ū−i :=

1 X Uj , N −1

V̄i,−i :=

j̸=i

1 X Vij , N −1 j̸=i

on the branch where bVii − cV̄i,−i > 0. In the symmetric equicorrelated state below, this condition holds because bVii − cV̄i,−i = (b − cρ)v > 0. Now let πi (p) = pi (a − bpi + cp̄−i ),

p̄−i :=

1 X pj . N −1 j̸=i

The auxiliary ρ-conduct game is the static game in which firm i chooses pi ∈ [Pmin , Pmax ] to maximize πi (p) + ρ

X

πj (p).

j̸=i

The equivalence below is an outcome equivalence: the OLS rule and this auxiliary conduct game have the same symmetric fixed-point price, although they need not have the same off-equilibrium best-response map. Lemma B.1 (Historical correlation as an implied conduct parameter). Fix N ≥ 2 and ρ ∈ [0, 1]. At any symmetric equicorrelated price-moments state with Ui = u, Vii = v > 0, and Vij = ρv for all

34

i ̸= j, the misspecified OLS rule reduces to  a + c(1 − ρ)u Bρ (u) = . 2(b − cρ) [Pmin ,Pmax ] 

Its symmetric self-consistent price is  pρ = min Pmax ,

a 2b − c(1 + ρ)

 .

This same pρ is the symmetric Nash price of the auxiliary ρ-conduct game. Proof. At the stated symmetric equicorrelated state, Ū−i = u, V̄i,−i = ρv, and bVii − cV̄i,−i = (b − cρ)v > 0. Substituting into the OLS pricing rule gives 

   (a + cu)v − cuρv a + c(1 − ρ)u Pi = = =: Bρ (u). 2(bv − cρv) 2(b − cρ) [Pmin ,Pmax ] [Pmin ,Pmax ] A symmetric self-consistent OLS price solves p = Bρ (p). Ignoring the price caps, so

2(b − cρ)p = a + c(1 − ρ)p,

p∗ρ =

a . 2b − c(1 + ρ)

As ρ increases from 0 to 1, this price moves from the Nash price a/(2b − c) to the monopoly price a/[2(b − c)]. Thus, under the maintained assumption that the lower cap does not bind at Nash, only the upper cap can bind, giving pρ = min



Pmax , p∗ρ



a = min Pmax , 2b − c(1 + ρ)

 .

It remains to check that the auxiliary conduct game has the same symmetric price. Differentiating firm i’s auxiliary objective with respect to pi gives  ∂  πi (p) + ρ ∂pi

 X

πj (p) = a − 2bpi + cp̄−i + ρcp̄−i = a − 2bpi + c(1 + ρ)p̄−i .

j̸=i

Therefore the unconstrained best response in the auxiliary game is pi =

a + c(1 + ρ)p̄−i . 2b

At a symmetric Nash equilibrium, pi = p̄−i = p, so (2b − c(1 + ρ))p = a, which gives the same un-capped price p∗ρ . Applying the same upper price cap yields pρ . Thus ρ gives a simple calibration of the implied degree of collusion: ρ = 0 rationalizes the 35

Nash price, ρ = 1 rationalizes the capped monopoly benchmark, and intermediate values rationalize intermediate symmetric prices.

B.2 Proof.

Proof of Theorem 3.3: symmetric-exploration price limit The proof reduces the symmetric price-moments ODE to one scalar trajectory. We then

use the fixed sign of the price-minus-mean gap to get convergence, relate terminal displacement to accumulated co-movement, and pass to the vanishing-exploration-noise limit. 2 I . By permutation equivariance of the price-moments ODE Fix σexp > 0 and set Σexp = σexp N

in Definition 3.1 and of the posted-price map (7), the symmetric initial condition U (1) = s1 and 2 I remains symmetric. Thus there are scalar functions4 u, v, C, p such that U (t) = u(t), V (1) = σexp i N

Vii (t) = v(t), Vij (t) = C(t) for i ̸= j, and Pi (t) = p(t). They satisfy u̇ =

p−u , t

v̇ = Ċ = (p − u)2 ,

u(1) = s,

2 v(1) = σexp ,

C(1) = 0,

2 . Since −bv + cC = −(b − c)C − bσ 2 < 0, the non-OLS branch never occurs. The so v − C ≡ σexp exp

common unclipped price is peσexp (u, C) =

2 ) + cσ 2 u a(C + σexp exp , 2 ) 2((b − c)C + bσexp

  p(t) = peσexp (u(t), C(t)) [P

min ,Pmax ]

.

Let e(t) := p(t) − u(t). At t = 1, the unclipped price is BR(s), and clipping preserves the sign of BR(s)−s. Since BR(s)−s = (a−(2b−c)s)/(2b), this sign is − sign(s−pNE ). By the sign-invariance lemma, e(t) ≥ 0 if s < pNE , e(t) ≤ 0 if s > pNE , and e(t) ≡ 0 if s = pNE . Hence u(t) is monotone and bounded, so u∞ (σexp ) := limt→∞ u(t) exists; also C(t) increases to some C∞ (σexp ) ∈ [0, ∞]. We first record the terminal relation between u∞ and C∞ . If C∞ = ∞, then peσexp (u(t), C(t)) → MNP p , so p(t) → p̄MNP ; since u converges and du/d log t = e has fixed sign, e(t) → 0 and therefore u∞ = p̄MNP . If instead u∞ < p̄MNP , then C∞ < ∞, the upper clip is eventually inactive, and u∞ = peσexp (u∞ , C∞ ). Solving this identity and using Ċ = e2 = t2 u̇2 gives 2 C∞ = σexp

2b − c u∞ − pNE , 2(b − c) pMNP − u∞

|u∞ − s|2 =

Z ∞

2

t−1 (tu̇(t)) dt

≤ C∞ .

1

Consequently, whenever u∞ < p̄MNP ,  |u∞ − s| ≤ σexp

2b − c u∞ − pNE 2(b − c) pMNP − u∞

1/2 .

The same relation also shows u∞ ≥ pNE ; otherwise the displayed formula for C∞ would be negative. We also need an upper bound. If pMNP ≥ Pmax , then u∞ ≤ Pmax = p̄MNP . If pMNP < Pmax , 4

The lowercase notation is local to the symmetric reduction. The full ODE variables remain the vector/matrix objects U (t), V (t), and P (t).

36

then peσexp (u, C) − u =

2 2(b − c)(pMNP − u)C + (2b − c)(pNE − u)σexp . 2 ) 2((b − c)C + bσexp

Thus u(t) ≥ pMNP + δ eventually would imply p(t) − u(t) ≤ −η eventually, uniformly over C(t) ≥ 0, contradicting convergence of u(t). Hence pNE ≤ u∞ (σexp ) ≤ p̄MNP for every σexp > 0. Now let σexp ↓ 0. If s = pNE , then e ≡ 0 and u∞ (σexp ) = s. If s ∈ (pNE , p̄MNP ), monotonicity gives u∞ (σexp ) ≤ s, and the factor in the width bound is uniformly bounded, so u∞ (σexp ) → s. If s = p̄MNP , the same conclusion follows by contradiction: any subsequence with u∞ (σk ) ≤ s − ε has a uniformly bounded width factor and is forced to converge to s. Therefore u∞ (σexp ) → s for all s ∈ [pNE , p̄MNP ]. If s < pNE , then u∞ (σexp ) ≥ pNE > s. A subsequence satisfying u∞ (σk ) ≤ p̄MNP − ε would again have a uniformly bounded width factor and would be forced to converge to s, a contradiction. Hence u∞ (σexp ) → p̄MNP . Finally, if s > p̄MNP , then necessarily p̄MNP = pMNP < Pmax ; the upper bound gives u∞ (σexp ) ≤ pMNP , and any subsequence bounded below pMNP by a fixed ε is ruled out by the same width-bound argument. Thus u∞ (σexp ) → pMNP = p̄MNP . For each fixed σexp > 0, the preceding terminal argument gives p(t) − u(t) → 0. Since the 2 I ) = p(α)1, we have symmetric trajectory satisfies P ODE (α; s1, σexp N 2 lim P ODE (α; s1, σexp IN ) = u∞ (σexp )1.

α→∞

Combining this identity with the scalar limits above gives exactly the two cases in Theorem 3.3.

C

Proofs for the Best-Response Cone Characterizations (Section 3.3)

This appendix proves the two auxiliary propositions stated in Section 3.3. Throughout, BR(x) = (a + cx)/(2b) and C + , C − are the best-response cones defined in (8)–(9). We write P (t) for the posted-price coordinate of the price-moments ODE and U (t) for the running price mean.

C.1

Proof of the duopoly limit-points proposition

We prove the closed-cone version of the statement: P ∞ (µ, Σexp ) ∈ C + ∪ C − . Consequently, if the limiting point is not on a one-sided best-response boundary, then P ∞ (µ, Σexp ) ∈ C + or P ∞ (µ, Σexp ) ∈ C − . The only point at which both best-response equalities hold is pNE 1. Proof.

Consider the duopoly case N = 2 and suppose P ∞ (µ, Σexp ) := lim P ODE (α; µ, Σexp ) = lim P (t) α→∞

t→∞

37

exists. Write P ∞ := P ∞ (µ, Σexp ). Since U (t) is the running average of posted prices, 1 Ui (t) = t

Z t

 Ui (1) +

 Pi (τ ) dτ

,

1

Cesàro convergence implies U (t) → P ∞ . Define the deviations from the true best responses ∆1 (t) := P1 (t) − BR(U2 (t)),

∆2 (t) := P2 (t) − BR(U1 (t)).

We first show that ∆1 (t) and ∆2 (t) always have the same weak sign. On the OLS branch, a direct rearrangement of the price-moments pricing formula gives  cV a + cU − bU 12 2 1  , Pe1 − BR(U2 ) = 2b bV11 − cV12

 cV a + cU − bU 12 1 2  . Pe2 − BR(U1 ) = 2b bV22 − cV12

The terms a + cU2 − bU1 and a + cU1 − bU2 are strictly positive because Ui ∈ [Pmin , Pmax ] and a − bPmax + cPmin > 0. Hence, on the OLS branch, both unclipped deviations have the sign of V12 . Clipping preserves this weak sign: since pNE ∈ [Pmin , Pmax ] and BR is increasing with fixed point pNE , we have for all x ∈ [Pmin , Pmax ].

Pmin ≤ BR(x) ≤ Pmax

Thus clipping at Pmin can only occur when the unclipped price is below BR, and clipping at Pmax can only occur when the unclipped price is above BR. Finally, on the non-OLS branch, the pricing rule sets Pi = Pmax ; this branch can only occur when the corresponding cross-covariance term is nonnegative, and again Pi − BR(Uj ) ≥ 0. Therefore, for every t, ∆1 (t)∆2 (t) ≥ 0. Passing to the limit and using continuity of BR yields   P1∞ − BR(P2∞ ) P2∞ − BR(P1∞ ) ≥ 0. Thus either both deviations are weakly positive or both are weakly negative. In the first case P ∞ ∈ C + ; in the second case P ∞ ∈ C − . If both deviations are zero, then P1∞ = BR(P2∞ ),

P2∞ = BR(P1∞ ).

Since BR(x) = (a + cx)/(2b) is affine with slope c/(2b) < 1, this fixed-point system has the unique solution P1∞ = P2∞ =

a = pNE . 2b − c

Hence the only point at which both best-response equalities bind is pNE 1. 38

C.2 Proof.

Proof of Proposition 3.2 It is convenient to recenter prices around Nash. Let zi := µi − pNE ,

λ :=

c r = . 2b 2

Since BR(pNE + z) − pNE = λz, the cone conditions become     X 1 zj for all i , C + = zi > λ   N −1

    X 1 C − = zi < λ zj for all i .   N −1

j̸=i

j̸=i

Write the outer band in the random interval prior as [P , P ] = [pNE −L, pNE +H], with L, H > 0. Let x, y be the two outer-band draws, and let ℓ := min{x, y} and u := max{x, y} as in the mainbody definition of the prior. Conditional on both recentered anchors being positive, equivalently ℓ > pNE , write zℓ := ℓ − pNE ,

zu := u − pNE ,

so 0 < zℓ < zu . The remaining N − 2 firms draw independently from [ℓ, u], so after recentering we may write zk = zℓ + (zu − zℓ )Yk ,

k = 3, . . . , N,

where the Yk are i.i.d. Unif[0, 1]. Let q :=

zℓ , zu

Y :=

N X

Yk .

k=3

Conditional on both recentered anchors being positive, the ratio q is uniform on [0, 1] and independent of Y . When all zi are positive, membership in C + is determined by the smallest coordinate zℓ : since the inequality

P zi > λ

j̸=i zj

N −1

is most stringent for the smallest zi , it is necessary and sufficient to check it at zi = zℓ . Thus µ ∈ C + whenever zℓ > λ

P zu + N k=3 zk . N −1

Dividing by zu and substituting zk /zu = q + (1 − q)Yk , this condition is equivalent to q > θ(Y ),

θ(Y ) :=

λ(1 + Y ) . (N − 1) − λ(N − 2 − Y )

39

Since q ∼ Unif[0, 1] independently of Y , Pr(µ ∈ C + | both recentered anchors positive) = E[1 − θ(Y )]. The function y 7→ θ(y) is concave on [0, N − 2], because  2 (N − 1) − λ(N − 1) 2λ θ′′ (y) = − 3 < 0. (N − 1) − λ(N − 2 − y) Therefore Jensen’s inequality gives 

N −2 E[θ(Y )] ≤ θ(EY ) = θ 2

 =

Nλ . 2(N − 1) − λ(N − 2)

Hence Pr(µ ∈ C + | both recentered anchors positive) ≥ 1−

2(N − 1)(1 − λ) Nλ = . 2(N − 1) − λ(N − 2) 2(N − 1) − λ(N − 2)

The same argument applies conditional on both recentered anchors being negative, equivalently u < pNE , after replacing zi by −zi , and gives the identical lower bound for C − . If π+ :=

H , L+H

π− :=

L , L+H

2 and the probability that both then the probability that both recentered anchors are positive is π+ 2 . Therefore recentered anchors are negative is π− 2 2 Pr(µ ∈ C + ∪ C − ) ≥ (π+ + π− )

2(N − 1)(1 − λ) . 2(N − 1) − λ(N − 2)

2 + π 2 ≥ 1/2. Thus Since π+ + π− = 1, we have π+ −

Pr(µ ∈ C + ∪ C − ) ≥

(N − 1)(1 − λ) . 2(N − 1) − λ(N − 2)

Pr(µ ∈ C + ∪ C − ) ≥

(N − 1)(2 − r) . 4(N − 1) − r(N − 2)

Substituting λ = r/2 yields

It remains to show that this lower bound is at least 1/4. Using λ = r/2 < 1/2, (N − 1)(1 − λ) 1 ≥ 2(N − 1) − λ(N − 2) 4 is equivalent to 4(N − 1)(1 − λ) ≥ 2(N − 1) − λ(N − 2), 40

or 2(N − 1) ≥ λ(3N − 2). This holds for every N ≥ 2 because λ < 1/2 implies λ(3N − 2) <

3N − 2 ≤ 2(N − 1). 2

Therefore  Pr µ ∈ C + ∪ C − ≥

(N − 1)(2 − r) 1 ≥ , 4(N − 1) − r(N − 2) 4

as claimed.

D

Additional general-N computational simulations (Section 4.1)

The duopoly heatmaps in Section 4.1 are useful because, when N = 2, the exploration-mean space is two-dimensional and the full terminal-price map can be swept directly. For general N , the terminal-price map is a function on an N -dimensional hypercube, so an exhaustive visualization is no longer available. This appendix therefore uses the same price-moments ODE to study two clustered exploration designs for general N . These designs are not intended to characterize the full hypercube. Rather, they evaluate the ODE on broad and interpretable families of exploration profiles in which firms experiment around a common local price band or a common anchor price. Throughout, unless otherwise noted, we use the same primitives as in Section 4.1: a = b = 1 and c = 1/2, so that pNE = 2/3 and pMNP = 1. The simulations in this appendix restrict the exploration covariance to the scalar form 2 Σexp = σexp IN ,

where σexp is the common exploration standard deviation. For N > 2, when reporting a scalar terminal price, we use the cross-firm average N

P̄ ODE (α; µ, Σexp ) :=

1 X ODE Pi (α; µ, Σexp ). N i=1

Under the scalar covariance specification above, this is N

2 P̄ ODE (α; µ, σexp IN ) =

1 X ODE 2 Pi (α; µ, σexp IN ). N i=1

D.1

Interval sampling

Motivation. The first design is the general-N analogue of the “similar exploration prices” emphasized by the cone analysis. Instead of attempting to sample uniformly from the full N -dimensional

41

exploration-mean space, we draw firms’ exploration means from a common interval. This focuses attention on clustered profiles, where firms explore within the same local price band. Design. For each boundary pair (ℓ, u) with 0.40 ≤ ℓ < u ≤ 1.00, we fix two exploration means at the endpoints and draw the remaining firms inside the interval: µ1 = ℓ,

µ2 = u,

iid

µi | (ℓ, u) ∼ Unif[ℓ, u],

i = 3, . . . , N.

This is the computational analogue of the random-interval prior in Definition 3.3: two firms determine a local price band, and the remaining firms explore within that same band. For each sampled 2 I profile, we solve the price-moments ODE with Σexp = σexp N and record the mean terminal price 2 I ). The figure uses N = 10 and σ P̄ ODE (α; µ, σexp exp = 0.10. Each boundary pair averages over N

100 independent draws of the interior N − 2 exploration means. For visualization, we symmetrize the heatmap: cells with µ2 < µ1 are filled using the value from the corresponding flipped pair (µ2 , µ1 ).

1.00

N = 10,

exp = 0.10

1.00

0.90 0.83

Mean price

1

0.80 0.70

0.67

0.60

0.50

0.50 0.40 0.40 0.50 0.60 0.70 0.80 0.90 1.00

0.33

2

Figure 8: ODE-implied mean terminal price under interval sampling. For each boundary pair 0.40 ≤ ℓ < u ≤ 1.00 on a grid with increments of 0.01, two exploration means are fixed at ℓ and u, and the remaining N − 2 means are drawn uniformly from [ℓ, u]. Each boundary pair averages over 100 draws, with N = 10 and scalar exploration covariance Σexp = 0.102 IN . White is centered at the Nash price; red and blue indicate supra-competitive and sub-competitive mean prices, respectively. The heatmap is then symmetrized by assigning cells with µ2 < µ1 the value from the flipped pair.

Results. Figure 8 shows that the cone calculation is conservative. Across boundary pairs, 78.6% have mean terminal price above Nash. At the simulation level, 74.2% of draws have total profit above the Nash total-profit benchmark, and 70.2% of boundary pairs have average total profit 42

above that benchmark. The overall mean terminal price is 0.758, compared with pNE = 0.667, and the range of mean terminal prices is [0.652, 1.000]. Thus, within this family of clustered generalN exploration profiles, sub-competitive outcomes remain close to Nash, while supra-competitive outcomes can reach the monopoly benchmark.

D.2

Center–dispersion sampling

Motivation. The second design parameterizes clustered exploration profiles by a common anchor price and a dispersion level. This separates the location of the exploration cluster from the amount of cross-firm heterogeneity. Small dispersion approximates symmetric exploration, while larger dispersion tests how robust the qualitative pattern is when firms’ exploration means are no longer nearly identical. Design. For an anchor price s and dispersion parameter ν, draw iid

µi ∼ Unif[s −

3ν, s +

3ν],

i = 1, . . . , N.

Thus s controls the location of the exploration cluster and ν controls cross-firm heterogeneity. The √ factor 3 normalizes ν as the standard deviation of the cross-firm exploration means. We solve the price-moments ODE with N = 10 and compare ν ∈ {0.02, 0.10} and scalar exploration covariances 2 I with σ Σexp = σexp exp ∈ {0.02, 0.10}. N

N = 10, = 0.02

1.00

N = 10, = 0.10

Monopoly price

Monopoly price

Mean final price

0.95 0.90 0.85 0.80 exp = 0.02 exp = 0.10

0.75 0.70 0.65

Nash price

0.4

exp = 0.02 exp = 0.10

Nash price

0.6

0.8 1.0 Anchor price

1.2

0.4

0.6

0.8 1.0 Anchor price

1.2

Figure 9: Mean terminal price under center–dispersion sampling. For √ √ each anchor price s, exploration means are drawn independently from Unif[s − 3ν, s + 3ν]. The panels compare cross-firm dispersion levels ν ∈ {0.02, 0.10}, and the lines compare scalar exploration covariances 2 I Σexp = σexp N with σexp ∈ {0.02, 0.10}, using N = 10. Dashed horizontal and vertical lines mark the Nash and monopoly benchmarks. 43

Results. The pattern in Figure 9 mirrors the symmetric-exploration benchmark. Prices are closest to Nash when the anchor is near pNE , and they rise substantially when the anchor is below Nash or above the monopoly benchmark. Increasing either the cross-firm dispersion ν or the common exploration standard deviation σexp attenuates the supra-competitive effect, but the qualitative shape remains stable. Taken together, these general-N simulations are not an exhaustive statement about the full exploration-mean space. They instead evaluate the ODE on broad families of clustered profiles. Within these families, supra-competitive outcomes arise more broadly than the analytical cone certificate alone would imply.

E

Details for the Empirically Calibrated Simulations (Section 5)

This appendix collects deferred details for the empirically calibrated simulation design described in Section 5. The demand and observed rents are calibrated from data, while the pricing dynamics are simulated counterfactually on top of that calibrated environment.

E.1

Demand-side parameterization

Recall the random-coefficients logit demand system from Section 5:   exp αh Pj,t + x′j βh + ξj w̃h sj,t (Pt ; ξ) = . P ′ 1+ N k=1 exp αh Pk,t + xk βh + ξk h=1 H X

This demand system converts any rent vector into predicted market shares by aggregating heterogeneous household choice probabilities. The covariates xj are observed rental characteristics (e.g., bedrooms, new-building indicators, and quality proxies). The renter-specific coefficients (αh , βh ) are set as in Calder-Wang and Kim (2024): αh < 0 is a decreasing linear function of log income, and βh uses their estimated interaction structure (Table 13 of Calder-Wang and Kim, 2024). The vertical-differentiation fixed effects {ξj }N j=1 are calibrated so that at observed rents p0 , the model matches the observed shares s0 : sj (p0 ; ξ) = sj0

for each j = 1, . . . , N.

These fixed effects absorb product-specific vertical quality not captured by observed characteristics, so the calibrated model exactly matches baseline shares.

E.2

Calibration of the shadow costs

We calibrate λ = (λj )N j=1 so that the observed rent vector p0 is the Nash equilibrium of the calibrated static game. These shadow costs are not observed accounting costs; they are supply-side primitives

44

chosen to rationalize observed rents as Nash. Formally, firm j solves  max(pj − λj ) sj (pj , p−j,0 ); ξ , pj

taking rivals’ rents as fixed at p−j,0 . We set λj so that the first-order condition holds at pj0 : 0 = sj (p0 ; ξ) + pj0 − λj

 ∂sj (p0 ; ξ), ∂pj

j = 1, . . . , N.

Because the demand side is calibrated so that sj (p0 ; ξ) = sj0 , this pins down a unique λj for each rental: λj = pj0 +

E.3

sj0 . ∂sj /∂pj (p0 ; ξ)

Explore-then-exploit pipeline: full equations

We provide the full equations for the explore-then-exploit dynamics. Exploration. Fix an exploration mean µj for each rental j. During t = 1, . . . , K, rental j posts a perturbed price i.i.d.

vj,t ∼ N (0, σ 2 ),

Pj,t = µj (1 + vj,t ),

clipped to [0, P̄ ], where P̄ is a large exogenous upper bound. Exploitation. For t ≥ K, each rental j estimates a naive binary logit using only its own priceshare history {(Pj,τ , sj,τ )}τ ≤t , treating competitors’ prices as unobserved: yj,τ := log

  s j,τ ≈ ηj,t + θj,t Pj,τ , 1 − sj,τ

τ ≤ t.

Here “naive” means that the fitted logit is one-product and own-price only, even though true shares depend on all rents. This yields (η̂j,t , θ̂j,t ) and a predicted share ŝj,t (p) = Λ(η̂j,t + θ̂j,t p), where Λ(z) = (1 + e−z )−1 is the logistic function. The next-period rent is set by myopically maximizing predicted profit: Pj,t+1 ∈ arg max (p − λj )ŝj,t (p),

t = K, . . . , K + T − 1.

p∈[0,P̄ ]

F

Omitted Lemmas in Proof of Theorem 3.2 (Section 6)

This appendix provides the omitted technical details for the proof of Theorem 3.2. The proofs follow the three-step structure used in Section 6, with the correlation-bias calculation placed in Step 2. The final two subsections adapt the same argument to the upper cone C + . The main difference is only the sign of the initial gap: on C + , prices initially move downward relative to historical averages,

45

so the invariant orthant is P −U ⪯ 0 rather than P −U ⪰ 0. Since covariance accumulation depends on products ei ej , the correlation argument is otherwise unchanged.

F.1 Proof.

Proof of Lemma 6.1 Let e(t) := P (t) − U (t). The claim is that if e(t0 ) ⪰ 0 for some t0 ≥ 1, then e(t) ⪰ 0 for

all t ≥ t0 . Assume for contradiction that the trajectory leaves RN + after t0 , and define the first exit time τ := inf{t ≥ t0 : ∃i with ei (t) < 0}. By continuity of e(·), we have e(τ ) ⪰ 0 and there exists at least one index i with ei (τ ) = 0. Fix such an i. We rule out each possible way in which the ith coordinate could be the first one to cross into the negative orthant. Clipped and non-OLS branches. First consider the branches on which the posted price is locally fixed at a cap. If Pi (τ ) = Pmin and ei (τ ) = 0, then Ui (τ ) = Pmin . As long as the lower clip remains active, Ui solves U̇i (t) =

Pmin − Ui (t) , t

with initial value Ui (τ ) = Pmin , and hence Ui (t) = Pmin on that right-neighborhood. Thus the lower clip cannot by itself make ei negative. If the lower clip ceases to bind, the posted price moves weakly upward from Pmin = Ui (τ ), which is also not a first exit from RN +. If Pi (τ ) = Pmax and ei (τ ) = 0, then Ui (τ ) = Pmax . While the upper clip, or the non-OLS branch Pi ≡ Pmax , remains active, feasibility gives Ui (t) ≤ Pmax and hence ei (t) = Pmax − Ui (t) ≥ 0. Therefore this branch cannot itself generate a negative gap. The only remaining possibility is that the trajectory leaves the upper clipped branch and crosses through the algebraic OLS boundary Pei = Ui . That crossing is covered by the energy calculation below. OLS boundary. It remains to rule out a crossing through an OLS boundary point with Bi := bVii − cV̄i,−i > 0,

Pei (τ ) − Ui (τ ) = 0.

This includes the interior OLS case and the instant at which a clip ceases to bind. Define  Ei := 2bVii Ui − BR(Ū−i ) − cUi V̄i,−i . A direct rearrangement of the formula for Pei gives the exact identity −Ei . Pei − Ui = 2Bi

46

Since Bi > 0, crossing from Pei − Ui ≥ 0 to Pei − Ui < 0 is equivalent to crossing from Ei ≤ 0 to Ei > 0. At time τ we have Ei (τ ) = 0. We now compute Ėi (τ ) under the information e(τ ) ⪰ 0 and ei (τ ) = 0. First, e(τ ) ⪰ 0 implies P U̇ (τ ) = e (τ )/τ ≥ 0 for all j, so Ū˙ (τ ) = 1 U̇ (τ ) ≥ 0. Also e (τ ) = 0 implies U̇ (τ ) = 0,

−i i i j̸=i j N −1 ˙ ⊤ 2 and from V̇ = (P − U )(P − U ) we get V̇ii (τ ) = ei (τ ) = 0 and V̄i,−i (τ ) = ei (τ )ē−i (τ ) = 0. j

j

Differentiating Ei along the ODE gives    Ėi = 2b V̇ii Ui − BR(Ū−i ) + 2bVii U̇i − BR′ (Ū−i ) Ū˙ −i − c U̇i V̄i,−i − cUi V̄˙ i,−i . Evaluating at τ and using V̇ii (τ ) = U̇i (τ ) = V̄˙ i,−i (τ ) = 0 gives Ėi (τ ) = −2bVii (τ ) BR′ (Ū−i (τ )) Ū˙ −i (τ ) = −c Vii (τ ) Ū˙ −i (τ ) ≤ 0. Therefore, at a boundary point where Ei (τ ) = 0, the derivative points toward {Ei ≤ 0} and cannot cross into {Ei > 0}. Equivalently, Pei − Ui cannot become negative immediately after τ , hence ei cannot become negative immediately after τ . Conclusion. In all cases, ei cannot be the first coordinate to turn negative, contradicting the definition of τ . Hence e(t) ⪰ 0 for all t ≥ t0 .

F.2 Proof.

Proof of Lemma 6.2 Write e(t) := P (t) − U (t). For i ̸= j, the ODE gives V̇ij (t) = ei (t)ej (t). Under e(t) ⪰ 0,

we have V̇ij (t) ≥ 0 for all t, so Vij is nondecreasing and Vij (t) ≥ Vij (1) = 0. Now assume e(t0 ) ≻ 0. Fix any t > t0 . The map (U, V ) 7→ P (U, V ) is continuous (it is a rational map on the OLS region, composed with coordinatewise clipping, and equals the constant Pmax on the non-OLS region), and (U, V ) is continuous in time because it solves an ODE with locally bounded right-hand side. Hence e(·) is continuous. Since e(t0 ) ≻ 0, there exist δt ∈ (0, t − t0 ] and εt > 0 such that ei (s) ≥ εt and ej (s) ≥ εt for all s ∈ [t0 , t0 + δt ]. Therefore Z t

Z t0 +δt ei (s)ej (s) ds ≥

Vij (t) = Vij (1) + 1

ε2t ds = ε2t δt > 0.

t0

Since t > t0 was arbitrary, Vij (t) > 0 for every i ̸= j and every t > t0 . Averaging over j ̸= i gives P V̄i,−i (t) = N 1−1 j̸=i Vij (t) > 0 for all i and all t > t0 .

F.3 Proof.

Proof of Lemma 6.3 Fix i ∈ [N ] and hold U and Vii fixed. Write x := V̄i,−i and A := a + c Ū−i . Using the

algebraic OLS price from Definition 3.1, write A Vii − c Ui x Pei (x) := , 2(bVii − cx)

on the domain bVii − cx > 0. 47

A Vii A = 2b = BR(Ū−i ). Part (1). If x = 0, then Pei (0) = 2bV ii Part (2). Differentiate Pei with respect to x using the quotient rule. Let N (x) := AVii − cUi x

and D(x) := 2(bVii − cx). Then N ′ (x) = −cUi and D′ (x) = −2c, so  (−cUi ) 2(bVii − cx) − AVii − cUi x (−2c) N ′ (x)D(x) − N (x)D′ (x) ′ e Pi (x) = = . D(x)2 4(bVii − cx)2 Expanding the numerator and canceling the x-terms gives (−2cUi bVii + 2cUi cx) + (2cAVii − 2c2 Ui x) = 2cVii (A − bUi ). Therefore,

2cVii (A − bUi ) . Pei′ (x) = 4(bVii − cx)2

On the domain bVii − cx > 0, the denominator is strictly positive. Also Vii > 0 because Vii (1) = (Σexp )ii > 0 and V̇ii = (Pi − Ui )2 ≥ 0. Finally, A − bUi = a + cŪ−i − bUi ≥ a − bPmax + cPmin > 0, because Ui , Ū−i ∈ [Pmin , Pmax ], Pmin > 0, and Pmax ≤ a/b. Hence Pei′ (x) > 0 and Pei is strictly increasing in V̄i,−i .

F.4 Proof.

Proof of Lemma 6.4 Assume (i) V̄i,−i (t) ≥ 0 for all i and t ≥ 1, and (ii) each coordinate Ui (t) is monotone in t.

Since Ui (t) ∈ [Pmin , Pmax ] for all t (because U is a running average of prices in [Pmin , Pmax ]), every monotone Ui has a finite limit; write Ui∞ := limt→∞ Ui (t) and U ∞ := (Ui∞ )N i=1 . Step 1: show Pi (t) ≥ BR(Ū−i (t)) for all i, t. Fix i and t. If the non-OLS branch is active, then Pi (t) = Pmax . Under the standing feasibility condition pNE ≤ Pmax (equivalently pNE ∈ [Pmin , Pmax ]) and monotonicity of BR, we have BR(Ū−i (t)) ≤ BR(Pmax ) ≤ Pmax = Pi (t). If instead the OLS branch is active, then bVii − cV̄i,−i > 0 and Lemma 6.3 gives Pei (U, V ) ≥ BR(Ū−i ) because V̄i,−i ≥ 0 and equality holds at V̄i,−i = 0. Clipping cannot reduce the value below BR(Ū−i ) since BR(Ū−i ) ≤ Pmax and Pi = [Pei ][P ,P ] (if BR(Ū−i ) < Pmin then Pi ≥ Pmin > BR(Ū−i )). Hence min

max

Pi (t) ≥ BR(Ū−i (t)) in all cases. Step 2: rule out mini Ui∞ < pNE . Let m∞ := mini Ui∞ and choose an index k with Uk∞ = m∞ . Suppose for contradiction that m∞ < pNE . Define the continuous function ψ(u, v) := BR(u)−v. We have ψ(m∞ , m∞ ) = BR(m∞ ) − m∞ > 0 because BR(u) − u is affine with unique zero at u = pNE . Therefore we can pick ε > 0 small enough that η := BR(m∞ − ε) − (m∞ + ε) > 0.

48

∞ ≥ m∞ , there exists T such that for all t ≥ T , U (t) ≤ m∞ +ε Since Uk (t) → m∞ and Ū−k (t) → Ū−k k

and Ū−k (t) ≥ m∞ − ε. Using Step 1 and monotonicity of BR, ek (t) := Pk (t) − Uk (t) ≥ BR(Ū−k (t)) − Uk (t) ≥ BR(m∞ − ε) − (m∞ + ε) = η,

t ≥ T.

Then U̇k (t) = ek (t)/t ≥ η/t for all t ≥ T , so integrating gives Uk (t) ≥ Uk (T ) + η log(t/T ), which diverges as t → ∞, contradicting convergence of Uk (t). Hence m∞ ≥ pNE , i.e. U ∞ ⪰ pNE 1.

F.5 Proof.

Proof of Lemma 6.5 Assume U ∞ ⪰ pNE 1 and Uj∞ = pNE for some j. In addition, in the C − analysis we have

V̄j,−j (t) ≥ 0 for all t ≥ 1 (indeed V̄j,−j (t) > 0 for t > 1 by Lemma 6.2 with t0 = 1). Suppose for contradiction that U ∞ ̸= pNE 1. Then there exists k ̸= j with Uk∞ > pNE , so ∞ > pNE . Because BR is strictly increasing and BR(pNE ) = pNE , we have BR(Ū ∞ ) > pNE ; define Ū−j −j  1 ∞ NE η := 2 BR(Ū−j ) − p > 0. ∞ and U (t) → pNE , for all sufficiently large t we have BR(Ū (t)) ≥ By convergence Ū−j (t) → Ū−j j −j

pNE + η and Uj (t) ≤ pNE + η/2. Also V̄j,−j (t) ≥ 0 implies Pj (t) ≥ BR(Ū−j (t)) (same argument as in the proof of Lemma 6.4). Hence for all sufficiently large t, ej (t) := Pj (t) − Uj (t) ≥ (pNE + η) − (pNE + η/2) = η/2. Therefore U̇j (t) = ej (t)/t ≥ (η/2)/t for all large t, which implies Uj (t) cannot converge to pNE . This contradiction shows that no such k exists, hence U ∞ = pNE 1.

F.6 Proof.

Proof of Lemma 6.6 Assume the hypotheses of Lemma 6.1, so e(t) := P (t) − U (t) ⪰ 0 for all t ≥ t0 . Fix t0 ≥ 1

and assume Ṗ (t0 ) ⪰ 0. Because P is defined by a piecewise-smooth map of (U, V ) (rational on the OLS region, constant on the non-OLS region, and coordinatewise clipped), it suffices to prove that on any open time interval on which (i) every firm is on the OLS branch (Bi > 0) and (ii) clipping is inactive (Pi = Pei ∈ (Pmin , Pmax )), the condition Ṗ ⪰ 0 is forward invariant. Once proved, crossing into the clipped or non-OLS regimes cannot create a negative Ṗi (in those regimes Pi is locally constant), so the result extends to all t ≥ t0 . Step 1: compute Ṗi and isolate its sign. Work on an interval where Pi = Pei ∈ (Pmin , Pmax ) and Bi := bVii − c V̄i,−i > 0 for every i. Write Ai := a + c Ū−i and recall Ai Vii − cUi V̄i,−i Pi = Pei = . 2Bi

49

Define also the (strictly positive) conditional expected demand at (Ui , Ū−i ): Di := Ai − bUi = a + c Ū−i − bUi > 0, which holds because Ui , Ū−i ∈ [Pmin , Pmax ] and a − bPmax + cPmin > 0. Differentiate Pi along the ODE. First, the reduced ODE gives U̇i = ei /t and Ū˙ −i = ē−i /t, where P ē−i := N 1−1 j̸=i ej . Also V̇ii = e2i and V̄˙ i,−i = ei ē−i because V̇ = ee⊤ . i −Ni Ḃi . Compute Let Ni := Ai Vii − cUi V̄i,−i , so Pi = Ni /(2Bi ). Then Ṗi = Ṅi B2B 2 i

c Ȧi = c Ū˙ −i = ē−i , t and

Ḃi = b V̇ii − c V̄˙ i,−i = be2i − cei ē−i = ei (bei − cē−i ),

c c Ṅi = Ȧi Vii + Ai V̇ii − c U̇i V̄i,−i − cUi V̄˙ i,−i = Vii ē−i + Ai e2i − V̄i,−i ei − cUi ei ē−i . t t

Substituting into Ṗi and grouping terms yields the factorization Ṗi =

c ℓi k i , 2t Bi2

ℓi := Bi + t ei Di ,

ki := Vii ē−i − V̄i,−i ei .

(One can verify this by expanding Ṅi Bi −Ni Ḃi , then using Bi = bVii −cV̄i,−i to rewrite the bracketed expression as cDi ki ; the remaining (c/t)Bi ki term comes from the Ȧi contribution.) Under ei ≥ 0, Bi > 0, and Di > 0, we have ℓi > 0. Hence on this interior OLS interval, sign(Ṗi ) = sign(ki ) for each i. Step 2: derive the dynamics of k and show k ⪰ 0 is forward invariant. Differentiate ki = Vii ē−i − V̄i,−i ei : k̇i = V̇ii ē−i + Vii ē˙ −i − V̄˙ i,−i ei − V̄i,−i ėi . Using V̇ii = e2i and V̄˙ i,−i = ei ē−i , the first and third terms cancel: V̇ii ē−i − V̄˙ i,−i ei = e2i ē−i −ei ē−i ei = 0. Thus k̇i = Vii ē˙ −i − V̄i,−i ėi . 1 P Now ėi = Ṗi − U̇i = Ṗi − eti and ē˙ −i = Ṗ −i − U̇ −i = Ṗ −i − ē−i j̸=i Ṗj . t , where Ṗ −i := N −1 Therefore k̇i = Vii Ṗ −i − V̄i,−i Ṗi −

 1 1 Vii ē−i − V̄i,−i ei = Vii Ṗ −i − V̄i,−i Ṗi − ki . t t

c Substituting Ṗj = 2tB 2 ℓj kj gives a linear system k̇ = A(t)k with Metzler (cooperative) off-diagonal j

c ii structure: for j ̸= i, the coefficient multiplying kj in k̇i is NV−1 ℓ ≥ 0. Hence the cone {k ⪰ 0} 2tB 2 j j

is forward invariant on this interval: if ki hits 0 while all kj ≥ 0, then k̇i ≥ 0. Step 3: conclude Ṗ ⪰ 0 persists. At time t0 we assume Ṗ (t0 ) ⪰ 0. On the interior OLS interval, ℓi (t0 ) > 0 so Ṗi (t0 ) ≥ 0 implies ki (t0 ) ≥ 0 for each i. By forward invariance, k(t) ⪰ 0 for later times on the interval, and then Step 1 gives Ṗ (t) ⪰ 0. 50

Finally, if at some later time the trajectory enters a regime where firm i is clipped (so Pi is locally constant) or enters the non-OLS branch (Pi ≡ Pmax ), then Ṗi ≥ 0 holds trivially there. Thus Ṗ ⪰ 0 holds for all t ≥ t0 .

F.7

Proof of Lemma 6.7

Proof. Let ei := Pi −Ui and ē−i := (N −1)−1

j̸=i ej . Use the notation from the proof of Lemma 6.6:

P

Bi := bVii − cV̄i,−i , Di := a + cŪ−i − bUi , ℓi := Bi + tei Di , and ki := Vii ē−i − V̄i,−i ei . At t = 1, U (1) = µ, V (1) = Σexp , and Σexp is diagonal, so Vii (1) = (Σexp )ii and V̄i,−i (1) = 0. Hence Bi (1) = b(Σexp )ii > 0 and the OLS branch is active. Also Pei (1) = BR(µ̄−i ). The clipping is inactive: the lower clip is inactive because BR(µ̄−i ) > µi ≥ Pmin , while the upper clip is inactive because BR(µ̄−i ) ≤ BR(Pmax ) < Pmax under pNE < Pmax . Hence Pi (1) = BR(µ̄−i ) and ei (1) = BR(µ̄−i ) − µi > 0 for every i. Thus ē−i (1) > 0 and ki (1) = (Σexp )ii ē−i (1) > 0. The factorization in Lemma 6.6 gives Ṗi = c ℓk 2tBi2 i i

on the smooth OLS interior branch. At t = 1, we have Bi (1) > 0, ki (1) > 0, and

ℓi (1) = Bi (1) + ei (1)Di (1) > 0, since Di (1) = a + cµ̄−i − bµi ≥ a + cPmin − bPmax > 0 by the standing price-bound assumptions. Therefore Ṗi (1+ ) > 0 for every i, so Ṗ (1+ ) ≻ 0.

F.8 Proof.

Proof of Lemma 6.8 Assume U (t) → pNE 1 and that for all t we have e(t) := P (t) − U (t) ⪰ 0 and Ṗ (t) ⪰ 0.

(1) Show P → pNE 1 and eventually P ⪯ pNE 1. Fix i. Since Ṗi ⪰ 0 and Pi (t) ∈ [Pmin , Pmax ], i implies the running-average identity the limit Pi∞ := limt→∞ Pi (t) exists. The ODE U̇i = Pi −U t

1 Ui (t) = Ui (1) + t

Z t

 Pi (s) ds .

1

Rt

∞ 1 Pi (s) ds → Pi , hence Ui (t) → Pi∞ as well. Since we also assume Ui (t) → pNE , it follows that Pi∞ = pNE for every i, i.e. P (t) → pNE 1.

If Pi (s) → Pi∞ , then the Cesàro mean converges to the same limit:

1 t

Because Ṗi ≥ 0 and Pi (t) → pNE , we must have Pi (t) ≤ pNE for all sufficiently large t (otherwise a nondecreasing function could not converge to a smaller limit). Thus P ⪯ pNE 1 eventually. (2) Show convergence of V and strict positivity of the limiting cross terms. For large t we have P (t) ⪯ pNE 1, so the running average also satisfies U (t) ⪯ pNE 1 and hence gi (t) := pNE −Ui (t) ≥ 0 for large t. Define also δi (t) := pNE − Pi (t) ≥ 0 for large t. Then ei = Pi − Ui = gi − δi ≤ gi . On the C − side (the context where this lemma is used), we also have V̄i,−i (t) ≥ 0 and therefore (by Lemma 6.3) Pi (t) ≥ BR(Ū−i (t)) on the OLS branch; on the non-OLS branch Pi = Pmax ≥ BR(Ū−i ) as well. For large t we already know Pi (t) ≤ pNE < Pmax , so we are eventually on the OLS branch

51

and clipping is inactive. Hence for large t we may use Pi ≥ BR(Ū−i ), which implies δi (t) = pNE − Pi (t) ≤ pNE − BR(Ū−i (t)) =

 c NE c p − Ū−i (t) = ḡ−i (t), 2b 2b

P where ḡ−i := N 1−1 j̸=i gj . P Let G(t) := N i=1 gi (t). Using U̇i = ei /t gives ġi = −(gi − δi )/t, so Ġ(t) = −

  X 1X 1 1 c X (gi − δi ) ≤ − G(t) − δi (t) ≤ − G(t) − ḡ−i (t) . t t t 2b i

But

i

i

i ḡ−i = G (each gj appears in exactly N − 1 averages and the factor 1/(N − 1) cancels), hence

P

Ġ(t) ≤ −

1 c 1− G(t). t 2b

c

c > 21 , we have Integrating yields G(t) ≤ C t−(1− 2b ) for some C > 0. Since c < b implies 1 − 2b

G ∈ L2 ([1, ∞)) and in particular ei (t) ≤ gi (t) ≤ G(t) implies ei ∈ L2 ([1, ∞)). Rt Now Vii (t) = Vii (1)+ 1 ei (s)2 ds converges to a finite limit Vii∞ because e2i is integrable. Similarly, Rt for i ̸= j, Vij (t) = Vij (1) + 1 ei (s)ej (s) ds converges because ei ej is integrable by Cauchy–Schwarz. Finally, to get strict positivity of the limiting cross terms, it suffices that there exists t0 such that e(t0 ) ≻ 0 (which holds on C − at t0 = 1 because V (1) = Σexp is diagonal). By continuity, R∞ ei (s)ej (s) > 0 on some interval of positive length, so 1 ei (s)ej (s) ds > 0 and therefore Vij∞ > 0 for ∞ > 0 for all i. all i ̸= j; averaging gives V̄i,−i

F.9 Proof.

Proof of Lemma 6.9 ∞ > 0. Suppose Assume we are on the C − side and that Vii (t) → Vii∞ and V̄i,−i (t) → V̄i,−i

for contradiction that U (t) → pNE 1. i As in Lemma 6.8, U̇i = Pi −U and convergence of U imply convergence of P and P (t) → pNE 1. t

In particular, for all sufficiently large t we have Pi (t) ∈ (Pmin , Pmax ) because pNE ∈ (Pmin , Pmax ) and Pi (t) → pNE . For large t, the non-OLS branch is impossible: if −bVii (t) + c V̄i,−i (t) ≥ 0 then by definition Pi (t) = Pmax , which contradicts Pi (t) → pNE < Pmax . Thus for all sufficiently large t we are on the OLS branch with Bi (t) := bVii (t) − c V̄i,−i (t) > 0. Also clipping is inactive for large t (since Pi (t) ∈ (Pmin , Pmax )), so Pi (t) = Pei (U (t), V (t)) eventually. Taking t → ∞ in the formula for Pei and using U (t) → pNE 1, V (t) → V ∞ gives pNE = Pi∞ = lim Pei (U (t), V (t)) = Pei (pNE 1, V ∞ ). t→∞

∞ > 0 and B ∞ > 0 allow us to apply Lemma 6.3(2) at U = pNE 1, yielding But V̄i,−i i

Pei (pNE 1, V ∞ ) > BR(pNE ) = pNE ,

52

a contradiction. Therefore U (t) ̸→ pNE 1, and combined with Lemma 6.4 we conclude U ∞ ≻ pNE 1.

F.10

Adaptation of Step 1 (Section 6.1) to C +

Step 1 on C − shows that P (1)−U (1) ≻ 0 and then uses a forward-invariance argument to keep P −U + in RN + . On C the only change is the sign: P (1) − U (1) ≺ 0, and we need the analogous forward N invariance for RN − rather than R+ . Once this is established, U is componentwise nonincreasing; the

covariance and correlation-bias argument is handled in Appendix F.11. F.10.1

Initial gap on C + : prices jump downward but remain synchronized

Since V (1) = Σexp is diagonal, V̄i,−i (1) = 0. Then the zero-cross-covariance case of the pricing rule in Definition 3.1 reduces to the best-response price at t = 1, and the cap is inactive under the standing price bounds. Hence Pi (1) = BR(µ̄−i ). Since µ ∈ C + means µi > BR(µ̄−i ) for every i, we obtain P (1) − U (1) = P (1) − µ ≺ 0. Thus all firms’ initial updates move in the same direction (down), which is the C + analogue of the synchronization used on C − . F.10.2

Forward invariance of the nonpositive orthant

− proof uses P − U ⪰ 0. The In the main text we only stated Lemma 6.1 for RN + , since the C

sign-reversed statement is the following. Claim (negative-orthant invariance). Fix t0 ≥ 1. If P (t0 ) − U (t0 ) ⪯ 0, then P (t) − U (t) ⪯ 0 for all t ≥ t0 . Let e := P − U and suppose, toward a contradiction, that the trajectory leaves RN − after t0 . Let τ := inf{t ≥ t0 : ∃i with ei (t) > 0}. At t = τ , we have e(τ ) ⪯ 0 and at least one coordinate satisfies ei (τ ) = 0. Fix such an i. The upper-cap and non-OLS cases require some care. If Pi (τ ) = Pmax and ei (τ ) = 0, then Ui (τ ) = Pmax . As long as the upper clip, or the non-OLS branch Pi ≡ Pmax , remains active, Ui solves U̇i (t) =

Pmax − Ui (t) t

with initial value Ui (τ ) = Pmax , so Ui (t) = Pmax and ei (t) = 0 on that right-neighborhood. If the upper cap ceases to bind, the posted price moves weakly downward from Pmax = Ui (τ ), which points into RN − , not out of it. Hence the upper cap/non-OLS branch cannot be the source of a positive first exit. If Pi (τ ) = Pmin and ei (τ ) = 0, then Ui (τ ) = Pmin . While the lower clip remains active, the same scalar argument gives Ui (t) = Pmin and ei (t) = 0. If the lower clip ceases to bind upward, any possible exit must pass through the algebraic OLS boundary Pei = Ui . Thus it remains only to rule

53

out an OLS-boundary crossing with Bi := bVii − cV̄i,−i > 0,

Pei (τ ) − Ui (τ ) = 0.

Use the same energy as in Lemma 6.1,  Ei := 2bVii Ui − BR(Ū−i ) − cUi V̄i,−i , for which

−Ei . Pei − Ui = 2Bi

Because Bi > 0, crossing from Pei − Ui ≤ 0 to Pei − Ui > 0 is equivalent to crossing from Ei ≥ 0 to Ei < 0. At the boundary time, Ei (τ ) = 0. Now e(τ ) ⪯ 0 implies U̇j (τ ) = ej (τ )/τ ≤ 0 for all j, so Ū˙ −i (τ ) ≤ 0. Also ei (τ ) = 0 implies U̇ (τ ) = 0, V̇ (τ ) = 0, and V̄˙ (τ ) = 0. The same derivative computation as in the proof of i

ii

i,−i

Lemma 6.1 gives Ėi (τ ) = −cVii (τ )Ū˙ −i (τ ) ≥ 0. Thus Ei cannot cross from ≥ 0 to < 0, equivalently Pei − Ui cannot cross from ≤ 0 to > 0. This contradicts the definition of τ . Therefore P − U ⪯ 0 is forward invariant on C + . As an immediate corollary, U̇ = (P − U )/t ⪯ 0, so U is componentwise nonincreasing on C + .

F.11

Adaptation of Step 2 (Section 6.2) to C +

On C − we used Step 2 to obtain a limit lower bound U ∞ ⪰ pNE 1, and Step 3 was needed to exclude convergence to the Nash boundary. On C + the direction of the initial price update is reversed (P (1) − U (1) ≺ 0), but the correlation mechanism is even simpler: the single-orthant gap from Appendix F.10 implies V̄i,−i (t) > 0 for all t > 1, and then Lemma 6.3 yields a strict pointwise lower bound Pi (t) > pNE for all t > 1. Since U is a running average of P and µ ∈ C + already implies µ ≻ pNE 1, we conclude U (t) ≻ pNE 1 for all t ≥ 1; combined with the pointwise bound on P (t), this gives P ODE (α; µ, Σexp ) ≻ pNE 1 for every α ∈ [1, ∞). Thus Step 3 is unnecessary on C + . F.11.1

A basic implication of C + : exploration means exceed Nash

We record a simple fact used repeatedly below: if µ ∈ C + , then µ ≻ pNE 1. Let m := mini µi and choose k with µk = m. Since µ̄−k ≥ m and BR is increasing, m = µk > BR(µ̄−k ) ≥ BR(m). But BR(u) − u is affine with unique zero at u = pNE , hence m > BR(m) implies m > pNE , and therefore µi ≥ m > pNE for all i. At t = 1 we also have V̄i,−i (1) = 0 because V (1) = Σexp is diagonal, so the zero-cross-covariance case of the pricing rule in Definition 3.1 gives Pi (1) = BR(µ̄−i ), with the cap inactive under the 54

standing price bounds. Since µ̄−i > pNE and BR is increasing with BR(pNE ) = pNE , this implies Pi (1) > BR(pNE ) = pNE for every i. F.11.2

Same-sign gaps imply positive cross-covariances, strictly for t > 1

With e(t) ⪯ 0 for all t ≥ 1, we have for i ̸= j, V̇ij (t) = ei (t)ej (t) ≥ 0, so each off-diagonal entry is nondecreasing and, since V (1) = Σexp is diagonal, satisfies Vij (t) ≥ Vij (1) = 0. To obtain strict positivity, fix any t > 1. Since e(1) ≺ 0 and e(·) is continuous, there exist δt ∈ (0, t − 1] and εt > 0 such that ei (s) ≤ −εt and ej (s) ≤ −εt for all s ∈ [1, 1 + δt ]. Hence Z t

Z 1+δt ei (s)ej (s) ds ≥

Vij (t) =

ε2t ds = ε2t δt > 0.

1

1

Averaging over j ̸= i yields V̄i,−i (t) > 0 for all i and all t > 1. This is exactly the correlation persistence property needed in the C + adaptation of Step 2. Remark (relation to Lemma 6.2). The proof of Lemma 6.2 uses only that V̇ij = ei ej ≥ 0 and that e is strictly nonzero at some time to make the integral strictly positive. These facts hold equally when e ⪯ 0, so one may view the C + argument above as the sign-reversed version of that lemma. F.11.3

Strict supra-competitive prices for t > 1 from correlation bias

By Appendix F.10, along the C + trajectory we have P (t) − U (t) ⪯ 0 for all t ≥ 1. By the previous subsection, the resulting covariance accumulation satisfies V̄i,−i (t) > 0 for all i and all t > 1. Fix t > 1 and a firm i. If the non-OLS branch is active, then Pi (t) = Pmax , which is strictly larger than pNE under the mild feasibility requirement pNE < Pmax .5 If instead the OLS branch is active, then bVii − cV̄i,−i > 0 and Lemma 6.3(2) yields since V̄i,−i (t) > 0.

Pei (t) > BR(Ū−i (t))

Now we only need the weak inequality BR(Ū−i (t)) ≥ pNE . This follows because P − U ⪯ 0 implies U̇ = (P − U )/t ⪯ 0, so U is componentwise nonincreasing and thus U (t) ⪰ U ∞ . Applying Lemma 6.4 (whose hypotheses are satisfied on C + because the previous subsection gives V̄i,−i ≥ 0 and Appendix F.10 gives U monotone) yields U ∞ ⪰ pNE 1, hence U (t) ⪰ pNE 1 and therefore Ū−i (t) ≥ pNE . Since BR is increasing and BR(pNE ) = pNE , we conclude BR(Ū−i (t)) ≥ pNE . Combining the last two displays gives Pei (t) > pNE . Finally, clipping cannot reduce Pi (t) = [Pei (t)][Pmin ,Pmax ] below pNE : if Pei (t) ≤ Pmax then Pi (t) = Pei (t) > pNE ; if Pei (t) > Pmax then 5

If one only assumes pNE ≤ Pmax , then the non-OLS branch gives Pi (t) ≥ pNE , and the strictness below still follows from the OLS branch for t > 1 (which is the relevant regime near the Nash price).

55

Pi (t) = Pmax > pNE . Thus in all regimes we have for every i and every t > 1.

Pi (t) > pNE F.11.4

Converting pointwise supra-competitive prices into P ODE (α; µ, Σexp ) ≻ pNE 1

Since U is the running average of the realized prices (cf. U̇ = (P − U )/t), we may write for each i and t ≥ 1, 1 Ui (t) = Ui (1) + t

Z t

 Pi (s) ds .

1

We already established Ui (1) = µi > pNE , and we have shown Pi (s) > pNE for all s > 1. Therefore Ui (t) is a convex combination of numbers strictly larger than pNE , hence Ui (t) > pNE for all t ≥ 1. Together with the initial bound Pi (1) > pNE and the pointwise bound Pi (t) > pNE for all t > 1, this yields P ODE (α; µ, Σexp ) ≻ pNE 1 for every α ∈ [1, ∞), proving Theorem 3.2(a) and explaining why Step 3 is not needed on C + .

G

Omitted Lemmas in Proof of Theorem 3.1 (Appendix A)

G.1

Proof of Lemma A.1

Proof.

For each firm i, the OLS regression uses only the four own-firm empirical moments t

P̄i,t :=

1X Pi,s , t

t

Q̄i,t :=

s=1

1X Qi,s , t

t

P 2 i,t :=

s=1

1X 2 Pi,s , t

t

P Qi,t :=

s=1

In the full moment state X̄t = (P̄t , Q̄t , P P t , P Qt ), these are the local coordinates   loc X̄i,t := P̄i,t , Q̄i,t , P P ii,t , P Qi,t = P̄i,t , Q̄i,t , P 2 i,t , P Qi,t . Define the clipping operator [x][Pmin ,Pmax ] := min{max{x, Pmin }, Pmax }. We first show that, for every exploitation period t ≥ K, loc Pi,t+1 = fP (X̄i,t ),

56

1X Pi,s Qi,s . t s=1

where

# " P̄i,t P Qi,t − P 2 i,t Q̄i,t     , if P Qi,t − P̄i,t Q̄i,t < 0, loc 2 P Qi,t − P̄i,t Q̄i,t fP (X̄i,t )= [Pmin ,Pmax ]    Pmax , otherwise.

Equivalently, in the notation of Appendix A, loc πi (X̄t ) = fP (X̄i,t ),

Pt+1 = π(X̄t ).

Fix a firm i ∈ [N ] and a time t ≥ K. Define 2 Ri,t := P 2 i,t − P̄i,t .

Si,t := P Qi,t − P̄i,t Q̄i,t ,

By nondegenerate exploration, Ri,t > 0 a.s. for all t ≥ K, so the OLS slope is well-defined. Since Qi,s > 0 a.s. for all s ≤ t, we also have Q̄i,t > 0 a.s. The OLS normal equations, with an intercept, for regressing Qi,s on (1, Pi,s ) over s ≤ t are α̂i,t P̄i,t + β̂i,t P 2 i,t = P Qi,t .

α̂i,t + β̂i,t P̄i,t = Q̄i,t ,

Subtracting P̄i,t times the first equation from the second gives β̂i,t Ri,t = Si,t , i.e. β̂i,t =

Si,t , Ri,t

α̂i,t = Q̄i,t − β̂i,t P̄i,t .

Let π̂i,t (p) := p(α̂i,t + β̂i,t p) for p ∈ [Pmin , Pmax ]. By the exploitation rule, Pi,t+1 = arg

max p∈[Pmin ,Pmax ]

π̂i,t (p).

Case 1: Si,t < 0 (equivalently β̂i,t < 0). Then π̂i,t is strictly concave, so its unique unconstrained maximizer is p⋆ = −α̂i,t /(2β̂i,t ), and the constrained maximizer is Pi,t+1 = [p⋆ ][Pmin ,Pmax ] . Using α̂i,t = Q̄i,t − β̂i,t P̄i,t and β̂i,t Ri,t = Si,t , p⋆ = −

Q̄i,t − β̂i,t P̄i,t 2β̂i,t

=

P̄i,t P Qi,t − P 2 i,t Q̄i,t P̄i,t Si,t − Q̄i,t Ri,t  . = 2Si,t 2 P Qi,t − P̄i,t Q̄i,t

This is the first branch of the claimed formula. Case 2: Si,t ≥ 0 (equivalently β̂i,t ≥ 0). Then π̂i,t is convex, or linear, on [Pmin , Pmax ], so maximizers lie at the endpoints. Moreover,  π̂i,t (Pmax ) − π̂i,t (Pmin ) = (Pmax − Pmin ) α̂i,t + β̂i,t (Pmax + Pmin )   = (Pmax − Pmin ) Q̄i,t + β̂i,t (Pmax + Pmin − P̄i,t ) > 0,

57

since Pmax > Pmin , Q̄i,t > 0, β̂i,t ≥ 0, and P̄i,t ≤ Pmax implies Pmax + Pmin − P̄i,t ≥ Pmin > 0. Hence the unique maximizer is Pmax , giving the second branch. Combining the two cases proves the stated deterministic mapping from the local empirical moments to Pi,t+1 , and hence proves Pt+1 = π(X̄t ). The continuity and Lipschitz statement used in the finite-horizon ODE-tracking argument follows from Lemmas G.3 and G.4 below. Those lemmas show that, on the fixed tube around the mean-field trajectory, each local block PQ PP (xPi , xQ i , xii , xi )

lies in a domain where the OLS denominator is bounded away from zero and the pricing map is Lipschitz. Therefore π is continuous and Lipschitz on the finite-horizon neighborhoods used in Lemma A.4.

G.2 Proof.

Proof of Lemma A.2 Recall that

t

Xt = (Pt , Qt , Pt Pt⊤ , Pt ⊙ Qt ),

X̄t =

1X Xs . t s=1

Let Ft := σ (Pj,s , Qj,s )j∈[N ], 1≤s≤t



be the history filtration. Fix t ≥ K. By Lemma A.1, Pt+1 = π(X̄t ), so Pt+1 is Ft -measurable. Write p := π(X̄t ). Then, by the demand model and E[εi,t+1 | Ft ] = 0, E[Qi,t+1 | Ft ] = a − bpi +

c X pj = di (p). N −1 j̸=i

Since Pt+1 is Ft -measurable, we also have E[Pt+1 | Ft ] = p,

⊤ E[Pt+1 Pt+1 | Ft ] = pp⊤ ,

and, for each i, E[Pi,t+1 Qi,t+1 | Ft ] = pi E[Qi,t+1 | Ft ] = pi di (p). Therefore  E[Xt+1 | Ft ] = π(X̄t ), d(π(X̄t )), π(X̄t )π(X̄t )⊤ , π(X̄t ) ⊙ d(π(X̄t )) = f (X̄t ), which proves the claim.

58

G.3 Proof.

Proof of Lemma A.3 (SA recursion and start condition) Recall

t

Xt := (Pt , Qt , Pt Pt⊤ , Pt ⊙ Qt ),

X̄t :=

1X Xs . t s=1

Let Ft := σ (Pj,s , Qj,s )j∈[N ], 1≤s≤t



be the history filtration. Step 1: Running-mean identity and SA decomposition. For every t ≥ 1, X̄t+1 =

t+1  1 X t 1 1 Xs = X̄t + Xt+1 = X̄t + Xt+1 − X̄t . t+1 t+1 t+1 t+1 s=1

For t ≥ K, define ξt+1 := Xt+1 − f (X̄t ), Then X̄t+1 = X̄t +

h(X̄t ) := f (X̄t ) − X̄t .

 1 h(X̄t ) + ξt+1 , t+1

t ≥ K.

By Lemma A.2, E[ξt+1 | Ft ] = E[Xt+1 | Ft ] − f (X̄t ) = 0. Step 2: Componentwise form of the noise. Fix t ≥ K. By Lemma A.1, Pt+1 = π(X̄t ), hence Pt+1 is Ft -measurable and satisfies Pmin ≤ Pi,t+1 ≤ Pmax for every i. By the demand model and E[εi,t+1 | Ft ] = 0, E[Qi,t+1 | Ft ] = a − bPi,t+1 +

c X Pj,t+1 . N −1 j̸=i

Moreover, since Pt+1 is Ft -measurable, ⊤ ⊤ E[Pt+1 Pt+1 | Ft ] = Pt+1 Pt+1 ,

and E[Pi,t+1 Qi,t+1 | Ft ] = Pi,t+1 E[Qi,t+1 | Ft ]. Thus the noise coordinates satisfy P = 0, ξt+1

PP ξt+1 = 0N ×N ,

59

and, for each i, Qi ξt+1 = Qi,t+1 − E[Qi,t+1 | Ft ] = εi,t+1 ,

P Qi ξt+1 = Pi,t+1 εi,t+1 .

Step 3: Uniform conditional second-moment bound. Using the preceding identities, ∥ξt+1 ∥22 =

N  X

2

2

(εi,t+1 ) + (Pi,t+1 εi,t+1 )



2 ≤ (1 + Pmax )

i=1

N X

ε2i,t+1 .

i=1

2 Taking conditional expectation and using E[ε2i,t+1 | Ft ] ≤ σenv yields

2 E[∥ξt+1 ∥22 | Ft ] ≤ (1 + Pmax )

N X

2 2 E[ε2i,t+1 | Ft ] ≤ N (1 + Pmax )σenv .

i=1

Thus we may take 2 2 )σenv . σξ2 := N (1 + Pmax

Step 4: Exploration start condition. Under the exploration law, (X1 , . . . , XK ) are i.i.d. with mean µX := E[X1 ] and covariance ΣX := Var(X1 ). Hence K i h1 X Xs = µX , E[X̄K ] = E K s=1

and Var(X̄K ) = Var Finally,

G.4 Proof.

K 1 X

K

s=1

Xs



K

1 X 1 = 2 Var(Xs ) = ΣX . K K s=1

   tr(ΣX ) E ∥X̄K − µX ∥22 = tr Var(X̄K ) = . K

Proof of Lemma A.4 (Mean-field ODE convergence) Fix τ ∈ [1, ∞) and a sequence {Km }m≥1 with Km → ∞ and {nm }m≥1 satisfying nm ≥ Km

and nm /Km → τ . Fix α+ > τ . For all sufficiently large m, nm /Km ≤ α+ ; throughout we work on the finite window [1, α+ ]. Let M : [1, ∞) → RN × RN × RN ×N × RN solve the mean ODE Ṁ (t) =

1 h(M (t)), t

M (1) = µX .

(12)

We will show X̄nm − → M (τ ) by comparing the SA recursion to an Euler scheme and then to M , P

using a fixed Lipschitz tube around the ODE trajectory on [1, α+ ]. 60

Lemma G.1 (Mean-field bounds on M (t) on [1, α+ ]). Define Qmin := a − bPmax + cPmin > 0,

Qmax := a − bPmin + cPmax .

Then for all t ∈ [1, α+ ] and all i, j ∈ [N ], Qmin ≤ MiQ (t) ≤ Qmax ,

Pmin ≤ MiP (t) ≤ Pmax ,

Pmin Qmin ≤ MiP Q (t) ≤ Pmax Qmax .

2 2 Pmin ≤ MijP P (t) ≤ Pmax ,

For any state x, the induced next-period price vector is π(x) ∈ [Pmin , Pmax ]N ,

Proof of Lemma G.1

and the induced expected-demand vector is d(π(x)) ∈ RN . Since πi (x) ∈ [Pmin , Pmax ] for all i, we have di (π(x)) ∈ [Qmin , Qmax ] for all i. Writing the mean ODE (12) coordinatewise, for each i, j ∈ [N ],  1 πi (M (t)) − MiP (t) , t

ṀiQ (t) =

 1 πi (M (t))πj (M (t)) − MijP P (t) , t

ṀiP Q (t) =

ṀiP (t) = ṀijP P (t) =

 1 di (π(M (t))) − MiQ (t) , t  1 πi (M (t)) di (π(M (t))) − MiP Q (t) . t

Multiplying by t and integrating from 1 to t gives the identities 1 MiP (t) = t



MiP (1) +

Z t

 πi (M (s)) ds ,

1 MiQ (t) = t

1

1 MijP P (t) = MiP Q (t) =

1 t



MiQ (1) +

Z t

 di (π(M (s))) ds ,

1

  Z t PP Mij (1) + πi (M (s))πj (M (s)) ds ,

t 

1

MiP Q (1) +

Z t

 πi (M (s))di (π(M (s))) ds .

1

Now use the bounds πi (M (s)) ∈ [Pmin , Pmax ] and di (π(M (s))) ∈ [Qmin , Qmax ] for all s, so that 2 2 πi (M (s))πj (M (s)) ∈ [Pmin , Pmax ],

πi (M (s))di (π(M (s))) ∈ [Pmin Qmin , Pmax Qmax ].

The initial condition M (1) = µX satisfies the same bounds, because it is the corresponding exploration moment vector. Each display above therefore implies that the corresponding coordinate of M (t) remains within the stated interval for all t ≥ 1, and in particular for all t ∈ [1, α+ ]. Lemma G.2 (Own-price variance floor along the mean ODE). For each i, define 2 RiM (t) := MiiP P (t) − MiP (t) . Then tRiM (t) is nondecreasing and RiM (t) ≥

(Σexp )ii RiM (1) = t t 61

for all t ≥ 1.

In particular, for all t ∈ [1, α+ ], RiM (t) ≥ (Σexp )ii /α+ . Proof of Lemma G.2

ṀiP (t) =

From the ODE coordinates,  1 πi (M (t)) − MiP (t) , t

ṀiiP P (t) =

 1 πi (M (t))2 − MiiP P (t) . t

Differentiate RiM (t) = MiiP P (t) − (MiP (t))2 : ṘiM (t) = ṀiiP P (t) − 2MiP (t)ṀiP (t)  1 = πi (M (t))2 − MiiP P − 2MiP (πi (M (t)) − MiP ) t  1 = (πi (M (t)) − MiP )2 − (MiiP P − (MiP )2 ) t  1 (πi (M (t)) − MiP )2 − RiM (t) . = t Therefore

 d tRiM (t) = tṘiM (t) + RiM (t) = (πi (M (t)) − MiP (t))2 ≥ 0, dt

so tRiM (t) is nondecreasing and tRiM (t) ≥ RiM (1). Finally, RiM (1) = MiiP P (1) − (MiP (1))2 = Varexp (Pi,1 ) = (Σexp )ii .

Lemma G.3 (Lipschitz pricing map under (Q, Var(P )) floors). Fix constants ηQ > 0 and ηV > 0, and define the block-domain  S(ηQ , ηV ) :=

x = (xP , xQ , xP P , xP Q ) ∈ R4 : 0 ≤ xP ≤ Pmax + 1, ηQ ≤ xQ ≤ Qmax + 1, 0 ≤ xP P ≤ (Pmax + 1)2 , 0 ≤ xP Q ≤ (Pmax + 1)(Qmax + 1),  PP P 2 x − (x ) ≥ ηV .

Then fP : S(ηQ , ηV ) → [Pmin , Pmax ] is globally Lipschitz: there exists LP = LP (Pmin , Pmax , Qmax , ηQ , ηV ) < ∞ such that for all x, y ∈ S(ηQ , ηV ), |fP (x) − fP (y)| ≤ LP ∥x − y∥2 . Proof of Lemma G.3

Fix x ∈ S(ηQ , ηV ) and define v(x) := xP P − (xP )2 (≥ ηV ), β(x) :=

s(x) , v(x)

s(x) := xP Q − xP xQ ,

α(x) := xQ − β(x)xP . 62

Set the shorthand P+ := Pmax + 1 and Q+ := Qmax + 1. Step 1: explicit bounds and Lipschitz constants for (α, β). On S(ηQ , ηV ) we have 0 ≤ xP ≤ P+ , ηQ ≤ xQ ≤ Q+ , and 0 ≤ xP Q ≤ P+ Q+ , hence |s(x)| ≤ |xP Q | + |xP xQ | ≤ 2P+ Q+ =: Bs . Moreover, ∇v = (−2xP , 0, 1, 0) and ∇s = (−xQ , −xP , 0, 1), so ∥∇v∥2 ≤

q 4P+2 + 1 =: Lv ,

∥∇s∥2 ≤

q

Q2+ + P+2 + 1 =: Ls .

By the mean value theorem, for all x, y ∈ S(ηQ , ηV ), |v(x) − v(y)| ≤ Lv ∥x − y∥2 ,

|s(x) − s(y)| ≤ Ls ∥x − y∥2 .

Using v(·) ≥ ηV and |s(·)| ≤ Bs , we obtain  |β(x) − β(y)| ≤

Ls Bs Lv + 2 ηV ηV

 ∥x − y∥2 =: Lβ ∥x − y∥2 ,

|β(x)| ≤

Bs =: Bβ . ηV

Finally, |α(x)| ≤ |xQ | + |β(x)| |xP | ≤ Q+ + Bβ P+ =: Bα , and  |α(x) − α(y)| ≤ |xQ − y Q | + |β(x)xP − β(y)y P | ≤ 1 + Bβ + P+ Lβ ∥x − y∥2 =: Lα ∥x − y∥2 . Step 2: Lipschitzness of fP . Let [u][Pmin ,Pmax ] := min{max{u, Pmin }, Pmax }. With g(α, β) := −α/(2β), the pricing map can be written as  Pmax , fP (x) =   g(α(x), β(x)) [P

β(x) ≥ 0, min ,Pmax ]

, β(x) < 0.

If β(x) < 0, then α(x) = xQ − β(x)xP ≥ xQ ≥ ηQ , since xP ≥ 0. Define δ :=

ηQ > 0. 2Pmax

If β(x) ∈ [−δ, 0), then g(α(x), β(x)) =

ηQ α(x) ≥ = Pmax , 2|β(x)| 2δ

hence fP (x) = Pmax whenever β(x) ≥ −δ.

63

On the region {β ≤ −δ}, g is smooth and 1 ∂g ≤ , ∂α 2δ

Bα ∂g ≤ 2. ∂β 2δ

Thus, by the mean value theorem, for β(x), β(y) ≤ −δ, |g(α(x), β(x)) − g(α(y), β(y))| ≤

1 Bα |α(x) − α(y)| + 2 |β(x) − β(y)|. 2δ 2δ

Since u 7→ [u][Pmin ,Pmax ] is 1-Lipschitz, the same bound holds for |fP (x) − fP (y)|, giving  |fP (x) − fP (y)| ≤

Lα Bα Lβ + 2δ 2δ 2

 ∥x − y∥2 ,

whenever β(x), β(y) ≤ −δ.

Step 3: mixed-region pairs and conclusion. If β(x) ≥ −δ and β(y) ≤ −δ, then fP (x) = Pmax and Pmax − fP (y) ≤ g(α(y), −δ) − g(α(y), β(y)) ≤

 Bα Bα Lβ Bα |β(y)| − δ ≤ 2 |β(x) − β(y)| ≤ ∥x − y∥2 . 2 2δ 2δ 2δ 2

The symmetric case is identical. Combining all cases, we may take LP :=

Lα Bα Lβ + , 2δ 2δ 2

which proves global Lipschitzness of fP on S(ηQ , ηV ). Lemma G.4 (Fixed Lipschitz tube for h). Let Σexp,min := mini∈[N ] (Σexp )ii and set ηV :=

Σexp,min , 2α+

ηQ :=

Qmin . 2

Let Γ := {M (t) : t ∈ [1, α+ ]}. Then there exists R ∈ (0, 1] and constants Lh , Hh < ∞ such that on the tube  ΓR := x ∈ RN × RN × RN ×N × RN : dist(x, Γ) ≤ R , the following hold: (i) For every x ∈ ΓR and every i, the local block Q PP PQ P xloc i := (xi , xi , xii , xi )

lies in S(ηQ , ηV ).

64

(ii) h is Lh -Lipschitz and bounded by Hh on ΓR : ∥h(x) − h(y)∥2 ≤ Lh ∥x − y∥2

∀x, y ∈ ΓR ,

sup ∥h(x)∥2 ≤ Hh . x∈ΓR

Proof of Lemma G.4

Step 1: margins on the path. By Lemma G.1, for all t ∈ [1, α+ ], MiQ (t) ≥ Qmin = 2ηQ .

By Lemma G.2, for all t ∈ [1, α+ ], MiiP P (t) − (MiP (t))2 ≥

(Σexp )ii ≥ 2ηV . α+

Also Lemma G.1 gives the remaining coordinate bounds on Γ. Step 2: choose R so margins persist. Define the blockwise maps Q q(xloc i ) := xi ,

PP P 2 v(xloc i ) := xii − (xi ) .

On the bounded set 0 ≤ xPi ≤ Pmax + 1 we have P ∥∇v(xloc i )∥2 = ∥(−2xi , 0, 1, 0)∥2 ≤

p 4(Pmax + 1)2 + 1 =: L(1) v ,

(1)

so v is Lv -Lipschitz there, while q is 1-Lipschitz. Choose ( R := min 1, ηQ ,

ηV (1)

Lv

) , Pmin , Pmin Qmin

.

If x ∈ ΓR , pick t ∈ [1, α+ ] with ∥x − M (t)∥2 ≤ R. Then for each i, loc ∥xloc i − Mi (t)∥2 ≤ R,

hence Q loc loc xQ i ≥ Mi (t) − ∥xi − Mi (t)∥2 ≥ 2ηQ − R ≥ ηQ ,

and  loc loc (1) xPii P − (xPi )2 ≥ MiiP P (t) − (MiP (t))2 − L(1) v ∥xi − Mi (t)∥2 ≥ 2ηV − Lv R ≥ ηV . Moreover, loc xPi ≥ MiP (t) − ∥xloc i − Mi (t)∥2 ≥ Pmin − R ≥ 0,

and xPi ≤ Pmax + R ≤ Pmax + 1.

65

The same coordinatewise comparison gives 0 ≤ xPii P ≤ (Pmax + 1)2 ,

0 ≤ xPi Q ≤ (Pmax + 1)(Qmax + 1),

after shrinking by the above choice of R, because Γ satisfies the bounds in Lemma G.1 and each coordinate is 1-Lipschitz in Euclidean distance. This proves (i). Step 3: Lipschitzness of h on the tube. By (i) and Lemma G.3, each local block map loc xloc i 7→ fP (xi ) is LP -Lipschitz on ΓR . Therefore the vector price map π satisfies

∥π(x) − π(y)∥22 =

N X

loc 2 2 |fP (xloc i ) − fP (yi )| ≤ LP

i=1

N X

loc 2 2 2 ∥xloc i − yi ∥2 ≤ LP ∥x − y∥2 ,

i=1

so ∥π(x) − π(y)∥2 ≤ LP ∥x − y∥2 . Define the linear expected-demand map d : [Pmin , Pmax ]N → RN by di (p) := a − bpi +

c X pj . N −1 j̸=i

Its matrix has diagonal entries −b and off-diagonal entries c/(N − 1), hence its induced 1- and ∞-norms are both b + c, so ∥d(p) − d(p′ )∥2 ≤ (b + c)∥p − p′ ∥2 . Now define Φ : RN → RN × RN × RN ×N × RN by  Φ(p) := p, d(p), pp⊤ , p ⊙ d(p) . Using the bounds 0 ≤ pi ≤ Pmax and Qmin ≤ di (p) ≤ Qmax , one checks that for all p, p′ ∈ [Pmin , Pmax ]N , ∥Φ(p) − Φ(p′ )∥2 ≤ LΦ ∥p − p′ ∥2 , where one may take LΦ := Indeed,

q √ 2 1 + (b + c)2 + (2 N Pmax )2 + Pmax (b + c) + Qmax .

√ ∥pp⊤ − p′ p′⊤ ∥F ≤ ∥(p − p′ )p⊤ ∥F + ∥p′ (p − p′ )⊤ ∥F ≤ 2 N Pmax ∥p − p′ ∥2 ,

and ∥p ⊙ d(p) − p′ ⊙ d(p′ )∥2 ≤ Qmax ∥p − p′ ∥2 + Pmax ∥d(p) − d(p′ )∥2 .

66

Since f (x) = Φ(π(x)), it follows that ∥f (x) − f (y)∥2 ≤ LΦ ∥π(x) − π(y)∥2 ≤ LΦ LP ∥x − y∥2 . Finally, ∥h(x) − h(y)∥2 = ∥(f (x) − x) − (f (y) − y)∥2 ≤ ∥f (x) − f (y)∥2 + ∥x − y∥2 ≤ (LΦ LP + 1)∥x − y∥2 . Thus (ii) holds with Lh := LΦ LP + 1. Boundedness supx∈ΓR ∥h(x)∥2 < ∞ follows from continuity of h and compactness of ΓR , and we denote this bound by Hh . Step 1: log-time reparametrization and a deterministic Euler scheme. Define the logtime variable ℓ := log t, so t = eℓ , and the reparametrized mean-field trajectory M ℓ (ℓ) := M (eℓ ),

ℓ ≥ 0.

Since M (·) solves Ṁ (t) = 1t h(M (t)), the chain rule gives d ℓ M (ℓ) = h(M ℓ (ℓ)), dℓ

M ℓ (0) = M (1) = µX .

Thus M ℓ solves the autonomous ODE d ℓ M (ℓ) = h(M ℓ (ℓ)), dℓ

(13)

M ℓ (0) = µX .

Fix m and abbreviate K := Km and n := nm . For calendar times s ≥ K, define the accumulated step size

s X 1 SK,s := , r

SK,K = 0.

r=K+1

For all s ∈ {K, . . . , n}, Z s SK,s ≤

s n dx = log ≤ log ≤ log α+ , K K K x

for all sufficiently large m. Hence SK,s ∈ [0, log α+ ] and M ℓ (SK,s ) ∈ Γ := {M (t) : t ∈ [1, α+ ]}

for all s ∈ {K, . . . , n}.

Define the deterministic Euler scheme for (13) on the grid {SK,s }: YK,K := µX ,

YK,s+1 := YK,s + 67

1 h(YK,s ), s+1

s ≥ K.

Define the Euler error es := YK,s − M ℓ (SK,s ),

s ≥ K,

so eK = 0. We now prove the uniform bound (15). The argument uses only that on the tube ΓR from Lemma G.4, h is Lh -Lipschitz and bounded by Hh . Step 1a (local truncation error for the exact flow). For each s ≥ K, ℓ

Z SK,s+1

M (SK,s+1 ) − M (SK,s ) =

h(M ℓ (u)) du.

SK,s

Write ∆s := SK,s+1 − SK,s =

1 . s+1

Add and subtract ∆s h(M ℓ (SK,s )) to obtain M ℓ (SK,s+1 ) = M ℓ (SK,s ) + ∆s h(M ℓ (SK,s )) + ρs+1 , where ρs+1 :=

Z SK,s+1 

 h(M (u)) − h(M (SK,s )) du. ℓ

SK,s

Since M ℓ (u) ∈ Γ ⊆ ΓR and ∥h(·)∥ ≤ Hh on ΓR , for u ∈ [SK,s , SK,s+1 ], ∥h(M ℓ (u)) − h(M ℓ (SK,s ))∥2 ≤ Lh ∥M ℓ (u) − M ℓ (SK,s )∥2 ≤ Lh Hh (u − SK,s ), and therefore Z SK,s+1 ∥ρs+1 ∥2 ≤

Lh Hh (u − SK,s ) du = SK,s

1 Lh Hh 2 Lh Hh ∆s = . 2 2 (s + 1)2

Step 1b (error recursion, localized to the tube). Define the deterministic exit time from the tube τE := inf{s ∈ {K, K + 1, . . . , n} : YK,s ∈ / ΓR },

(inf ∅ := n + 1).

For s < τE we have YK,s ∈ ΓR and M ℓ (SK,s ) ∈ Γ ⊆ ΓR , so es+1 = YK,s+1 − M ℓ (SK,s+1 )     = YK,s + ∆s h(YK,s ) − M ℓ (SK,s ) + ∆s h(M ℓ (SK,s )) + ρs+1   = es + ∆s h(YK,s ) − h(M ℓ (SK,s )) − ρs+1 .

68

Taking norms and using Lipschitzness of h on ΓR gives, for all s < τE ,  1 Lh  Lh Hh . ∥es+1 ∥2 ≤ 1 + ∥es ∥2 + s+1 2 (s + 1)2

(14)

Step 1c (solve the recursion). Let as := ∥es ∥2 and c := Lh Hh /2. From (14), for all s < τE ,  Lh  c as+1 ≤ 1 + as + , s+1 (s + 1)2

aK = 0.

Unrolling yields, for K ≤ s ≤ τE ∧ n, s X

as ≤ c

r=K+1

s 1 Y  Lh  . 1 + r2 j j=r+1

Moreover, log(1 + Lh /j) ≤ Lh /j implies s  Y j=r+1

1+

L  h

j

 s X Lh   s Lh  n Lh Lh ≤ exp ≤ ≤ α+ , ≤ j r K j=r+1

for all sufficiently large m. Hence, for s ≤ τE ∧ n, Lh ∥es ∥2 ≤ c α+

n X r=K+1

Define CE :=

1 . r2

Lh Hh Lh α+ . 2

Step 1d (the Euler iterates stay in the tube). Since have max

K≤s≤τE ∧n

∥es ∥2 ≤

Pn

r=K+1 r

−2 ≤

R∞ K

x−2 dx = 1/K, we

CE . K

Because Km → ∞, for all sufficiently large m we have CE /K ≤ R/2. Fix such an m. Then for all s ≤ τE ∧ n, ∥YK,s − M ℓ (SK,s )∥2 = ∥es ∥2 ≤ R/2, and since M ℓ (SK,s ) ∈ Γ, this implies YK,s ∈ ΓR/2 ⊆ ΓR for all s ≤ τE ∧ n. Thus τE = n + 1, i.e., YK,s ∈ ΓR for all s ∈ {K, . . . , n}, and max ∥YK,s − M ℓ (SK,s )∥2 = max ∥es ∥2 ≤

K≤s≤n

K≤s≤n

CE . K

Step 2: SA vs. Euler (martingale perturbation) via localization. Let ds := X̄s − YK,s . 69

(15)

Define the stopping time τSA := inf{s ∈ {K, K + 1, . . . , n} : ∥ds ∥2 > R/2},

(inf ∅ := ∞),

and the stopped process d˜s := ds∧τSA . On {s < τSA } we have X̄s , YK,s ∈ ΓR , hence ∥h(X̄s ) − h(YK,s )∥2 ≤ Lh ∥ds ∥2 . Using Lemma A.3 and the Euler recursion for YK,s , for s ≥ K we have ds+1 = ds +

 1 1 h(X̄s ) − h(YK,s ) + ξs+1 . s+1 s+1

Taking conditional expectation of ∥d˜s+1 ∥22 given Fs and using E[ξs+1 | Fs ] = 0 yields   E ∥d˜s+1 ∥22 | Fs ≤

    L2h 1 2Lh + ∥d˜s ∥22 + E ∥ξs+1 ∥22 | Fs . 1+ 2 2 s + 1 (s + 1) (s + 1)

Using E[∥ξs+1 ∥22 | Fs ] ≤ σξ2 and (s + 1)−2 ≤ (s + 1)−1 , there is a constant C1 such that E∥d˜s+1 ∥22 ≤ Iterating and using

  σξ2 C1 2 ˜ 1+ E∥ds ∥2 + . s+1 (s + 1)2 n X

n X 1 ≤ log α+ , r

r=K+1

r=K+1

gives

1 1 ≤ , 2 r K

  1 2 2 ˜ E∥dn ∥2 ≤ C2 E∥dK ∥2 + , K

where C2 < ∞ depends only on C1 , σξ2 , α+ . Since dK = X̄K − µX , Lemma A.3 gives E∥dK ∥22 = O(1/K), hence E∥d˜n ∥22 = O(1/K). Therefore

  4 R ˜ Pr(τSA ≤ n) ≤ Pr ∥dn ∥2 ≥ ≤ 2 E∥d˜n ∥22 −→ 0. 2 R

On {τSA > n} we have dn = d˜n , so P

dn − → 0.

70

(16)

→ 0, Step 3: conclude X̄nm → M (τ ). By (15) and dn − P

X̄n − M ℓ (SK,n ) = (X̄n − YK,n ) + (YK,n − M ℓ (SK,n )) − → 0. | {z } | {z } P

en

dn

It remains to identify the limit of M ℓ (SK,n ). Using the integral test for harmonic sums, Z n n X dx dx 1 1 ≤ ≤ + , r K K x K x

Z n

r=K+1

so SK,n = log

n K



1 +O K

 ,

hence e

SK,n

   n 1 = −→ τ. · exp O K K

By continuity of M (·), M ℓ (SK,n ) = M (eSK,n ) → M (τ ). Therefore P

X̄nm − → M (τ ), proving Lemma A.4.

G.5 Proof.

Proof of Lemma A.5 Let M (·) be the mean-field ODE solution from Lemma A.4, and write  M (t) = M P (t), M Q (t), M P P (t), M P Q (t) .

By definition of the mean-field moment coordinates, for t ≥ 1 the price mean M P (t) ∈ RN and raw price second-moment matrix M P P (t) ∈ RN ×N satisfy Ṁ P =

P − MP , t

Ṁ P P =

PP⊤ − MPP , t

where P (t) = π(M (t)). Define U (t) := M P (t),

ΣP (t) := M P P (t) − U (t)U (t)⊤ ,

V (t) := t ΣP (t).

Deriving the (U, V ) dynamics. From U = M P we immediately have U̇ =

P −U . t

For V = tΣP , the product rule gives V̇ = ΣP + t Σ̇P .

71

Also Σ̇P = Ṁ P P − U̇ U ⊤ − U U̇ ⊤ . Substituting the ODEs for Ṁ P P and U̇ and using M P P = ΣP + U U ⊤ yields t Σ̇P = P P ⊤ − (ΣP + U U ⊤ ) − (P − U )U ⊤ − U (P − U )⊤ = (P − U )(P − U )⊤ − ΣP . Therefore  V̇ = ΣP + (P − U )(P − U )⊤ − ΣP = (P − U )(P − U )⊤ . Expressing P (t) as a function of (U, V ). Fix i ∈ [N ] and write Ū−i :=

1 X Uj , N −1

V̄i,−i :=

j̸=i

1 X Vij . N −1 j̸=i

The mean-field moment coordinates inherit the linear moment identities from the demand model: MiQ = a − b Ui + c Ū−i , and MiP Q = a Ui − b MiiP P + c M P P i,−i ,

M P P i,−i :=

1 X PP Mij . N −1 j̸=i

Since we have

1 M P P = U U ⊤ + V, t 1 MiiP P = Ui2 + Vii , t

1 M P P i,−i = Ui Ū−i + V̄i,−i . t

Substituting into MiP Q gives MiP Q = Ui MiQ +

 1 − bVii + c V̄i,−i , t

so MiP Q − MiP MiQ =

−bVii + c V̄i,−i . t

Also MiiP P − (MiP )2 =

Vii . t

Hence the OLS slope estimate in moment form satisfies β̂i =

MiP Q − MiP MiQ −bVii + c V̄i,−i = . P P P 2 Vii Mii − (Mi )

72

Thus: • If −bVii + c V̄i,−i ≥ 0, equivalently β̂i ≥ 0, then predicted profit pi (α̂i + β̂i pi ) is nondecreasing at the relevant upper endpoint in the sense established in Lemma A.1, and the selected price is Pmax . • If −bVii + c V̄i,−i < 0, equivalently β̂i < 0, the unconstrained maximizer is −α̂i /(2β̂i ). Using α̂i = MiQ − β̂i Ui ,

MiQ = a − bUi + cŪ−i ,

and the moment identities above, a direct substitution gives (a + c Ū−i )Vii − c Ui V̄i,−i  Pei (U, V ) = . 2 bVii − c V̄i,−i Enforcing feasibility gives Pi = [Pei (U, V )][Pmin ,Pmax ] . This is exactly the pricing map stated in Lemma A.5, so P (t) = π(M (t)) is the posted-price coordinate of the price-moments ODE. Initial conditions. Because U = M P , we have U (1) = M P (1) = µ. Also,  V (1) = 1 · M P P (1) − U (1)U (1)⊤ = (µµ⊤ + Σexp ) − µµ⊤ = Σexp . Therefore (U, V ) solves U (1) = µ,

V (1) = Σexp ,

U̇ =

P −U , t

V̇ = (P − U )(P − U )⊤ ,

with the stated price map. Finally, the recovery formulas follow from the same identities: MiP (t) = Ui (t),

MiQ (t) = a − bUi (t) + cŪ−i (t),

1 M P P (t) = U (t)U (t)⊤ + V (t), t and MiP Q (t) = Ui (t)MiQ (t) +

 1 −bVii (t) + cV̄i,−i (t) . t

These dynamics and the posted-price coordinate P (t) coincide with Definition 3.1.

73

Record · ID 192410 · SHA-256 e5064e8f5cc1e238
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.