arXiv:2609.06351v1 [eess.SY] 6 Sep 2026
A Theory of Information Architecture for Networked Decisions: Freshness, Locality, and Coordination Scott Moeller
Bhaskar Krishnamachari
Independent Researcher, Portsmouth, RI, USA [email protected]
University of Southern California [email protected]
August 2026
Abstract Networked systems face a tradeoff between the scope of the information behind a decision and its freshness: a broader view of the system supports better coordination, but assembling and communicating it takes time, so it arrives older. We study this tradeoff for a team of agents that repeatedly choose actions to minimize a shared quadratic cost driven by an environment that evolves on its own, unaffected by the agents’ actions. An information architecture specifies what each agent observes, from where, and with what delay; we measure an architecture by the cost it loses relative to a decision made with complete, current information. We first compare fresh local observations with a complete but delayed global view. When every component of the environment decorrelates at a common exponential rate, this comparison reduces to a closed-form threshold on the ratio of delay to coherence time. We then study intermediate architectures in which each agent acts on a time-aligned, and therefore older, snapshot of a wider neighborhood. The optimal neighborhood radius occurs where the marginal value of added scope equals the marginal cost of lost freshness. In a canonical spatial model, this radius is set by the spatial-correlation, decision-relevance, and temporal-propagation lengths. Throughout, architecture performance is governed by the predictability of the optimal decision rather than of the raw state.
1
Introduction
Many networked systems must repeatedly make coupled decisions using information that is distributed across space and changes over time. A wireless transmitter may know its current local channel state but not the current interference conditions elsewhere in the network; an edge server may observe its own workload immediately while a system-wide view is available only after telemetry has been collected and disseminated; and a distributed sensing system may have fresh measurements from nearby sensors while observations from distant sensors arrive with greater latency. In each case, broad information can improve coordination, but obtaining it takes time. The resulting design choice is not simply between centralized and distributed control, but between information architectures that differ simultaneously in spatial scope and temporal freshness. This tension appears in several literatures under different performance models. Work on stale information in distributed systems shows that old global state can lose much of its value [1], and Age of Information formalizes freshness as a system resource [2]. Distributed optimization studies computation under communication constraints [3], organizational economics studies adaptation to local information versus coordination [4], and networked-control work examines how communication scope, topology, and propagation delay interact with closed-loop performance. 1
Distributed-control work establishes controller localization and models finite-speed or distancedependent communication [5–8]. Most directly, Ballotta, Jovanović, and Schenato show that scopedependent delay can make sparse feedback outperform all-to-all control [9]; Section 7.2 develops the comparison with this literature. The distinction pursued here is the decision problem and the resulting analytical object. The environment evolves exogenously, and each epoch is a static team decision. Dynamic feedback control, in which the controller changes future plant state, lies outside this setting. For a fixed information architecture, performance is the Q-weighted loss in reproducing the current full-information decision, where Q is the cost’s Hessian, defined in Section 2.1. We then vary spatial scope and information age jointly. Under a model in which every component of the environment decorrelates at the same exponential rate (the common-rate model), this prediction-theoretic formulation yields an exact composition identity for spatial omission and temporal staleness, a fresh-local versus stale-global crossover in terms of decision-relevant predictability, and a marginal-balance characterization of the synchronized-snapshot radius. In the canonical field model, these quantities produce an explicit radius law involving temporal coherence, spatial correlation, decision-relevance length, and propagation speed. This paper develops a theory for this tradeoff. We consider a collection of agents operating in an exogenously evolving stochastic environment θt = (θ1,t , . . . , θn,t ). At each decision epoch the agents select a joint action xt = (x1,t , . . . , xn,t ) that incurs a common cost f (xt , θt ). An information architecture specifies the information available to each agent when this decision is made. Thus a fresh-local architecture may allow agent i to condition on the current θi,t , a neighborhood architecture may provide the states of agents within a given spatial radius, and a global architecture may provide the entire state vector only after a delay τ . More generally, an architecture specifies who knows what, from where, and at what age. Our analysis uses the classical fixed-information framework of static team decision theory initiated by Radner [10]: a static team has multiple decision makers sharing one objective, each acting on its own information, whose actions do not affect what anyone later observes. In the quadratic model considered here, whose Hessian Q is constant and state-independent, completing the square reduces optimization under any fixed information architecture to a Hilbert-space approximation problem: the team-optimal policy under any fixed architecture is the Q-orthogonal projection of the full-information action, and architecture regret is the squared projection distance (Section 2.4). We study an outer design problem: comparison and selection within a structured family in which information scope and age vary jointly. Classical team and information-structure work has also treated information and organizational networks as design objects [11, 12]. The narrower question pursued here is how scope-dependent age changes the value of a synchronized-snapshot radius family when information is evaluated through prediction error of the full-information decision. Our first result answers the two extreme points of this tradeoff. We compare a fresh-local architecture, in which each agent observes its current local state, against a stale-global architecture, in which every agent has access to the complete system state with delay τ . In general, the loss from staleness is governed by the prediction error of the current decision-relevant state from information τ units old. For a Gauss–Markov environment with temporal coherence time T , and for an affine fullinformation decision, stale-global regret grows as 1−e−2τ /T toward the open-loop regret. Fresh-local information is preferred once this temporal loss exceeds the share of decision-relevant variation not captured locally. In a canonical shared-resource quadratic model, where τ is the global-information delay, T is the temporal coherence time, q is local stiffness, and γ is coordination strength, the large-system threshold is τ ⋆ 1 γ = log 1 + , (1) T 2 q 2
making explicit how stronger decision coupling increases the amount of staleness that is worth tolerating in exchange for global coordination. The local–global comparison concerns two pure benchmark regimes. Our second result considers agents embedded in a graph or spatial domain and a synchronized-snapshot family in which each agent bases its decision on a time-aligned neighborhood snapshot of radius r. Increasing r replaces a narrower, fresher snapshot by a broader snapshot of greater age τ (r). The resulting architecture faces two competing sources of error: a spatial omission error, caused by excluding decision-relevant state beyond radius r, and a temporal staleness error, caused by the delay required to acquire the broader information. We characterize this tradeoff using the spatial and temporal predictability of the underlying stochastic field. In a canonical model, the optimal radius is determined jointly by spatial correlation, temporal coherence, the scope–delay relation, and the coupling structure of the decision objective. The result yields comparative statics with direct architectural interpretations: faster temporal variation favors smaller and fresher information radii; faster information propagation supports broader coordination; and stronger long-range decision coupling increases the value of larger information scope. Spatial smoothness determines how rapidly additional spatial observations become redundant and hence how much information radius is needed to approximate global decisions. Under tractable spatio-temporal covariance and propagation models, these effects combine into dimensionless scaling laws for r⋆ . The framework separates information architecture from the algorithm used to implement it. It consequently complements canonical distributed-optimization formulations, where the communication structure is specified first and the principal question is how to compute efficiently over it [3]. We study a structured outer comparison over a scope–age family. Similarly, unlike freshness metrics that assign value to an update primarily through its age [2], the value of an observation here depends on how much it changes the downstream decision. This perspective also suggests extensions to heterogeneous information radii, sparse long-range information “shortcuts,” low-dimensional global summaries, and budget-constrained information acquisition. Our model is restricted to static information structures: the stochastic environment may be temporally and spatially correlated, but agents’ current actions do not alter the information subsequently observed by other agents. This boundary is consequential. Once actions influence future observations, they act simultaneously as control actions and implicit signals; Witsenhausen’s counterexample [13] shows that this interaction breaks the projection structure our analysis relies on. This theory is most directly applicable to repeated networked optimization against an exogenously evolving environment, for example, stylized quadratic or locally quadratic models of rate, power, compute, or sensing decisions driven by changing demand, channel, or resource conditions. Fully constrained engineering formulations, controlled queueing, and decentralized Markov decision processes lie beyond this baseline. In summary, this paper makes two technical contributions and one conceptual contribution: 1. We derive an exact fresh-local versus stale-global crossover law for two pure information regimes, expressed through decision-relevant temporal predictability and the share of decision variation captured locally. 2. For a synchronized-snapshot radius family under the common-rate model, we derive an exact spatio-temporal prediction-error composition law and use it to characterize the optimal radius by a marginal balance between spatial-information gain and freshness loss. In a canonical field model, this yields an explicit scaling law in the spatial-correlation, decision-relevance, and temporal-propagation lengths.
3
3. We show that, within the analyzed framework, architecture performance is governed by decision-relevant predictability, not raw state predictability. The desired architecture depends jointly on environmental statistics, decision sensitivity and coupling, information geometry, and freshness; the communication graph and decision geometry therefore need not coincide. To situate the contribution precisely: scope-dependent communication delay and optimal architectures of intermediate sparsity both appear in prior networked-control work [9, 14–16]. What is new here is the decision problem and the resulting analytical object: an exogenous environment, a stagewise team decision, and an architecture valued by its prediction error for the full-information decision. This formulation yields the composition identity between spatial omission and temporal staleness, the fresh-local versus stale-global crossover, and the coordination-radius laws that follow. Section 2 formalizes architectures and regret. Section 3 compares the fresh-local and stale-global endpoints, and Sections 4– 5 develop the radius family and its optimal coordination scale. Section 6 reports numerical checks, Section 7 situates the work, and Section 8 concludes.
2
Networked Decisions and Information Architectures
We consider a network of decision-making agents operating in a stochastic environment whose state is distributed across the network and evolves over time. Our formulation separates three distinct objects: the network geometry, which defines relationships and distances among agents; the decision problem, which determines how their actions and states interact in the objective; and the information architecture, which determines what information is available to each agent and with what delay. These three structures need not coincide. The environment evolves exogenously: agents react to the state but their current actions do not alter the law governing its future evolution. The model therefore describes repeated informationconstrained decisions against an exogenous environment. Under the fixed-information static-team formulation, the quadratic model used here has a constant, state-independent Hessian and admits the projection characterization developed below. Table 1 collects the principal notation used throughout the paper; we introduce each object in context below.
2.1
Network, Environment, and Decision Model
There are n decision-making agents indexed by V = {1, . . . , n}, represented as the vertices of a graph G = (V, E).
(2)
The graph provides a notion of network or spatial proximity. We write dG (i, j) for the graph distance between nodes i and j, and define the radius-r neighborhood of node i as Nr (i) ≜ {j ∈ V : dG (i, j) ≤ r}.
(3)
Depending on the application, G may represent a physical communication network, a geographic neighborhood relation, or another structure governing the cost or latency of obtaining nonlocal information. 1 1
The graph model is adopted for concreteness. The framework can more generally be formulated for agents embedded in a metric space, or for any collection of agents equipped with a meaningful notion of neighborhoods and information-propagation cost.
4
Let θ = {θt }t∈T be a stochastic process defined on a probability space (Ω, F, P), where T may be discrete or continuous. At time t, the environment state is θt = (θ1,t , . . . , θn,t ) ∈ Rm ,
θi,t ∈ Rmi ,
n X
mi = m.
(4)
i=1
The component θi,t denotes the state locally associated with node i. Depending on the application, it may represent a local channel condition, workload or demand, sensing quality, resource availability, or another externally evolving quantity relevant to the decision. For the basic results in this section, we assume that θt is stationary and square-integrable: E∥θt ∥2 < ∞.
(5)
No independence, Gaussianity, Markov property, or specific form of spatial or temporal correlation is required. More structured stochastic models will be introduced later when we derive explicit architecture laws. Agent i chooses an action xi,t ∈ Rdi , and the joint action is xt = (x1,t , . . . , xn,t ) ∈ Rd ,
d=
n X
di .
(6)
i=1
The agents share a common per-stage cost 1 f (x, θ) = x⊤ Qx − b(θ)⊤ x + c(θ), 2
(7)
where Q ∈ Rd×d is symmetric and positive definite, b : Rm → Rd is measurable, and c : Rm → R is an additive state-dependent term. We assume E∥b(θt )∥2 < ∞,
E|c(θt )| < ∞.
(8)
The matrix Q describes coupling among the agents’ actions. If Q is partitioned according to x = (x1 , . . . , xn ), an off-diagonal block Qij captures direct interaction between the decisions of agents i and j. The mapping b(θ) provides another form of coupling: the preferred action of agent i may depend on states associated with other agents. We emphasize that this decision coupling is distinct from the graph G. The graph describes network proximity and information propagation, whereas Q and b(·) determine which states and actions are relevant to one another in the objective. Nearby agents need not be strongly coupled, and strong decision coupling may exist between distant nodes. If the complete current state θt is available without information constraints, the unique optimal joint action is x⋆t ≜ x⋆ (θt ) = Q−1 b(θt ). (9) We call x⋆t the clairvoyant or full-information decision. For the affine specializations used to obtain closed forms, we write b(θ) = Bθ + b0 ,
B ∈ Rd×m ,
b0 ∈ Rd , 5
K ≜ Q−1 B ∈ Rd×m .
(10)
The matrix B maps state into the linear term of the objective, while K is the resulting decisionsensitivity operator. The general projection results do not require this affine form. Completing the square in (7) gives f (x, θ) − f (x⋆ (θ), θ) =
1 (x − x⋆ (θ))⊤ Q (x − x⋆ (θ)) . 2
(11)
Hence, for the quadratic model, excess decision cost is exactly a weighted squared error in reproducing the full-information action.
2.2
Information Architectures
An information architecture specifies what information is available to each agent when its decision is made. At a fixed decision epoch, let Gi ⊆ F denote the σ-algebra representing the information available to agent i. An information architecture is the tuple A = (G1 , . . . , Gn ). (12) A decision rule for agent i is admissible only if it is measurable with respect to Gi . This abstraction accommodates information patterns that differ in spatial scope, temporal freshness, and degree of aggregation. Fresh-local information.
Agent i observes its own current state and local history: loc Gi,t = σ(θi,s : s ≤ t) .
(13)
This architecture has minimal spatial scope and no imposed information delay. Stale-global information. Every agent has access to the complete state history, but only up to time t − τ : glob Gi,t (τ ) = σ(θs : s ≤ t − τ ) , i = 1, . . . , n. (14) This architecture has maximal spatial scope but potentially substantial staleness. Radius-r information. graph distance r:
More generally, agent i may obtain state information from nodes within (r) Gi,t = σ θj,t−τ (r) : j ∈ Nr (i) .
(15)
The radius family studied below is a synchronized-snapshot architecture. At radius r, the decision is formed from a time-aligned neighborhood snapshot whose age is τ (r); increasing r replaces, rather than augments, the fresher narrow snapshot. This restriction models applications requiring temporal consistency or a common decision epoch. If fresh local and delayed remote information can instead be freely combined, the resulting hybrid architecture contains more information and weakly dominates either source alone; such mixed-age architectures are outside the principal family analyzed here. A closely related common-delay-by-scope abstraction appears in Ballotta, Jovanović, and Schenato [9]; see Section 7.2. Here τ (r) denotes the latency required to acquire information over radius r. We focus on systems for which r2 ≥ r1 =⇒ τ (r2 ) ≥ τ (r1 ), (16) 6
so that broader information is systematically less fresh. The synchronized-snapshot radius family interpolates between its local and global endpoints under the stated temporal model. At the smallest radius, each decision relies on a narrow, rapidly available snapshot. As r increases, that snapshot is replaced by a broader but older view. When r reaches the diameter of a finite graph, (15) is a delayed global snapshot, not the complete delayed global history in (14). Local information plus a global summary. An architecture need not transport raw remote state. Agent i may instead combine fresh-local information with a low-dimensional statistic zt = h(θt ) of the global state: hyb Gi,t = σ(θi,s : s ≤ t) ∨ σ(zs : s ≤ t − τz ) .
(17)
Examples include congestion prices, aggregate load indicators, or other summaries intended to communicate the globally decision-relevant portion of the state. The hybrid always contains the fresh-local information set. It dominates the stale-global benchmark only if the delayed summary preserves all delayed global information; an arbitrary low-dimensional summary need not. For a given architecture A, define the admissible joint-policy set S(A) ≜ u = (u1 , . . . , un ) : ui ∈ L2 , ui is Gi -measurable, i = 1, . . . , n . (18) Thus an information architecture constrains the information on which each decision may depend; it does not prescribe a specific algorithm for computing that decision. The sets Gi are exogenously specified: current actions neither change the law of θt nor other agents’ information, as discussed in Section 1.
2.3
Architecture Regret
We evaluate an architecture by the best expected performance achievable under its information restrictions. Define J(A) ≜ inf E [f (u, θt )] , (19) u∈S(A)
and let J ⋆ ≜ E [f (x⋆t , θt )]
(20)
denote the expected full-information cost. The architecture regret is R(A) ≜ J(A) − J ⋆ .
(21)
Here “regret” denotes a static expected performance gap after optimizing over all policies admissible under the architecture, distinct from cumulative regret in sequential decision making or online learning. Because the environment is stationary, this expected per-stage regret is independent of the particular decision epoch, and the expected finite-horizon average equals the same one-stage expectation. Almost-sure convergence of sample-path time averages additionally requires ergodicity. At this stage, R(A) measures only the decision-quality consequence of an information restriction. We do not separately charge for communication, sensing, or computation. The physical cost of obtaining broader information enters through the feasible architecture and, in particular, through the scope–latency relation τ (r). 7
If two architectures A and A′ satisfy Gi ⊆ Gi′ ,
i = 1, . . . , n,
(22)
then every policy feasible under A is also feasible under A′ , and therefore R(A′ ) ≤ R(A).
(23)
Thus additional information cannot hurt when it is available with the same freshness. The interesting architectural tradeoff arises because increasing spatial scope generally also increases information age.
2.4
Fixed-Architecture Quadratic Projection
The classical fixed-information static-team formulation supplies coupled person-by-person optimality conditions [10]. In this quadratic model, completing the square turns those conditions into the projection characterization below. We use this fixed-architecture result as the inner optimization primitive for the outer scope–age comparison. Related affine-gain projection formulations appear in Krainak, Speyer, and Marcus [17]. Let H = L2 (Ω, F, P; Rd ) (24) and equip H with the weighted inner product h i ⟨u, v⟩Q ≜ E u⊤ Qv ,
∥u∥2Q ≜ ⟨u, u⟩Q .
(25)
Since Q ≻ 0, the induced norm is equivalent to the ordinary L2 norm. Lemma 1 (Fixed-architecture quadratic projection; Radner-type specialization). For any static information architecture A, the admissible policy set S(A) defined in (18) is a closed linear subspace of H. Under the quadratic cost (7), the unique optimal policy implementable under A is u⋆A = ΠS(A) x⋆t ,
(26)
where ΠS(A) denotes orthogonal projection with respect to (25). Consequently, R(A) =
1 ⋆ 2 x − ΠS(A) x⋆t Q . 2 t
(27)
This identity specializes the classical static-team formulation [10] to the quadratic model used here. Proof. Linearity of S(A) follows because Gi -measurability is preserved under linear combinations. Closedness follows because an L2 limit of Gi -measurable random variables admits a Gi -measurable version. For any admissible policy u, (11) gives i 1 h E [f (u, θt ) − f (x⋆t , θt )] = E (u − x⋆t )⊤ Q(u − x⋆t ) 2 1 = ∥u − x⋆t ∥2Q . (28) 2 Minimizing this expression over the closed subspace S(A) is the standard Hilbert-space projection problem, yielding (26) and (27). 8
Corollary 1 (Coupled conditional normal equations). The unique team optimum in Lemma 1 satisfies i = 1, . . . , n. E Qu⋆A − b(θt ) i Gi = 0, (29) With Q partitioned into action blocks, this is equivalently X Qii u⋆i + Qij E[u⋆j | Gi ] = E[bi (θt ) | Gi ].
(30)
j̸=i
These are the coupled person-by-person stationarity equations associated with the classical staticteam formulation [10]. Strict convexity makes them sufficient for the unique team optimum here. With heterogeneous information and off-diagonal Q, one generally does not have u⋆i = E[x⋆i | Gi ]. When all components share a common information sigma-algebra, deterministic Q makes the joint solution the ordinary vector conditional expectation. Remark 1 (Scope of the projection identity). The fixed-architecture projection identity requires square integrability, a deterministic positive-definite matrix Q, unconstrained Euclidean actions, and a closed linear space of admissible policies. It does not require Gaussianity, an affine map b(θ), Markov dynamics, temporal stationarity, or independent observations. Those stronger assumptions enter only in later closed-form prediction and composition results. Lemma 1 turns an information architecture into a geometric object. The full-information policy The architecture restricts the system to the policy subspace S(A), and its regret is precisely the squared distance from the ideal policy to the closest rule implementable using the available information. This interpretation also identifies the relevant notion of environmental predictability. An architecture need not reconstruct the complete state θt . It need only provide enough information to predict the components that affect x⋆t . Section 4.4 develops the corresponding notion of decisionrelevant predictability. The radius-dependent architecture in (15) makes the central tradeoff clear. Increasing r provides each agent with a broader view of the network and, if freshness were fixed, could only improve performance. In a physical network, however, larger r generally also increases τ (r). The architecture family Ar,τ (r) r≥0 (31) x⋆t specifies what the system would ideally do if the current global state were known.
therefore trades spatial information against temporal freshness. We first study the two endpoints of this tradeoff: fresh-local information and stale-global information. We then ask the more general architecture-design question: what information radius minimizes regret when the value of broader spatial knowledge must be balanced against the additional delay required to obtain it? The notation admits a direct networked interpretation through illustrative mappings. The component θi may represent an externally evolving local condition, for example, channel quality, demand, workload intensity, sensing conditions, or resource availability, while xi is a corresponding continuous decision such as rate, power, compute allocation, or sensing effort. The matrices Q and B, or equivalently the decision-sensitivity operator K = Q−1 B, encode how other agents’ decisions and remote state affect the desired action; this coupling structure need not coincide with the physical or information graph over which observations travel. The essential boundary is exogeneity. An external arrival or workload process may be modeled as part of θ, whereas a queue length whose evolution depends on earlier service decisions is not exogenous in the sense required here. 9
Table 1: Core notation. Quantities labeled as shares are normalized by R∞ . Symbol
Meaning
Symbol
n θt Q
Meaning
number of agents exogenous environment state positive-definite action-coupling matrix K = Q−1 B decision-sensitivity operator A information architecture
G = (V, E) network or geometry graph xt joint action B affine state-to-objective map
R∞ ηT (τ ) ηST (r, τ )
ηS (0) ηS (r) T
τ (r) ℓs LT = vT
3
open-loop regret temporal unpredictability normalized joint spatio-temporal regret information latency at radius r spatial correlation length temporal propagation length
x⋆t R(A)
v ℓc r⋆
full-information decision architecture regret; includes Rloc and Rglob (τ ) fresh-local omission share radius-r zero-delay spatial omission temporal coherence time effective propagation speed decision-relevance length optimal synchronized-snapshot radius
Fresh-Local versus Stale-Global Information
We begin with the two endpoints of the freshness–scope tradeoff introduced in Section 2. A freshlocal architecture gives each agent immediate access to its own state but no current information about remote states. A stale-global architecture gives every agent complete system-wide information, but only after a delay τ . This section compares two pure regimes; as noted in Section 2.2, an architecture free to combine both information sources would weakly dominate either one. If both architectures had the same delay, global information could only help: the corresponding information sets would contain the local ones. The comparison becomes nontrivial precisely because spatial scope and freshness move in opposite directions. This section characterizes that tradeoff first for an arbitrary stationary environment through a temporal prediction-error function, and then obtains an exact crossover law under a Gauss–Markov model.
3.1
The Two Architectures
At decision time t, the fresh-local architecture is loc loc Aloc = G1,t , . . . , Gn,t ,
loc Gi,t = σ(θi,s : s ≤ t) .
(32)
Its regret is Rloc ≜ R(Aloc ).
(33)
Because this information is current, Rloc does not depend on the global information delay τ . Its magnitude instead reflects how much of the full-information decision can be inferred from local information alone. The stale-global architecture at delay τ is Aglob (τ ) = (Gtτ , . . . , Gtτ ) ,
Gtτ ≜ σ(θs : s ≤ t − τ ) .
(34)
All agents therefore share the same complete but delayed view of the environment. We denote its regret by Rglob (τ ) ≜ R(Aglob (τ )). (35) 10
For reference, define also the open-loop architecture, whose decision is the best constant, stateindependent action. Its optimal action is E[x⋆t ], and its regret is i 1 h R∞ ≜ E ∥x⋆t − E[x⋆t ]∥2Q . (36) 2 Throughout this section we assume R∞ > 0, excluding the trivial case in which the clairvoyant decision is deterministic.
3.2
General Predictability Characterization
The stale-global architecture is simple because all agents share one common information set. Its admissible policy space is therefore L2 (Ω, Gtτ , P; Rd ). Since Q is deterministic, the Q-weighted orthogonal projection of x⋆t onto this space is simply its conditional expectation. Proposition 1 (Common-information projection and stale-global prediction error). For any stationary environment satisfying the assumptions of Section 2, u⋆glob (τ ) = E [x⋆t | Gtτ ] ,
(37)
and i 1 h Rglob (τ ) = E ∥x⋆t − E [x⋆t | Gtτ ]∥2Q 2 1 = E [tr (Q Cov (x⋆t | Gtτ ))] . 2
(38) (39)
Proof. By Lemma 1, the optimal stale-global policy is the orthogonal projection of x⋆t onto the set of Gtτ -measurable joint actions. For any Gtτ -measurable v, h i E (x⋆t − E[x⋆t | Gtτ ])⊤ Qv = 0, so this projection is the conditional expectation in (37). Substitution into (27) yields (38), and the conditional-covariance form follows from the standard mean-square prediction identity. Proposition 1 reveals that the effect of staleness is fundamentally a prediction problem. Delay matters only to the extent that information available at time t − τ fails to predict the decision that would be optimal at time t. This motivates the normalized temporal unpredictability function h i ⋆ − E[x⋆ | G τ ]∥2 E ∥x t t t Q Rglob (τ ) h i . = ηT (τ ) ≜ (40) R∞ E ∥x⋆ − E[x⋆ ]∥2 t
t
Q
Thus Rglob (τ ) = ηT (τ )R∞ .
(41)
The function ηT (τ ) ∈ [0, 1] measures the fraction of decision-relevant uncertainty that remains when the newest available global state is τ old. It is nondecreasing in τ because the available σ-algebra becomes smaller as the delay increases. At τ = 0, ηT (0) = 0, 11
(42)
since the current complete state is available and x⋆t is therefore known exactly. For mixing processes, ηT (τ ) → 1 as τ → ∞; a stationary process with perfectly persistent latent components is the exception. Similarly, define the fresh-local omission share ηS (0) ≜
Rloc ∈ [0, 1]. R∞
(43)
The quantity ηS (0) measures the fraction of decision-relevant variation omitted by fresh-local information, with prediction error weighted by its consequence for the decision objective. The captured share is 1 − ηS (0). These two quantities give an immediate general crossover characterization. Theorem 1 (General freshness–locality crossover). For any stationary environment satisfying the assumptions above, Rloc < Rglob (τ ) ⇐⇒ ηT (τ ) > ηS (0). (44) Proof. By definitions (41) and (43), Rglob (τ ) = ηT (τ )R∞ ,
Rloc = ηS (0)R∞ .
Since R∞ > 0, comparing the two gives (44). Although Theorem 1 is algebraically simple, it separates the two fundamentally different sources of architectural loss. The quantity ηS (0) measures what is lost by restricting information spatially, while ηT (τ ) measures what is lost by allowing global information to age in time. Fresh-local information is preferable precisely when the temporal information lost by waiting for a global view exceeds the decision-relevant information unavailable locally.
3.3
Gauss–Markov Temporal Model
We now specialize the temporal unpredictability function to obtain a closed form. Suppose the environment follows the stationary multivariate Ornstein–Uhlenbeck process r 1 2 1/2 dθt = − θt − θ̄ dt + Σ dWt , (45) T T ∞ where T > 0 is the temporal coherence time, θ̄ is the stationary mean, and Σ∞ is the stationary covariance matrix. The process may exhibit arbitrary instantaneous cross-correlation among agents through Σ∞ ; the simplifying assumption is that all temporal modes share the same decay time T . Assume further that b(θ) = Bθ + b0 , (46) so that the full-information decision is affine: x⋆ (θ) = Q−1 Bθ + Q−1 b0 . Let ρ≜ denote the dimensionless staleness ratio.
12
τ T
(47)
(48)
The OU Markov property gives [18, eqs. (8)–(10), (14)] E [θt | Gtτ ] = θ̄ + e−ρ θt−τ − θ̄ , Cov (θt | Gtτ ) = 1 − e−2ρ Σ∞ .
(49) (50)
We can therefore evaluate the stale-global regret exactly. Theorem 2 (Exact stale-global regret). Under (45)–(46), Rglob (τ ) = 1 − e−2τ /T R∞ ,
(51)
1 ⊤ −1 tr B Q B Σ∞ . 2
(52)
where R∞ = Equivalently,
ηT (τ ) = 1 − e−2τ /T .
(53)
h
(54)
Proof. From (47) and (50), Cov (x⋆t | Gtτ ) = Q−1 B
i 1 − e−2τ /T Σ∞ B ⊤ Q−1 .
Substituting into (39) gives Rglob (τ ) =
1 1 − e−2τ /T tr B ⊤ Q−1 BΣ∞ , 2
(55)
which yields (51). The same trace expression without the exponential factor is exactly the open-loop regret (36), giving (52). The result has a simple interpretation. When τ ≪ T , the environment changes little over the information delay and τ τ . (56) Rglob (τ ) = 2 R∞ + o T T When τ ≫ T , the delayed global history becomes essentially uninformative about the current decision and Rglob (τ ) → R∞ . Staleness therefore matters relative to the timescale on which decision-relevant state changes. For any fixed τ > 0, the same identity gives Rglob (τ ) → R∞ as T ↓ 0 and Rglob (τ ) → 0 as T → ∞.
3.4
Fresh-Local versus Stale-Global Crossover
Combining Theorem 1 with (53) yields the first main architectural result. Theorem 3 (Fresh-local versus stale-global crossover). Under the Gauss–Markov and affinedecision assumptions of Theorem 2, suppose 0 ≤ ηS (0) < 1. Fresh-local information strictly outperforms stale-global information if and only if τ 1 1 > log . T 2 1 − ηS (0)
13
(57)
Equivalently, defining ρ⋆ ≜
1 1 log , 2 1 − ηS (0)
(58)
we have Rloc < Rglob (τ )
ρ > ρ⋆ .
⇐⇒
(59)
For ηS (0) = 1, define ρ⋆ = +∞ in the extended-real sense; fresh-local information then never strictly outperforms stale-global information at a finite delay. Proof. By Theorem 1, fresh-local wins exactly when 1 − e−2ρ > ηS (0). Equivalently, e−2ρ < 1 − ηS (0), which gives (57). The limiting cases are informative. If ηS (0) = 0, fresh-local information is sufficient to reproduce the full-information decision, and the fresh-local architecture dominates for every τ > 0. If ηS (0) = 1, local information provides no improvement over open loop, and stale-global information is never strictly worse at any finite delay. More generally, Theorem 3 identifies two dimensionless determinants of the architecture choice. The staleness ratio τ /T measures the loss of temporal relevance incurred by global information, while ηS (0) measures the share of decision-relevant variation omitted locally. Within the comparison of the two pure benchmarks, fresh-local has lower regret precisely when the two loss fractions cross.
3.5
Role of Decision Coupling
The fresh-local omission share ηS (0) depends jointly on the spatial statistics of the environment and on the coupling structure of the decision problem. To isolate the role of decision coupling, consider a symmetric shared-resource model with n ≥ 2 scalar actions, f (x, θ) =
n X q i=1
2
x2i − θi xi
γ + 2
n X
!2 xi
,
q > 0, γ ≥ 0.
(60)
i=1
Here q determines the local curvature or stiffness of each decision, while γ determines the strength of the shared coordination penalty. In the notation of Section 2, Q = qI + γ11⊤ ,
b(θ) = θ.
(61)
Assume that the state components θi are independent, mean-zero stationary OU processes with common variance σ 2 and common coherence time T .
14
The full-information action is θi x⋆i = − q
n
X γ θj . q(q + γn)
(62)
j=1
Thus, even though the linear reward θi xi is local, coupling through the common resource makes agent i’s optimal action depend on the states of all other agents. Proposition 2 (Local explainability under aggregate coupling). For the model (60), the optimal fresh-local policy is θi uloc , (63) i = q+γ and nσ 2 (q + γ(n − 1)) , 2q(q + γn) nσ 2 γ 2 (n − 1) . Rloc = 2q(q + γ)(q + γn) R∞ =
(64) (65)
Consequently, the fresh-local omission share is ηS (0) =
γ 2 (n − 1) . (q + γ) (q + γ(n − 1))
(66)
Proof. Let µj = E[uj ]. Because the component processes are independent, for j ̸= i the full local history satisfies loc E[uj | Gi,t ] = E[uj ] = µj , loc -measurable. The conditional normal equation (30) therefore gives while θi,t is Gi,t
(q + γ)ui + γ
X
µj − θi,t = 0.
j̸=i
Taking expectations yields qµi + γ
X
µj = 0.
j
P
Summing over i implies j µj = 0, hence every µi = 0 and ui = θi,t /(q+γ), which proves (63). This argument uses independence, zero means, square integrability, and strict convexity; Gaussianity is not required for the local-policy formula. The expressions (64) and (65) follow by substituting the full-information and local policies into the quadratic regret expression (27). Substituting them into (43) and simplifying yields (66). Combining Proposition 2 with Theorem 3 gives an explicit threshold: 1 (q + γ) (q + γ(n − 1)) ρ (γ, n) = log . 2 q(q + γn)
(67)
τ > ρ⋆ (γ, n). T
(68)
⋆
Therefore Rloc < Rglob (τ )
⇐⇒
15
As n → ∞, 1 γ ρ (γ, n) −→ log 1 + . 2 q ⋆
(69)
The local–global architecture choice therefore collapses, in the large-system limit, to a comparison between two dimensionless quantities: τ T |{z}
and
information staleness
1 γ log 1 + . 2 q | {z }
(70)
value of coordination
When γ = 0, the decisions decouple and ηS (0) = 0: current local information is sufficient, so any positive global delay makes the fresh-local architecture preferable. As γ increases, each agent’s optimal action depends more strongly on remote states, ηS (0) increases, and the crossover moves to larger τ /T . Thus stronger coordination coupling makes the system willing to tolerate increasingly stale-global information before abandoning global coordination in favor of fresh-local decisions. The result provides the first endpoint characterization of the broader freshness–scope problem. The next sections move beyond the binary local–global comparison by allowing the information radius itself to vary, and ask how spatial predictability and temporal staleness jointly determine the optimal coordination scale.
4
Spatial and Temporal Predictability
The preceding section compared the two endpoints of the freshness–scope tradeoff: fresh-local information and stale-global information. We now introduce a family of intermediate architectures in which each agent may use information from within a spatial radius r. The purpose of this section is to characterize the two forms of predictability that determine the value of such an architecture. Temporal predictability determines how much decision-relevant information is lost while observations are being communicated. Spatial predictability determines how much of the globally optimal decision can be inferred without observing the entire network.
4.1
Spatial Structure of the Environment
We first describe the spatial statistics that determine what nearby observations reveal about remote state. For simplicity of exposition, first suppose that each node has a scalar local state, so that θt = (θ1,t , . . . , θn,t )⊤ . The vector-valued case follows by replacing scalar covariances with block covariances. Let θ̄ ≜ E[θt ], ΣS ≜ Cov(θt ) ⪰ 0
(71)
denote the stationary mean and spatial covariance. The entry (ΣS )ij measures the instantaneous statistical dependence between the states associated with nodes i and j. On a path, line, or Euclidean spatial domain, a canonical covariance model is dG (i, j) 2 Cov (θi,t , θj,t ) = σ exp − , (72) ℓs 16
where ℓs > 0 is a spatial correlation length. Larger ℓs means that states remain strongly correlated over greater network distance. In one spatial dimension, (72) is the exponential, or Ornstein– Uhlenbeck, covariance model [19]. For an arbitrary graph, a convenient alternative is to define the covariance spectrally through the graph Laplacian LG . For example, −ν ΣS = σ 2 κ2s I + LG , κs > 0, ν > 0, (73) is positive semidefinite and places more state energy in low graph-frequency modes. Such constructions are closely related to graph signal smoothness and Gaussian Markov random-field models [20, 21]. A basic graph-smoothness statistic is h i EG ≜ E (θt − θ̄)⊤ LG (θt − θ̄) = tr (LG ΣS ) .
(74)
For an undirected weighted graph, (θt − θ̄)⊤ LG (θt − θ̄) =
2 1X wij (θi,t − θ̄i ) − (θj,t − θ̄j ) . 2
(75)
i,j
Small EG therefore indicates that nearby nodes tend to have similar states. Neither ℓs nor EG by itself determines architecture regret. A state component may be difficult to infer spatially yet have little effect on the desired decision. We therefore introduce a decisionrelevant spatial predictability function below.
4.2
Temporal Structure
Section 3 defined the normalized temporal unpredictability function ηT (τ ) =
Rglob (τ ) . R∞
(76)
It measures the fraction of full-information decision variation that cannot be predicted from global information whose age is τ . The general framework does not require a particular temporal process. The function ηT (τ ) may be obtained empirically or calculated from any specified stochastic model. For the exact results in this and the following section, however, we use the common-rate Gauss–Markov model r 1 2 1/2 dθt = − θt − θ̄ dt + Σ dWt . (77) T T S Its stationary space-time covariance is Cov (θt , θs ) = e−|t−s|/T ΣS .
(78)
Thus ΣS specifies arbitrary instantaneous spatial dependence, while T specifies a common temporal coherence time. Equation (78) is a separable space-time model. Separability is used here because it produces an exact composition law between spatial omission and temporal staleness. More general nonseparable covariance models can be accommodated by working directly with a joint prediction-error function, although the factorization below will not generally remain exact; see, for example, [22]. 17
4.3
Joint Spatio-Temporal Information
For radius r, define the zero-delay information available to agent i at time t as (r)
Ii,t ≜ σ(θj,t : j ∈ Nr (i)) .
(79)
The corresponding zero-delay radius-r architecture is (r) (r) Ar,0 ≜ I1,t , . . . , In,t .
(80)
A delayed synchronized-snapshot radius-r architecture provides the same type of information, shifted backward by τ : (r) (r) Ar,τ ≜ I1,t−τ , . . . , In,t−τ . (81) Such common-time snapshots arise, for example, when a coordinated optimizer requires a temporally consistent system estimate at a common decision epoch, or when asynchronous measurements cannot be combined without explicit time alignment into a coherent state vector. For comparison, define the distinct history information set (r) Iei,t ≜ σ(θj,s : j ∈ Nr (i), s ≤ t) .
(82)
The principal architecture-design claims use the synchronized-snapshot family (81). Theorem 4 also applies to the consistently shifted history family, with its own zero-delay omission function. Let Ft ≜ σ(θs : s ≤ t) (83) and define the centered admissible-policy subspace S0 (A) ≜ {u ∈ S(A) : E[u] = 0}.
(84)
Let S−τ shift a stationary sample path backward by τ , and define the time-shift operator by (Uτ Z)t (ω) ≜ Zt (S−τ ω).
(85)
Stationarity makes Uτ an isometry under the Q-weighted L2 norm. A delayed architecture family is shift-consistent if Uτ S0 (Ar,0 ) = S0 (Ar,τ ). (86) The synchronized-snapshot family and the consistently shifted history family both satisfy this property. We write Pr,τ for the Q-orthogonal projection onto S0 (Ar,τ ). Lemma 2 (Snapshot–history sufficiency under a common temporal rate). Suppose the process is Gaussian and Cov(θt , θs ) = e−|t−s|/T ΣS . (87) For an observed component set N with complement N̄ and nonsingular ΣN N , the current unobserved state is conditionally independent of the entire past observed history given the current observed snapshot: (88) θN̄ ,t ⊥⊥ σ(θN,s : s < t) | θN,t . In particular, for every s < t, Cov θN̄ ,t , θN,s | θN,t
−(t−s)/T = e−(t−s)/T ΣN̄ N − ΣN̄ N Σ−1 ΣN N = 0. NN e
18
(89)
For affine current-state decisions, any team policy based on neighborhood histories can be replaced by a policy based only on the corresponding current neighborhood snapshots without increasing quadratic team cost. Thus the two variants have the same zero-delay architecture regret, including for coupled Q. Proof. For arbitrary s1 , . . . , sk < t, let Z = (θN,s1 , . . . , θN,sk ). Applying the Gaussian conditional covariance formula blockwise gives Cov(θN̄ ,t , Z | θN,t ) = 0,
(90)
because every block is exactly the cancellation displayed in (89). Joint Gaussianity therefore makes θN̄ ,t conditionally independent of every finite collection of past observed values given θN,t . Cylinder events generate the observed-history sigma-algebra (equivalently, use rational times and path separability), so a monotone-class argument yields (88). For the team-level claim, let Hi be agent i’s neighborhood history and Ii its current neighborhood snapshot. Given any history-feasible square-integrable policy u, define vi ≜ E[ui | Ii ].
(91)
The conditional independence just proved implies E[ui | θt ] = E[ui | Ii ] = vi ; hence v = E[u | θt ] componentwise and v is snapshot-feasible. Because x⋆t and v are measurable with respect to θt , conditioning on θt makes the cross term vanish and gives E∥x⋆t − u∥2Q = E∥x⋆t − v∥2Q + E∥u − v∥2Q .
(92)
Thus Rao–Blackwellization to current snapshots cannot increase team cost. Since snapshot policies are already history-feasible, the optimal regrets are equal. Remark 2 (Scope of snapshot–history sufficiency). If ΣN N is singular, the same statement uses the Moore–Penrose inverse on the covariance support. Heterogeneous temporal rates and nonseparable space–time processes need not have this sufficiency property. Define the normalized zero-delay spatial omission function ηS (r) ≜
R(Ar,0 ) . R∞
(93)
The quantity ηS (r) is the fraction of decision-relevant variation that cannot be captured using current information within radius r. Since zero-delay radius architectures are nested, r2 ≥ r1
=⇒
S(Ar1 ,0 ) ⊆ S(Ar2 ,0 ),
(94)
and therefore 0 ≤ ηS (r2 ) ≤ ηS (r1 ) ≤ 1.
(95)
If radius zero observes precisely the local current component, then ηS (0) =
Rloc . R∞
(96)
Under the common-rate Gaussian-affine assumptions, Lemma 2 identifies this snapshot endpoint with the history-based fresh-local benchmark; the identification is specific to those assumptions. On a finite graph with diameter D, complete current state is available at r = D, so ηS (D) = 0. 19
(97)
4.4
Decision-Relevant Smoothness
Suppose, as in (10), that b(θ) = Bθ + b0 .
(98)
Then x⋆t = Kθt + k0 ,
K ≜ Q−1 B,
k0 ≜ Q−1 b0 .
(99)
The matrix K is the decision-sensitivity operator: its block Kij describes how the state at node j influences the full-information decision at node i. The covariance of the clairvoyant decision is Σx ≜ Cov(x⋆t ) = KΣS K ⊤ .
(100)
Thus spatial state modes that lie approximately in the null space of K have little decision value, even if they are statistically unpredictable. By contrast, uncertainty in a state mode strongly amplified by K can dominate architecture regret. The function ηS (r) in (93) incorporates all three relevant ingredients: state spatial statistics
+
decision sensitivity
+
information geometry.
(101)
It is therefore more informative for architecture design than a raw correlation length alone. The next result shows that, under the common-rate Gauss–Markov model, ηS (r) and temporal unpredictability combine in an exact law. Theorem 4 (Spatio-temporal predictability composition law). Suppose the environment follows (77), the full-information decision is affine as in (99), and the radius-dependent information family is shift-consistent. Then, for every radius r and delay τ , R(Ar,τ ) = 1 − e−2τ /T R∞ + e−2τ /T R(Ar,0 ).
(102)
R(Ar,τ ) = 1 − e−2τ /T (1 − ηS (r)) . R∞
(103)
Equivalently,
In terms of the temporal unpredictability ηT (τ ) = 1 − e−2τ /T , ηST (r, τ ) = ηT (τ ) + (1 − ηT (τ )) ηS (r), where ηST (r, τ ) ≜
(104)
R(Ar,τ ) . R∞
Proof. Let x̄⋆ ≜ E[x⋆t ],
Xt ≜ x⋆t − x̄⋆ .
(105)
Because constant policies are feasible under every architecture, architecture regret depends only on the centered process Xt . Under (77) and (99), Xt = e−τ /T Xt−τ + ξt,τ , (106)
20
where and
E [ξt,τ | Ft−τ ] = 0
(107)
E ∥ξt,τ ∥2Q = 1 − e−2τ /T E ∥Xt ∥2Q .
(108)
Let Pr,τ be the Q-orthogonal projection defined above onto the centered policy subspace induced by Ar,τ . Every policy in this subspace is measurable with respect to Ft−τ . Hence (107) implies that ξt,τ is orthogonal to the entire admissible subspace. By linearity of orthogonal projection, Pr,τ Xt = e−τ /T Pr,τ Xt−τ .
(109)
Xt − Pr,τ Xt = e−τ /T (Xt−τ − Pr,τ Xt−τ ) + ξt,τ .
(110)
Consequently,
The two terms on the right-hand side are orthogonal, so E ∥Xt − Pr,τ Xt ∥2Q = e−2τ /T E ∥Xt−τ − Pr,τ Xt−τ ∥2Q + E ∥ξt,τ ∥2Q .
(111)
By stationarity and shift consistency, 1 E ∥Xt−τ − Pr,τ Xt−τ ∥2Q = R(Ar,0 ), 2
(112)
Indeed, time translation maps the zero-delay centered policy subspace isometrically onto the corresponding subspace at t − τ , so the projection residual has the same distribution and weighted norm. Moreover, 1 E ∥Xt ∥2Q = R∞ . (113) 2 Substitution of (108) into (111) gives (102); normalization by R∞ gives (103) and (104). Theorem 4 gives a precise meaning to the two architectural errors. The term 1 − e−2τ /T R∞ is the temporal innovation that no information available at time t − τ can predict. The remaining fraction e−2τ /T of the current decision is temporally predictable, but only the fraction 1 − ηS (r) of that component can be reconstructed from radius-r information. Spatial information can only help with the portion of the current decision that remains temporally predictable when it arrives.
5
Optimal Coordination Scale
We now make information radius an architectural design variable. For each r, an agent may base its action on observations available within radius r, but those observations arrive after latency τ (r). The objective is to select the spatial scope that minimizes architecture regret. On a graph, admissible radii are generally integers. We first analyze the continuous relaxation r ∈ [0, D], where D is the maximum allowed information radius. On a finite graph, the exact optimizer is obtained by minimizing over the admissible integer radii. Rounding the continuous optimum to a neighboring integer is justified only when the interpolated regret curve is unimodal. 21
5.1
Radius-Dependent Information Architectures
We now formalize the radius family and the latency law that couples scope to age. Let τ : [0, D] → [0, ∞)
(114)
be a nondecreasing information-propagation law satisfying τ (0) = 0.
(115)
Architecture Ar,τ (r) provides each agent with information from a synchronized radius-r snapshot delayed by τ (r). Proposition 3 (Nested information cannot create a strict interior minimum without explicit cost). If radius-indexed information sets satisfy Gi (r1 ) ⊆ Gi (r2 ) for every agent i whenever r1 ≤ r2 , then R(r2 ) ≤ R(r1 ).
(116)
Therefore, without an explicit information-acquisition cost, a strict interior minimum requires a non-nested architecture such as the synchronized-snapshot family studied here. Proof. The nesting assumption implies S(A(r1 )) ⊆ S(A(r2 )). Information monotonicity, Eq. (23), then gives the claim. The synchronized-snapshot family is non-nested in radius: increasing r replaces a narrower, fresher snapshot with a broader, older one, and this replacement is what permits a strict interior minimum without an explicit information-acquisition cost. Define R(r) ≜ R Ar,τ (r) . (117) By Theorem 4, h i R(r) = R∞ 1 − e−2τ (r)/T (1 − ηS (r)) .
(118)
On a finite graph, moving from radius r to r + 1 strictly improves performance exactly when (119) e−2τ (r+1)/T 1 − ηS (r + 1) > e−2τ (r)/T 1 − ηS (r) , or, provided 1 − ηS (r) > 0, equivalently 2 τ (r + 1) − τ (r) 1 − ηS (r + 1) log > . 1 − ηS (r) T
(120)
The multiplicative comparison (119) remains the primary statement when 1−ηS (r) = 0. Increasing the radius by one hop is beneficial precisely when the discrete proportional gain in decision-relevant spatial information exceeds the freshness loss incurred over that hop. At radius zero, R(0) = R∞ ηS (0) = Rloc . (121) On a finite graph, if D = diam(G) and radius D provides the entire state, then R(D) = R∞ 1 − e−2τ (D)/T = Rglob (τ (D)).
(122)
Under the common-rate Gaussian-affine model, the snapshot endpoints in (121)–(122) are equivalent to the Section 3 benchmarks by Lemma 2. 22
5.2
Spatial-Omission versus Temporal-Staleness Tradeoff
Equation (118) can be written as R(r) = RT (r) + RS (r),
(123)
RT (r) ≜ 1 − e−2τ (r)/T R∞ ,
(124)
RS (r) ≜ e−2τ (r)/T ηS (r)R∞ .
(125)
where
The temporal term increases with information radius whenever τ (r) does. The zero-delay spatial omission fraction ηS (r) decreases with radius, although its contribution is discounted by the temporal relevance factor e−2τ (r)/T . To characterize the optimum, define the marginal temporal decay rate mT (r) ≜
2τ ′ (r) T
(126)
and the proportional marginal spatial information gain mS (r) ≜ −
ηS′ (r) . 1 − ηS (r)
(127)
The latter quantity is the rate at which the spatially explained share 1 − ηS (r) grows: mS (r) =
d log (1 − ηS (r)) . dr
(128)
Theorem 5 (Continuous marginal-balance characterization). Under the assumptions of Theorem 4, additionally assume that τ (r) and ηS (r) are continuously differentiable on [0, D], with τ ′ (r) ≥ 0,
ηS′ (r) ≤ 0,
0 ≤ ηS (r) < 1.
Then R′ (r) = R∞ e−2τ (r)/T (1 − ηS (r)) [mT (r) − mS (r)] .
(129)
Consequently, every interior optimal radius r⋆ ∈ (0, D) satisfies −
ηS′ (r⋆ ) 2τ ′ (r⋆ ) = . 1 − ηS (r⋆ ) T
(130)
Suppose additionally that mS (r) is strictly decreasing and mT (r) is nondecreasing. Then R(r) is unimodal and exactly one of the following holds: 1. If mS (0) ≤ mT (0),
(131)
mS (D) ≥ mT (D),
(132)
then r⋆ = 0. 2. If then r⋆ = D. 23
3. Otherwise, there is a unique r⋆ ∈ (0, D) satisfying (130). Proof. Differentiating (118) gives ′ R′ (r) −2τ (r)/T 2τ (r) ′ =e (1 − ηS (r)) + ηS (r) R∞ T = e−2τ (r)/T (1 − ηS (r)) [mT (r) − mS (r)] ,
(133)
which proves (129) and the interior first-order condition. Under the additional assumptions, the difference mT (r) − mS (r) is strictly increasing. It can therefore change sign at most once. If it is nonnegative at r = 0, regret is nondecreasing and radius zero is optimal. If it is nonpositive at r = D, regret is nonincreasing and radius D is optimal. Otherwise it crosses zero exactly once, from negative to positive, yielding a unique interior minimizer. Within the common-rate synchronized-snapshot model, Theorem 5 gives the marginal-balance characterization
expand the information radius until
marginal spatial value = marginal freshness loss.
(134)
The first-order condition applies to any differentiable latency law, but uniqueness requires the stated single-crossing assumptions. Concave aggregation laws can make mT (r) decrease with radius, in which case multiple crossings or boundary optima cannot be excluded; convex congestion laws act in the opposite direction. It also gives comparative statics without committing to a specific covariance model. Under the single-crossing conditions of the theorem: 1. Increasing T lowers mT (r) pointwise and therefore moves the optimum toward a larger radius. 2. Increasing communication latency, for example by replacing τ (r) with aτ (r) for a > 1, raises mT (r) and moves the optimum toward a smaller radius. 3. Any increase in decision coupling that raises mS (r) pointwise makes additional spatial information more valuable and moves the optimum toward a larger radius. 4. Any increase in spatial predictability that lowers mS (r) pointwise makes additional radius less valuable and moves the optimum toward a smaller radius. The final two statements are single-crossing conditions on the decision-relevant function ηS (r); they are not automatic consequences of changing an arbitrary raw covariance parameter.
5.3
An Exponential Spatial-Predictability Regime
A simple but useful closed form is obtained when the zero-delay spatial omission fraction decays exponentially: 0 < η0 < 1. (135) ηS (r) = η0 e−2r/ℓc ,
24
Here η0 is the decision variation not explained by radius-zero information, and ℓc is a decisionrelevance length: larger ℓc means that decision-relevant information remains distributed over a greater distance. Assume information propagates at effective speed v > 0, so that r . v
(136)
LT ≜ vT
(137)
τ (r) = The distance
is the distance over which information can propagate during one temporal coherence time. Proposition 4 (Optimal radius under exponential spatial predictability). Under (135) and (136), the unique optimal radius on [0, ∞) is LT ℓc log η0 1 + r = , 2 ℓc + ⋆
(138)
where [z]+ ≜ max{z, 0}. If the architecture imposes a maximum radius D, the optimizer is ⋆ rD = min {D, r⋆ } .
(139)
Proof. Under (135), mS (r) =
2η0 e−2r/ℓc , ℓc 1 − η0 e−2r/ℓc
(140)
2 . LT
(141)
ℓc . ℓc + L T
(142)
while mT (r) = The marginal-balance equation gives ⋆
η0 e−2r /ℓc =
Solving for r⋆ yields (138). If the logarithm is nonpositive, R′ (0) ≥ 0 and radius zero is optimal. Unimodality follows from Theorem 5. The dimensionless form is r⋆ /ℓc = 21 [log(η0 [1 + LT /ℓc ])]+ . Thus optimal scope depends on the ratio between the temporal propagation length LT and the decision-relevance length ℓc , together with the fraction η0 of decision variation unavailable locally.
5.4
A Canonical Spatio-Temporal Field Model
We now derive the exponential form (135) from a concrete stochastic field and decision model. Consider a dense one-dimensional network, represented by locations z ∈ R. This is also the continuum approximation of a sufficiently long path graph away from its boundaries. Let θ(z, t) be a zero-mean Gaussian field with covariance |z − z ′ | |t − s| ′ 2 E θ(z, t)θ(z , s) = σ exp − exp − . (143) ℓs T The parameter ℓs is the spatial correlation length, and T is the temporal coherence time. 25
Suppose the full-information decision at location z is the exponentially weighted spatial aggregate Z ∞
x⋆ (z, t) =
−∞
wℓc (u)θ(z + u, t) du,
(144)
where
|u| 1 exp − . (145) wℓc (u) ≜ 2ℓc ℓc The decision-relevance length ℓc controls how broadly remote states influence the desired local action. For example, (144) is generated pointwise by the quadratic tracking-loss density q fz (x, θ) = [x(z) − x⋆ (z, θ)]2 , q > 0. (146) 2 All regrets in this translation-invariant infinite-line model are understood per unit length, equivalently as the limit of spatially averaged regret on expanding finite intervals. This convention avoids assigning an infinite total cost to a stationary field on R. Equivalently, in the notation of Section 2, Q = qI and b(θ) = qKℓc θ, where Kℓc is convolution with wℓc . Thus the canonical model places nonlocal decision dependence in b(θ). Direct action coupling through a non-diagonal Q changes the specific form of ηS (r) but not the general composition law or Theorem 5. At zero delay, the radius-r observation at location z is Or (z, t) ≜ σ (θ(z + u, t) : |u| ≤ r) .
(147)
Proposition 5 (Spatial omission in the canonical line model). For the field (143), decision (144), and observation (147), the normalized zero-delay spatial omission error is ℓc 2r ηS (r) = exp − . (148) ℓc + 2ℓs ℓc In particular, η0 = ηS (0) =
ℓc . ℓc + 2ℓs
(149)
1 . ℓs
(150)
Proof. By spatial stationarity, take z = 0. Let κ≜
1 , ℓc
λ≜
At a fixed time, θ(·, t) is a stationary one-dimensional Ornstein–Uhlenbeck field and hence is Markov in the spatial coordinate [21]. Conditioned on the field over [−r, r], uncertainty in the two exterior regions is carried by independent left and right spatial innovations. For the right exterior region, the conditional residual covariance at distances u, v ≥ 0 beyond the boundary r is h i σ 2 e−λ|u−v| − e−λ(u+v) . (151) The corresponding contribution to the conditional variance of x⋆ (0, t) is Z Z κ2 σ 2 −2κr ∞ ∞ −κ(u+v) e e VR (r) = 4 0 0 h i × e−λ|u−v| − e−λ(u+v) du dv =
σ 2 κλ e−2κr . 4(κ + λ)2 26
(152)
The left exterior region contributes the same amount, so Var (x⋆ (0, t) | Or (0, t)) =
σ 2 κλ e−2κr . 2(κ + λ)2
(153)
A direct covariance calculation gives the unconditional variance Var (x⋆ (0, t)) =
σ 2 κ (2κ + λ) . 2(κ + λ)2
(154)
Since the loss in (146) is squared error, the normalized spatial regret is the ratio of (153) to (154). Therefore λ ℓc ηS (r) = e−2r/ℓc , (155) e−2κr = 2κ + λ ℓc + 2ℓs which proves the claim. Remark 3 (Continuum-field interpretation). The canonical line model can be formulated directly in the L2 Hilbert space of stationary random fields with the per-unit-length inner product, or as the bulk limit of finite periodic lattices whose circumference tends to infinity. Because the canonical loss has Q = qI, the radius-constrained pointwise optimizer is the conditional expectation of x⋆ (z, t) given the radius-r field observation, and per-unit-length regret is the corresponding conditional variance. Proposition 5 follows directly in the field formulation or as this finite-lattice limit. The expression has the expected limiting behavior. If ℓc → 0, the desired action becomes purely local and ηS (r) → 0 for every r ≥ 0. If ℓs → ∞, the state becomes nearly constant over space and a local observation is sufficient to predict the remote field, again driving ηS (r) to zero. Conversely, spatially rough states and long-range decision dependence increase the value of nonlocal information.
5.5
Main Result: Closed-Form Optimal Coordination Radius
Combining Proposition 4 and Proposition 5 yields the second main result. Theorem 6 (Optimal synchronized-snapshot radius in the canonical model). Suppose the state field obeys (143), the full-information decision is given by (144), and information from radius r arrives after r τ (r) = , v > 0. (156) v Let LT = vT (157) be the temporal propagation length. Then the optimal information radius on the infinite line satisfies the dimensionless law r⋆ 1 1 + LT /ℓc = log . (158) ℓc 2 1 + 2ℓs /ℓc + Equivalently, ℓc ℓc + LT log . r = 2 ℓc + 2ℓs + ⋆
(159)
The optimal radius is strictly positive if and only if LT > 2ℓs . Within the interior regime, r⋆ : 27
(160)
1. increases with the temporal coherence time T ; 2. increases with the information-propagation speed v; 3. decreases with the spatial correlation length ℓs ; and 4. increases with the decision-relevance length ℓc . Proof. Substituting η0 =
ℓc ℓc + 2ℓs
and LT = vT into (138) gives r⋆ = ℓ2c [log((ℓc + LT )/(ℓc + 2ℓs ))]+ , which is (159). The logarithm is positive exactly when LT > 2ℓs , proving (160). In the interior regime, ∂r⋆ ℓc = > 0, (161) ∂LT 2(ℓc + LT ) and
ℓc ∂r⋆ =− < 0. ∂ℓs ℓc + 2ℓs
(162)
Since LT = vT , the first derivative establishes monotonicity in both v and T . Finally, when LT > 2ℓs , Z 1 LT ℓc ⋆ r = ds. 2 2ℓs ℓc + s
(163)
The integrand is strictly increasing in ℓc , proving that r⋆ increases with the range over which remote state affects the desired decision. Theorem 6 identifies three physical length scales: ℓs |{z}
spatial correlation
ℓc |{z}
decision relevance
L = vT | T {z }
.
(164)
temporal propagation
The temporal propagation length LT , equivalently a freshness horizon measured in distance, is how far information can travel before the environment substantially changes. The spatial correlation length ℓs measures how much of the remote field can already be predicted locally. The decisionrelevance length ℓc measures how far remote state continues to affect the desired action. The threshold LT > 2ℓs has a direct interpretation in the canonical model. If information cannot propagate beyond approximately two spatial correlation lengths within one coherence time, then the additional state it would reveal is not worth the staleness incurred in obtaining it, and fresh local information is optimal. Once the temporal propagation length exceeds that scale, a positive information radius becomes beneficial. The preferred radius then grows with the excess propagation horizon and with the spatial range of the decision dependence. Two limiting cases sharpen this interpretation. The radius approaches zero as ℓc ↓ 0, because the desired decision becomes local, and it is zero for sufficiently large ℓs (in particular as ℓs → ∞), because remote state is then predictable locally. On a finite graph, Theorem 4 and minimization over integer radii remain exact, with the onehop comparison (119)–(120) as the discrete form of marginal balance; Eq. (159) is the closed-form counterpart on the homogeneous line.
28
6
Numerical Consistency Checks and Finite-Graph Illustrations
This section checks the implementations against the closed-form predictions and then shows how the same quantities behave on finite graphs. Experiments 1 and 2 check temporal prediction and the fresh-local–stale-global endpoints; Experiments 3 and 4 isolate and sweep the continuous radius tradeoff; Experiment 5 separates state statistics from decision-relevant omission; and Experiment 6 evaluates exact integer radii on several noncanonical graph instances.
6.1
Numerical Methodology and Synthetic Environments
The numerical objects are architecture regrets, so each calculation uses the same information structure as its analytical counterpart. We work with an exogenous common-rate Gaussian-affine environment: actions do not alter later states. For x⋆ = Kθ + k0 , with K = Q−1 B, we evaluate R(A) = 21 E ∥x⋆ − u⋆A ∥2Q from conditional covariance and, where appropriate, from independently sampled quadratic losses. The solvers respect the information structure: the common delayed snapshot in Experiment 1 permits a joint conditional mean even for non-diagonal Q; Experiment 2 uses the exact fresh-local rule; Experiments 3 and 4 evaluate the scalar composition; and Experiments 5 and 6 use Q = I so agentwise Gaussian conditional variances are exact. Experiment 5 holds the graph and covariance fixed while changing decision sensitivity. The graph covariances are scaled to average marginal variance one. Pointwise Monte Carlo means use 95% normal confidence intervals. The Experiment 2 boundary intervals apply the delta method to the paired fresh-local/open-loop regret means, so their uncertainty reflects that the two regrets use the same draws. Table 2 summarizes the reference settings. The repository stores fixed configurations, processed tables, selected raw arrays, and graph covariance/distance inputs where applicable, together with run metadata and diagnostics; it is not a complete archive of every random draw. The exact system construction, graph seeds, conditioning safeguards, boundary conventions, radius grids, crossover estimator, and interval calculation are specified in the reproduction documentation. Robustness beyond the common-rate Gaussian-affine model is discussed in Section 8. Table 2: Reference numerical settings. The shared-resource sweep fixes q = 1; graph covariances in Experiments 5–6 are scaled to average marginal variance one. Quantities not being swept retain the experiment-specific values in the configuration files.
6.2
Experiment
principal sweep
fixed setting/size
evaluation
1. Temporal pairs 2. Local–global 3. Scalar radius 4. Canonical map 5. Decision relevance 6. General graphs
τ /T ∈ [0, 3] (25 points) 9 γ/q values in [0, 10] r/ℓc ∈ [0, 2] (1,001 points) (vT /ℓc , ℓs /ℓc ); 72 × 68 ℓc = 0.7, 2, 6 path, ring, grid, geometric
A–D; n = 8, 20, 50, 32 n = 5, 10, 25, 100; q = 1 3 regimes; D/ℓc = 2 r/ℓc ∈ [0, 4]; ∆r = 0.0016 64-node ring; T = 6 64 nodes; T = 5
20,000/point 30,000/setting 50,000/point direct scalar 12,000/check exact covariance
Temporal Scaling and the Local–Global Crossover
The first two experiments test two different consequences of common-rate temporal predictability: the exact stale-global law (51), and the endpoint crossover threshold (67).
29
1.0
Rglob/R∞
0.8 0.6 0.4 1 − e −2τ/T
0.2
System C System D
System A System B
0.0 0.0
0.5
1.0
1.5
2.0
2.5
3.0
staleness τ/T Figure 1: Temporal scaling collapse from exact stationary OU pairs. Markers are Monte Carlo estimates for four systems with different dimensions, spatial covariances, objective matrices, and decision operators; bars are 95% confidence intervals. The solid curve is the exact law 1 − e−2τ /T . Normalization by R∞ removes system-specific scale. Experiment 1 asks whether the normalized regret of a delayed global snapshot depends only on ρ = τ /T , even when spatial covariance, objective coupling, and the decision operator change. We used the four systems in Table 2 and drew 20,000 exact stationary OU pairs (θt−τ , θt ) at each of 25 values of ρ. The first member has covariance ΣS and the second is generated as e−ρ θt−τ + (1 − e−2ρ )1/2 ε, with an independent innovation ε of covariance ΣS . For each pair, the delayed conditional mean was evaluated with the known common-rate coefficient and the quadratic loss was computed in the system’s Q metric. Across the 100 system–parameter combinations, the Monte Carlo curves collapse onto 1−e−2τ /T (Fig. 1). The largest absolute normalized deviation was 0.0155 and the root-mean-square deviation was 0.00471. This supports the exact temporal scaling for these stationary common-rate pairs. Experiment 2 asks whether the finite-n threshold in Eq. (67) describes the point at which fresh local information becomes preferable to a stale global snapshot. We used the shared-resource model Q = qI + γ11⊤ , q = 1, with independent OU components. For each of 36 (n, γ/q) settings, the fresh-local and open-loop regrets were evaluated on the same 30,000 current-state draws. The empirical boundary was then inferred by inserting these two estimates into the OU temporal law, ! b∞ 1 R ⋆ ρb = log . b∞ − R bloc 2 R Accordingly, Fig. 2 reports the temporal-law boundary inferred from paired local/open-loop losses. 30
1.4
1.4
finite n = 100 n→∞
1.2
crossover ρ ⋆
fresh-local
1.2
staleness τ/T
1.0
0.8
closed form exact n = 5 exact n = 10
exact n = 25 exact n = 100
1.0 0.8 0.6 0.4 0.2
paired MC ±95% CI
0.0
0.6
0.01
ρ̂ ⋆ − ρ ⋆
0.4
0.2 stale-global
0.00
−0.01
0.0 10
−1
10
0
10
1
zero coupling: residual = 0
10−1
100
101
γ/q
coupling γ/q
Figure 2: Fresh-local versus stale-global crossover. Left: the phase map uses the exact finiten = 100 boundary (solid) and the large-n limit (dashed), Eqs. (67) and (69). Upper right: exact theory curves and inferred Monte Carlo boundaries (crosses) for four system sizes, with paired delta-method 95% intervals. Lower right: empirical minus theoretical boundary, on a scale that reveals the sampling differences and their uncertainty. The zero-coupling setting is evaluated at ρ⋆ = 0 but omitted from the logarithmic positive-coupling axes. The mean absolute boundary error was 0.00111 and the maximum was 0.00557. Finite size matters most under strong coupling: at n = 5 and γ/q = 10, the exact threshold is 0.109 below the large-system limit. As n grows, the inferred boundary approaches 12 log(1 + γ/q). Stronger decision coupling therefore makes older global information worth tolerating in this independent-component model.
6.3
Emergence of an Optimal Coordination Radius
Experiment 3 illustrates why expanding the information radius can help at first and hurt later: reduced spatial omission competes with increasing staleness. We used ηS (r) = η0 e−2r/ℓc and τ (r) = r/v on 0 ≤ r ≤ D = 2, with local (η0 , vT ) = (0.25, 0.2), interior (0.7, 4), and maximum allowed radius (0.9, 100) regimes (all with ℓc = 1). At each radius, independent standard-normal draws were scaled by the two prescribed component variances and added before squaring. This is a scalar loss check: no spatial field, graph, or neighborhood conditional distribution is sampled. The grid minima were 0, 0.626, and 2; the corresponding analytical minima after imposing the cap were 0, 0.62638, and 2. At the interior grid minimum, the marginal-slope gap was |mS (r̂⋆ ) − mT (r̂⋆ )| = 4.77 × 10−4 , and the largest Monte Carlo deviation from the prescribed scalar curve was about 0.020. The maximum allowed radius case is not a global endpoint: its unconstrained minimizer lies beyond D, while ηS (D) = η0 e−2D/ℓc > 0. Thus the cap still omits decisionrelevant state. The three curves distinguish immediate dominance of staleness, an interior balance, and a benefit from expansion throughout the allowed range.
31
normalized regret component
local η0 = 0.25, vT/ℓc = 0.2
interior η0 = 0.7, vT/ℓc = 4
maximum allowed radius η0 = 0.9, vT/ℓc = 100
1.0 0.8 total R/R∞ temporal spatial theory r ⋆ (capped)
0.6 0.4
grid minimum scalar variance check
0.2 0.0 0.0
0.5
1.0
1.5
dimensionless radius r/ℓc
2.0
0.0
0.5
1.0
1.5
dimensionless radius r/ℓc
2.0
0.0
0.5
1.0
1.5
2.0
dimensionless radius r/ℓc
Figure 3: Emergence of local, interior, and maximum allowed radius optima under ηS (r) = η0 e−2r/ℓc and τ (r) = r/v. Curves show total regret and its temporal and residual-spatial terms from Eq. (123); open markers with 95% intervals are independent scalar-innovation estimates. Dashdotted lines and crosses mark the analytical constrained and grid minimizers. At the maximum allowed radius, positive omission remains: ηS (D) > 0.
6.4
Scaling Laws and Regime Map
Experiment 4 asks whether the canonical closed form predicts both the zero-radius region and the interior radius over a two-parameter sweep, rather than only at selected examples. We swept 4,896 points over vT /ℓc ∈ [0.1, 20] and ℓs /ℓc ∈ [0.05, 5], using ℓc = 1 as the unit of length. Independently of the closed form, every point was minimized by direct scalar evaluation of Eq. (103) on the same fixed grid rj = 0.0016j, j = 0, . . . , 2500, so r ∈ [0, 4] for every parameter pair. No covariance simulation or adaptive radius range was used. Separate one-dimensional sweeps tested the four comparative statics. The gray region in Fig. 4 selects fresh local information; above the threshold vT = 2ℓs , positive scope becomes worthwhile. Moving upward increases the distance information can travel within a coherence time and raises the preferred radius. Moving right increases spatial correlation, making remote observations more redundant and lowering the preferred radius. The slices show these changes as ordinary radius curves. The mean absolute grid-versus-closed-form error was 2.25 × 10−4 and the maximum was 8.00 × 10−4 , below the 0.0016 grid step. Ordinary grid error occurs throughout the positive-radius interior because the minimizer is discretized; it is not confined to the threshold. There were five zero/positive classification disagreements, all adjacent to the threshold where the continuous optimum fell below the first positive grid point. The one-dimensional sweeps were monotone in all four predicted directions: r⋆ increased with T , v, and ℓc , and decreased with ℓs . This validates the scalar search and canonical scaling within the stated range.
6.5
State Predictability versus Decision-Relevant Predictability
Experiment 5 asks whether the architecture depends on which state modes affect the decision, rather than only on raw state correlation. We held a 64-node ring, its Laplacian covariance, T = 6, and the latency law (0.12 time units per hop) fixed, and changed only the decision-sensitivity operator Kij = e−dG (i,j)/ℓc /∥e−dG (i,·)/ℓc ∥2 , using ℓc ∈ {0.7, 2, 6} and Q = I. Exact Gaussian conditioning gives the decision-weighted omission ηS (r) at every graph radius. The raw state-correlation statistic
32
1.25
vT/ℓc
1.00
0.
100
0.75
25
0.50
0.25
grid minimum r = 0
0.00
closed form ℓs/ℓc = 0.099
1.2
ℓs/ℓc = 0.39
1.0
ℓs/ℓc = 1
0.8 0.6 0.4 0.2
fixed-grid minimum
0.0 10−1
100
101
vT/ℓc
share (%)
1.25
positive grid r ̂ ⋆ /ℓc
0.7
0.
5
5
1
101
optimal radius r ⋆ /ℓc
1.4 1.50
40
MAE 0.14 max 0.50
20
exact threshold vT = 2ℓs
10−1 10−1
0 −0.75 −0.50 −0.25
100
0.00
0.25
0.50
0.75
signed grid error /(Δr)
ℓs/ℓc
Figure 4: Dimensionless optimal-radius law from direct scalar minimization. Left: the numerical map, with the numerical zero-radius region shown in gray and the dashed curve marking the theoretical transition vT = 2ℓs . Upper right: selected ℓs /ℓc slices comparing grid minimizers with the closed form in Eq. (158). Lower right: the histogram of (r̂⋆ − r⋆ )/∆r over all 4,896 points; its zero spike includes local-regime points and the remaining grid errors are of order a half-step. was computed separately from the fixed covariance; it is not the same quantity as ηS (r). Each curve is normalized by its own R∞ = 12 tr(KΣS K ⊤ ), so normalized regrets compare unexplained fractions rather than equal absolute cost scales. Increasing ℓc from 0.7 to 2 to 6 increased ηS (0) from 0.0460 to 0.258 to 0.572 and moved the discrete optimum from 1 to 2 to 5 hops, even though the raw state correlation was identical. Independent 12,000-sample conditioning checks differed from the exact omission by at most 2.39 × 10−3 (the largest pointwise standard error was 3.20×10−3 ). Thus changing only decision sensitivity in a fixed environment produced substantially different preferred architectures.
6.6
Finite Graphs beyond the Canonical Geometry
Experiment 6 asks whether the conditional-variance construction and discrete radius comparison remain usable when the geometry is not the homogeneous line assumed by the canonical formula. We considered one 64-node graph of each of a path, ring, two-dimensional grid, and random geometric type; the geometric graph uses the fixed configured seed. For each instance we formed ΣS = (κ2s I + LG )−ν , scaled it to average marginal variance one, and computed every ηS (r) directly from the Gaussian conditional covariance of the desired action given the radius-r neighborhood. The reference setting uses Q = I, T = 5, ℓc = 2, and latency 0.12 per hop. No exponential omission fit or canonical line formula was used, and no graph ensemble average was taken. Under the reference setting, Fig. 6 gives exact integer optima of 2, 2, 3, and 2 hops for the path, ring, grid, and geometric instance, respectively. In these four sampled instances, increasing T over the configured values 1.5, 2.5, 4, 6, 9, 14 changed the optimum from 1 to 3 hops on the path 33
0.6
ℓc = 2.0
0.4
raw state correlation
ℓc = 6.0
0.6 0.5
R(r)/R∞
decision omission ηS(r)
0.5
1.0
ℓc = 0.7
0.7
0.3 0.2
0.4 0.3 0.2
0.1
0.8
0.6
0.4
0.2
0.1 0.0
0.0 0
3
6
9
graph radius r
12
15
0
3
6
9
graph radius r
12
15
0
3
6
9
12
15
graph radius r
Figure 5: Decision relevance with the graph and state covariance held fixed. The decision-weighted omission and normalized composed regret are shown for three row-ℓ2 -normalized operators with Q = I. Raw state correlation is shown in a separate context panel and is identical across operators; it is not an omission estimate. Each regret is normalized by its operator-specific R∞ , and the minimizing integer radius is marked. Changing only the decision-relevance range moves the optimum from 1 to 5 hops. and ring, from 1 to 4 on the grid, and from 1 to 3 on the geometric graph. The sampled ℓc sweep from 0.7 to 6 moved the optima from (0, 0, 1, 1) to (5, 5, 5, 3) in the same graph order. Increasing per-hop latency weakly decreased the optimum, as did decreasing κs , which increases spatial smoothness after variance normalization. These comparative statics describe the four specified graph instances. Since radii are integer-valued, the appropriate comparison is the one-hop regret difference in Eq. (119). For example, the geometric graph’s three-hop regret exceeds its two-hop minimum by only 0.00362R∞ ; the archived neighbor differences quantify such shallow minima.
34
(a) Spatial omission
(b) Regret
0.4
path
grid
ring
geometric
0.5
grid (r ⋆ = 3)
ring (r ⋆ = 2)
geometric (r ⋆ = 2)
0.4
R(r)/R∞
0.3
ηS(r)
path (r ⋆ = 2)
0.2 0.1
0.3
0.2
0.0 0
2
4
6
8
10
12
14
16
0
2
4
graph radius r (c) Coherence-time sweep path
grid
ring
geometric
8
10
12
14
16
(d) Decision-range sweep 5
optimal integer radius r ⋆
optimal integer radius r ⋆
4
6
graph radius r
3
2
1
path
grid
ring
geometric
4 3 2 1 0
1.5 2.5
4.0
6.0
9.0
14.0
sampled coherence time T
0.7 1.2
2.0
3.5
6.0
sampled decision range ℓc
Figure 6: Finite noncanonical graphs. Top: ηS (r) computed from exact Gaussian conditional covariances and the resulting regret under one common normalized setting; lines join integer-radius evaluations. Bottom: exact integer optima at the sampled coherence times and decision ranges; no transition locations between those settings are inferred. Distinct marker shapes identify coincident graph results. The canonical line closed form is not used.
35
7
Related Work
We organize the closest literatures around three distinctions: iterative computation versus information architecture, controlled dynamics versus an exogenous environment, and age-based freshness versus decision-relevant value. These distinctions frame the more specific connections developed below.
7.1
Static Team Decision Theory
Radner established foundational optimality conditions for static teams, in which decision makers share a common objective but act on different information [10]. Under appropriate convexity and regularity conditions, the resulting stationarity conditions are sufficient for global team optimality; in the quadratic Gaussian specialization, the optimal decision rules are affine [10, 11]. In the deterministic-Hessian quadratic model used here, these fixed-information optimality conditions admit the equivalent Hilbert-space formulation developed in Section 2.4: for a fixed information architecture A, the team-optimal policy is the Q-orthogonal projection of the full-information action onto the closed subspace S(A) of implementable policies. This projection identity specializes the classical static-team framework to our quadratic model; Radner’s theorem is not stated in this literal form. Krainak, Speyer, and Marcus subsequently relaxed the sufficient conditions associated with Radner’s stationarity result and extended the theory to an exponential-quadratic criterion with jointly Gaussian state and observations [23]. In a companion paper, they formulated affine linear-quadratic-Gaussian and linear-exponential-Gaussian team problems as constrained parameter optimizations and represented optimal team gains as projections of the corresponding centralized gains [17]. Classical team theory also treated information itself as an organizational design object. Marschak and Radner analyzed the value of information, delayed information and the tradeoff between timeliness and completeness, and formulated team organization in terms of optimal networks [11]. More recently, Summers, Li, and Kamgarpour formulated information-structure design as the selection of additional measurements or communication links jointly with the corresponding team decision strategies [12]. Afshari and Mahajan characterized quadratic Gaussian static teams whose observations decompose into information common to all agents and agent-specific local information [24]. They subsequently applied this decomposition to team-optimal decentralized filtering over communication graphs with delayed information sharing, where sufficiently old measurements constitute common information while more recent measurements remain local [25]. Rusmevichientong and Van Roy studied large teams in which each agent observes the local cost structure only within a graph radius, using centralized performance as a benchmark [26]. Gamarnik, Goldberg, and Weber studied decentralized decisions based on local information in random decision networks and related their near-optimality to a correlation-decay, or long-range-independence, property [27]. Relative to these works, the distinctive object studied here is a structured family of information architectures in which spatial scope and temporal freshness are coupled through the scope–latency relation τ (r). Each spatial scope induces an acquisition delay, and the resulting architecture is valued through its prediction error for the full-information decision. This coupling yields the freshness–locality crossover and coordination-scale laws developed in Sections 3 and 5.
7.2
Distributed Optimization, Communication Constraints, and Scope
Distributed optimization studies how agents compute a common solution using local computation and message exchange. Classical asynchronous gradient methods allow delayed interprocessor
36
communication while retaining convergence guarantees under suitable boundedness conditions [28], while distributed subgradient methods establish convergence when agents exchange information locally over time-varying network topologies [3]. In these canonical formulations, the communication process constrains an iterative computation of an optimizer. Our setting separates that computational question from information architecture: the full-information decision at each epoch is the benchmark, and the object of design is the information available when that decision must be made. Networked and decentralized control place communication restrictions directly in the feedback loop, including sampling, packet loss, communication topology, and delay [29]. For spatially invariant plants, early work showed that optimal controllers can exhibit inherent spatial localization [5], while subsequent work imposed explicit finite-speed communication structure. Arbelaiz, Bamieh, Hosoi, and Jadbabaie characterize the spatial decay of optimal estimator gains, using the decay rate as a proxy for the distance over which measurements are useful in a spatially distributed dynamic estimation problem [30]. Voulgaris, Bianchini, and Bamieh studied H2 control under delayed communication requirements [6]; Bamieh and Voulgaris formulated distance-dependent information propagation through funnel causality [7]; and Fardad and Jovanović developed state-space models for distributed controllers with finite communication speed [8]. These works are important antecedents to our use of spatial scope and propagation time, but they optimize dynamic feedback for controlled plants rather than information-constrained stage decisions in an exogenous environment. Communication structure has also been treated as a design variable. Langbort and Gupta study interconnection topology when communication is assigned a topology-dependent cost [31], and Matni co-designs communication-delay structure and H2 control using regularization [32]. In wireless consensus, Vanka, Gupta, and Haenggi explicitly trade communication radius against interference-induced delay and show that the preferred connectivity depends on network geometry [15]. Most directly, Ballotta, Jovanović, and Schenato parameterize control architectures by communication scope: each agent receives measurements from nodes within a chosen number of communication hops, with an architecture-dependent delay that increases with that scope. After optimizing the feedback gains for each scope, they show that sparse architectures can outperform all-to-all feedback [9]. Ballotta and Gupta obtain an analogous conclusion when maximizing consensus convergence speed under hop-dependent communication delay [14]. These results establish that scope-dependent delay can itself induce a sparse or intermediate optimal control architecture. Our setting differs in the decision problem: an exogenous-state static team, with the current full-information action as the benchmark and an architecture valued by the prediction error it induces for that action. This yields the decision-relevant spatial omission function ηS (r), temporal unpredictability ηT (τ ), their exact common-rate composition law, and the fresh-local versus stale-global and coordination-radius laws derived above. Ballotta, Arbelaiz, Gupta, Schenato, and Jovanović study a different but complementary question for spatially invariant dynamic plants: how delayed measurements from different spatial locations alter the spatial locality of the optimal feedback kernel [16]. Their results operate through closed-loop plant dynamics and controller localization, whereas ours operate through prediction of an exogenously varying full-information decision.
7.3
Freshness and Age of Information
Age of Information (AoI) formalizes the timeliness of status information through the age of the freshest received update, ∆(t) = t − u(t). Seminal work showed that minimizing AoI is distinct from maximizing throughput or minimizing packet delay and initiated a large literature on update generation, queueing, and scheduling [2,33]. Subsequent work generalized linear age to applicationspecific nonlinear age penalties [34]. Connections to remote estimation make the distinction between 37
age and task loss explicit: for a Gauss–Markov Ornstein–Uhlenbeck process, estimation error under signal-agnostic sampling is a nonlinear function of age, whereas signal-aware sampling can exploit the realized estimation error and outperform age-optimal policies [18]. More recent work has pushed freshness metrics toward task-aware objectives. Shisher et al. showed that remote-inference loss can be nonmonotonic in age and designed communication policies around inference performance [35]; subsequent work considers correlated sources for which inference penalties depend jointly on multiple sources’ ages [36]. Li et al. go further toward decision-aware freshness by jointly optimizing sampling and remote decision making under stale observations [37]. Those works optimize when or which updates reach a fixed remote receiver; we instead choose the information architecture available to a decentralized team: what is observed, from what spatial scope, and at what age. Because broader scope itself incurs greater acquisition delay, spatial omission and temporal staleness are coupled by the architecture. This leads to the fresh-local versus stale-global crossover and optimal coordination-scale results developed here.
7.4
Centralization, Delegation, and Organizational Economics
A closely related organizational-economics literature asks how dispersed information, communication constraints, and decision rights determine organizational form. Aoki compares hierarchical and horizontal information structures when local conditions are uncertain and central monitoring or response is limited [38]. In the information-processing tradition, Radner studies organizations of limited-capacity processors trading off processing resources against decision delay [39]; Bolton and Dewatripont derive communication networks from processing and communication costs [40]; and Van Zandt studies real-time decentralized processing in which computational delay constrains the use of recent information [41]. Dessein and Santos likewise study adaptation to local information versus coordination under imperfect communication [42]. A complementary delegation literature makes decision rights endogenous when information and objectives are dispersed. Aghion and Tirole distinguish formal from real authority [43], while Dessein studies delegation as an alternative to strategic communication [44]. Most closely related to our centralization framing, Alonso, Dessein, and Matouschek compare centralized and decentralized coordination when managers privately observe local conditions [4]. Our model abstracts from incentive conflicts and endogenous information acquisition: agents share a common objective and the environment evolves exogenously. The tradeoff studied here is spatio-temporal: broader synchronized snapshots can improve coordination but arrive older. This yields a pure-benchmark crossover and, in the canonical spatial model, an interior optimum within the synchronized-snapshot family.
7.5
Spatial Statistics and Graph Signal Processing
Spatial statistics provides the closest statistical antecedent to our spatial-predictability model. Kriging casts spatial prediction as an optimal linear prediction problem under spatial dependence [45]. For lattice data, Besag’s conditional formulation represents spatial dependence through local Markov neighborhoods [46]; Lindgren, Rue, and Lindström later connect Matérn Gaussian fields to sparse Gaussian Markov random fields through stochastic partial differential equation representations [21]. Space–time statistics treats separability as an optional modeling restriction: Gneiting constructs broad classes of stationary nonseparable covariance functions [22]. These works provide the statistical context for our covariance-based spatial model and its separable and nonseparable distinction. Graph signal processing extends analogous ideas to irregular domains. Shuman et al. use graphLaplacian eigenvectors to define a graph Fourier domain and associate low graph frequencies with
38
signals that vary slowly over the graph [20]. Graph-sampling theory asks when bandlimited graph signals can be exactly recovered from partial observations [47], while stationary graph signal processing gives spectral descriptions and Wiener-type estimators for random graph signals [48]. Our objective is different: we do not value accurate reconstruction of θt per se. A radius-r architecture is evaluated by its loss in reproducing the full-information decision x⋆t , and increasing r also increases information age through τ (r). Thus spatial omission and temporal staleness are priced jointly, and graph-smooth modes matter only insofar as they are decision-relevant.
8
Discussion and Conclusions
What the theory establishes. The central constraint in networked decision making is that information breadth and freshness generally cannot be chosen independently. Nearby observations can be used quickly but may omit remote state needed for coordination; a broader view can improve coordination but loses predictive value while it is collected and disseminated. Architecture regret measures this loss relative to the full-information decision after optimizing over every policy implementable under the architecture. Theorem 3 characterizes the fresh-local versus stale-global endpoints. Theorem 5 identifies the optimal synchronized-snapshot radius through marginal balance, and Theorem 6 gives its closedform scaling law in the canonical model. Normalization by open-loop regret separates the overall scale of the decision problem from the fraction of decision-relevant variation preserved by an architecture, enabling comparisons across modeled systems. Why decision-relevant predictability matters. Raw state predictability is not the correct design criterion. Architecture regret depends on predicting the full-information decision x⋆ , not on reconstructing every component of θ. A remote state component can be hard to predict but architecturally unimportant when it has little influence on the desired action. Conversely, a small residual uncertainty can matter when the decision-sensitivity operator K strongly amplifies that mode. The relevant predictability therefore combines environmental statistics, decision sensitivity and coupling, information geometry, and freshness (Section 6.5). This distinction also separates physical communication geometry from decision geometry. Graph distance determines which observations can be acquired and their age on arrival; Q, B, and K determine whether those observations change the desired joint action. Two nearby nodes may be weakly coupled, while distant nodes may interact strongly through a shared constraint or common decision mode. An architecture based only on physical distance can therefore communicate extensively about irrelevant variation while omitting a remote mode with high decision value. Scope and limitations. The paper establishes an exogenous-state, static-team baseline. Environmental state may be correlated across space and time, but current actions neither alter its future law nor change what another agent later observes. Within this boundary, the fixed-architecture quadratic projection gives an exact benchmark for each fixed information pattern, and the paper’s new results compare a structured scope–age family using that benchmark. The principal radius family uses synchronized snapshots. Under the common-rate Gaussian model, Lemma 2 makes current snapshots sufficient relative to observed histories for the affine current-decision prediction problem; the equivalence is specific to that model. The fixed-epoch projection does not require stationarity, although stationarity makes expected performance time-invariant and ergodicity is needed for sample-path long-run averages.
39
The projection identity assumes quadratic costs, a deterministic positive- definite Hessian, and unconstrained Euclidean actions. The explicit stale-global and composition laws additionally require affine full-information decisions and a common-rate Gaussian or Gauss–Markov process, while the canonical radius formula uses a separable homogeneous field and linear propagation delay. Nonquadratic or discrete decisions, state-dependent Hessians, and feasibility constraints such as nonnegativity, capacity, or simplex restrictions require different analytical tools. Outside the stated assumptions, one should work directly with decision-relevant prediction-error functions or suitable bounds rather than assume the same factorization. The experiments do not probe temporal-rate heterogeneity, nonseparable covariance, nonGaussian evolution, nonlinear decisions, or mixed-age architectures; each remains an extension. Finally, the common-rate assumption suppresses possible alignment between temporal coherence, spatial scale, and decision relevance. If long-range decision-relevant modes evolve more slowly than local modes, a scalar common-rate approximation may understate the value of broad delayed information; the reverse alignment may overstate it. There is no universal direction of bias without specifying the mode weights and the procedure used to select an effective coherence time. With heterogeneous temporal modes, the scalar composition law generally gives way to a joint spatiotemporal decision-prediction error. Most important extensions. A first extension is to price information acquisition explicitly. The present model represents communication and sensing constraints mainly through the feasible scope–delay relationship. Adding an acquisition cost or budget could produce sparse, nonuniform topologies and compressed summaries that preserve decision-relevant modes with less latency. A second extension is adaptation to unknown or nonstationary statistics. Temporal coherence, spatial predictability, and decision sensitivity may need to be estimated online while the architecture changes on a slower timescale than operational decisions. The challenge is to account for estimation uncertainty while retaining a stable architecture-selection rule. The most important extension is action-dependent state evolution and dynamic teams, where actions alter future state and may signal information to other agents. Extending informationarchitecture design beyond the exogenous-state baseline is therefore not a routine modification. Witsenhausen’s counterexample [13] explains why: once control and signaling interact, the fixedsubspace projection characterization generally no longer solves the team problem. Design implication. Taken together, the theory yields a practical design heuristic: expand information scope as long as the marginal decision value of broader information exceeds the predictive value lost in obtaining it. AI Use Statement. The development of this paper made substantial use of AI tools, including Claude Code, ChatGPT/Codex, Gemini, and Grok, for drafting and revising text, developing code and figures, mathematical formulation and proofs. The human authors formulated the original problem, made substantial contributions throughout the paper’s development, reviewed the manuscript, and accept full responsibility for its contents. Code and Data Availability. The manuscript source, numerical code, experiment configurations, processed reference results, and figure-generation pipeline are available in the public TINA repository at https://github.com/ANRGUSC/TINA [49].
40
References [1] M. Mitzenmacher, “How useful is old information?” IEEE Transactions on Parallel and Distributed Systems, vol. 11, no. 1, pp. 6–20, Jan. 2000. [2] S. Kaul, R. Yates, and M. Gruteser, “Real-time status: How often should one update?” in 2012 Proceedings IEEE INFOCOM. IEEE, 2012, pp. 2731–2735. [3] A. Nedic and A. Ozdaglar, “Distributed subgradient methods for multi-agent optimization,” IEEE Transactions on Automatic Control, vol. 54, no. 1, pp. 48–61, 2009. [4] R. Alonso, W. Dessein, and N. Matouschek, “When does coordination require centralization?” American Economic Review, vol. 98, no. 1, pp. 145–179, 2008. [5] B. Bamieh, F. Paganini, and M. A. Dahleh, “Distributed control of spatially invariant systems,” IEEE Transactions on Automatic Control, vol. 47, no. 7, pp. 1091–1107, Jul. 2002. [6] P. G. Voulgaris, G. Bianchini, and B. Bamieh, “Optimal H2 controllers for spatially invariant systems with delayed communication requirements,” Systems & Control Letters, vol. 50, no. 5, pp. 347–361, Dec. 2003. [7] B. Bamieh and P. G. Voulgaris, “A convex characterization of distributed control problems in spatially invariant systems with communication constraints,” Systems & Control Letters, vol. 54, no. 6, pp. 575–583, Jun. 2005. [8] M. Fardad and M. R. Jovanović, “Design of optimal controllers for spatially invariant systems with finite communication speed,” Automatica, vol. 47, no. 5, pp. 880–889, May 2011. [9] L. Ballotta, M. R. Jovanović, and L. Schenato, “Can decentralized control outperform centralized? the role of communication latency,” IEEE Transactions on Control of Network Systems, vol. 10, no. 3, pp. 1629–1640, Sep. 2023. [10] R. Radner, “Team decision problems,” The Annals of Mathematical Statistics, vol. 33, no. 3, pp. 857–881, 1962. [11] J. Marschak and R. Radner, Economic Theory of Teams, ser. Cowles Foundation Monographs. New Haven, CT: Yale University Press, 1972, no. 22. [12] T. Summers, C. Li, and M. Kamgarpour, “Information structure design in team decision problems,” IFAC-PapersOnLine, vol. 50, no. 1, pp. 2530–2535, 2017. [13] H. S. Witsenhausen, “A counterexample in stochastic optimum control,” SIAM Journal on Control, vol. 6, no. 1, pp. 131–147, 1968. [14] L. Ballotta and V. Gupta, “Faster consensus via a sparser controller,” IEEE Control Systems Letters, vol. 7, pp. 1459–1464, 2023. [15] S. Vanka, V. Gupta, and M. Haenggi, “Power-delay analysis of consensus algorithms on wireless networks with interference,” International Journal of Systems, Control and Communications, vol. 2, no. 1/2/3, pp. 256–274, 2010. [16] L. Ballotta, J. Arbelaiz, V. Gupta, L. Schenato, and M. R. Jovanović, “The role of communication delays in the optimal control of spatially invariant systems,” IEEE Transactions on Automatic Control, vol. 71, no. 2, pp. 978–993, Feb. 2026. 41
[17] J. C. Krainak, J. Speyer, and S. I. Marcus, “Static team problems–part II: Affine control laws, projections, algorithms, and the LEGT problem,” IEEE Transactions on Automatic Control, vol. 27, no. 4, pp. 848–859, 1982. [18] T. Z. Ornee and Y. Sun, “Sampling and remote estimation for the Ornstein–Uhlenbeck process through queues: Age of information and beyond,” IEEE/ACM Transactions on Networking, vol. 29, no. 5, pp. 1962–1975, 2021. [19] C. E. Rasmussen and C. K. Williams, Gaussian Processes for Machine Learning. MIT press Cambridge, MA, 2006. [20] D. I. Shuman, S. K. Narang, P. Frossard, A. Ortega, and P. Vandergheynst, “The emerging field of signal processing on graphs: Extending high-dimensional data analysis to networks and other irregular domains,” IEEE Signal Processing Magazine, vol. 30, no. 3, pp. 83–98, 2013. [21] F. Lindgren, H. Rue, and J. Lindström, “An explicit link between Gaussian fields and Gaussian Markov random fields: the stochastic partial differential equation approach,” Journal of the Royal Statistical Society: Series B (Statistical Methodology), vol. 73, no. 4, pp. 423–498, 2011. [22] T. Gneiting, “Nonseparable, stationary covariance functions for space–time data,” Journal of the American Statistical Association, vol. 97, no. 458, pp. 590–600, 2002. [23] J. C. Krainak, J. Speyer, and S. I. Marcus, “Static team problems–part i: Sufficient conditions and the exponential cost criterion,” IEEE Transactions on Automatic Control, vol. 27, no. 4, pp. 839–848, 1982. [24] M. Afshari and A. Mahajan, “Static teams with common information,” IFAC-PapersOnLine, vol. 50, no. 1, pp. 11 926–11 931, 2017. [25] ——, “Team optimal decentralized state estimation,” in 2018 IEEE Conference on Decision and Control (CDC). IEEE, 2018, pp. 5044–5050. [26] P. Rusmevichientong and B. Van Roy, “Decentralized decision-making in a large team with local information,” Games and Economic Behavior, vol. 43, no. 2, pp. 266–295, 2003. [27] D. Gamarnik, D. A. Goldberg, and T. Weber, “Correlation decay in random decision networks,” Mathematics of Operations Research, vol. 39, no. 2, pp. 229–261, 2014. [28] J. N. Tsitsiklis, D. P. Bertsekas, and M. Athans, “Distributed asynchronous deterministic and stochastic gradient optimization algorithms,” IEEE Transactions on Automatic Control, vol. 31, no. 9, pp. 803–812, Sep. 1986. [29] J. P. Hespanha, P. Naghshtabrizi, and Y. Xu, “A survey of recent results in networked control systems,” Proceedings of the IEEE, vol. 95, no. 1, pp. 138–162, Jan. 2007. [30] J. Arbelaiz, B. Bamieh, A. E. Hosoi, and A. Jadbabaie, “Optimal estimation in spatially distributed systems: How far to share measurements from?” IEEE Transactions on Automatic Control, vol. 70, no. 5, pp. 3226–3239, 2025. [Online]. Available: https://doi.org/10.1109/TAC.2024.3504257 [31] C. Langbort and V. Gupta, “Minimal interconnection topology in distributed control design,” SIAM Journal on Control and Optimization, vol. 48, no. 1, pp. 397–413, 2009.
42
[32] N. Matni, “Communication delay co-design in H2 -distributed control using atomic norm minimization,” IEEE Transactions on Control of Network Systems, vol. 4, no. 2, pp. 267–278, Jun. 2017. [33] R. D. Yates, Y. Sun, D. R. Brown, S. K. Kaul, E. Modiano, and S. Ulukus, “Age of information: An introduction and survey,” IEEE Journal on Selected Areas in Communications, vol. 39, no. 5, pp. 1183–1210, 2021. [34] Y. Sun, E. Uysal-Biyikoglu, R. D. Yates, C. E. Koksal, and N. B. Shroff, “Update or wait: How to keep your data fresh,” IEEE Transactions on Information Theory, vol. 63, no. 11, pp. 7492–7508, 2017. [35] M. K. C. Shisher, Y. Sun, and I.-H. Hou, “Timely communications for remote inference,” IEEE/ACM Transactions on Networking, vol. 32, no. 5, pp. 3824–3839, 2024. [36] M. K. C. Shisher, V. Tripathi, M. Chiang, and C. G. Brinton, “AoI-based scheduling of correlated sources for timely inference,” IEEE Transactions on Networking, vol. 34, pp. 2181– 2195, 2026. [37] A. Li, S. Wu, G. C. F. Lee, and S. Sun, “From freshness to effectiveness: Goal-oriented sampling for remote decision making,” IEEE Transactions on Information Theory, vol. 72, no. 4, pp. 2237–2276, Apr. 2026. [Online]. Available: https://doi.org/10.1109/TIT.2026.3663678 [38] M. Aoki, “Horizontal vs. vertical information structure of the firm,” The American Economic Review, vol. 76, no. 5, pp. 971–983, 1986. [39] R. Radner, “The organization of decentralized information processing,” Econometrica: Journal of the Econometric Society, vol. 61, no. 5, pp. 1109–1146, 1993. [40] P. Bolton and M. Dewatripont, “The firm as a communication network,” The Quarterly Journal of Economics, vol. 109, no. 4, pp. 809–839, 1994. [41] T. Van Zandt, “Real-time decentralized information processing as a model of organizations with boundedly rational agents,” The Review of Economic Studies, vol. 66, no. 3, pp. 633–658, 1999. [42] W. Dessein and T. Santos, “Adaptive organizations,” Journal of Political Economy, vol. 114, no. 5, pp. 956–995, 2006. [43] P. Aghion and J. Tirole, “Formal and real authority in organizations,” Journal of Political Economy, vol. 105, no. 1, pp. 1–29, 1997. [44] W. Dessein, “Authority and communication in organizations,” The Review of Economic Studies, vol. 69, no. 4, pp. 811–838, 2002. [45] N. Cressie, “The origins of kriging,” Mathematical Geology, vol. 22, no. 3, pp. 239–252, 1990. [46] J. Besag, “Spatial interaction and the statistical analysis of lattice systems,” Journal of the Royal Statistical Society: Series B (Methodological), vol. 36, no. 2, pp. 192–225, 1974. [47] S. Chen, R. Varma, A. Sandryhaila, and J. Kovačević, “Discrete signal processing on graphs: Sampling theory,” IEEE Transactions on Signal Processing, vol. 63, no. 24, pp. 6510–6523, 2015. 43
[48] N. Perraudin and P. Vandergheynst, “Stationary signal processing on graphs,” IEEE Transactions on Signal Processing, vol. 65, no. 13, pp. 3462–3477, Jul. 2017. [49] S. Moeller and B. Krishnamachari, “TINA: Reproduction package for “a theory of information architecture for networked decisions”,” GitHub repository, 2026, source code, data, figures, and manuscript materials. [Online]. Available: https://github.com/ANRGUSC/TINA
44