Anonymous sharing is pairwise phase-blind Neutrality at order two and third-order structure in a fleet of checkpointing jobs Brieuc Le Roux Tardif IMT Nord Europe [email protected]
arXiv:2607.28377v1 [eess.SY] 30 Jul 2026
July 31, 2026 Abstract Independent training jobs sharing a storage system write their checkpoints through the same finite bandwidth, and the resulting bursts of correlated I/O, together with the facility-level power transients that accompany them, are commonly described as a self-reinforcing “checkpoint storm”. We formalise the self-reinforcement as phase locking in a population of integrate-and-fire oscillators coupled through a shared resource, and show that within that model it fails for a reason that has nothing to do with checkpointing. Call a resource anonymous if the rate it delivers to an active user depends on how many users are active and not on which. For identical jobs whose write is shorter than their compute interval, an anonymous resource produces no pairwise coupling at all: the two-job return map of the phase gap is the identity, under storage contention, under a shared power cap and under both at once, so the two-body interaction on which the Kuramoto and Mirollo–Strogatz frameworks are built is not weak here but absent. Anonymity also freezes the firing order, for any fleet size and any cap, so no trajectory reaches the synchronous state from outside it. What survives is a third-order effect: where all N write windows overlap and the cap does not bind, the map is diagonal in the intervals between consecutive write starts, aj 7→ ((N − j)/j) aj , with reciprocal spectrum and unit determinant, which makes synchrony a fixed point with ⌈N/2⌉ − 1 expanding directions rather than an attractor. That determinant follows from anonymity and not from fairness: for any anonymous throughput f with f (n) ≤ n the spectrum becomes (N − j)f (j)/(j f (N − j)), whose product is still one. Numerically, a fleet launched at random neither locks nor clusters, and the only memory of its launch our observables resolve is that frozen order: after hundreds of cycles its configuration sits as far from its own launch as from an independent one relabelled into the same order sector, and no decay of its smallest gap is resolved. Absence of locking is not absence of bursts, the upper tail of the number of concurrent writers staying above its independent-phase value in every cell measured. Nothing erodes a stagger in the deterministic uncapped model; per-cycle jitter σ does, a margin m surviving 0.11 to 0.24 (m/σ)2 cycles across the fleets tested. Heterogeneous jobs behind a binding cap do acquire a genuine pairwise coupling, which is where the statement stops generalising.
1
Introduction
Large training runs checkpoint periodically, and a checkpoint is a burst: tens to hundreds of gigabytes leave the accelerators for a parallel file system while the GPUs sit idle. Two distinct things are called a checkpoint storm, and only one of them is our subject. Within a single distributed job, thousands of ranks serialise their shards at the same instant by construction, and the resulting burst is a design problem in the checkpointing stack itself [7]. Across independent jobs sharing one fabric, the bursts contend without any mechanism aligning them, and the operational question is whether they nonetheless drift into alignment. This paper is about the second question only; nothing here bears on the first, which is a sufficient cause of correlated I/O on its own. The electrical stakes are set by the first: instantaneous fleet-level swings of tens of megawatts are reported for a cluster of order 104 accelerators
1
[3, 9], and the sub-second timescale on which they matter for facility stability is the subject of a growing power-engineering literature [8]. The implicit model behind mitigation practice is that collisions between jobs are self-reinforcing: jobs that once collide keep colliding, so schedules must be deliberately staggered. Recent systems work treats the timing of a job’s synchronisation points as a schedulable quantity, choosing when to run DiLoCo outer merges against measured fleet pressure [10]; complementary work smooths the power transients of a single job [4]. Neither studies the free dynamics of the collision itself. The question this raises is dynamical: is the synchronised state an attractor? If it is, staggering is futile, because contention drags jobs back into phase. If it is repelling, contention is not the cause and a stagger has nothing to undo, which is not the same as a fleet that can be left alone: three properties have to be kept apart, asymptotic phase locking, transient clustering, and the extreme upper tail of the number of concurrent writers. We settle the first within the model, measure the second, and find the third above its independent-phase value even where the first two are absent (§6); storm below always names the first, which is the only one of the three the dynamical question is about. A companion paper [5] asks the same question for a different channel, a shared power envelope with a delayed throttle, and finds a coupling whose sign is set by the control lag: repulsive to leading order, attractive only once the lag exceeds half a cycle. The present paper removes the lag and the memory, keeps the sharing, and finds that the pairwise term does not merely change sign, it vanishes. The two are reconciled in §9: the coupling in [5] is a property of the controller, not of the sharing. The natural formalism is that of pulse-coupled oscillators: each job is an integrate-and-fire unit whose phase is its progress through the inter-checkpoint interval and whose discharge is the checkpoint write. For excitatory coupling with a concave rise, Mirollo and Strogatz proved that almost every initial condition synchronises [6], and inhibitory all-to-all coupling has since been shown to produce globally attracting synchrony as well [1]; applications have concentrated on biological populations and on clock synchronisation in wireless sensor networks [11]. The datacenter version differs in the coupling channel and, as we show, in the order of the interaction. Contributions. Graded by the strength of the supporting argument. (i) (Theorem, proved) For two identical jobs with d < T , the return map of the phase gap is the identity: under storage contention, under a shared power cap, and under both at once, for any cap severity. The pairwise coupling function is identically zero, not merely weak (§3). (ii) (Theorem, proved) For N identical jobs whose write windows all overlap, with the cap not binding, the return map is diagonal in interval coordinates with eigenvalues (N − j)/j: reciprocal spectrum, unit determinant on that branch, and synchrony a fixed point with ⌈N/2⌉−1 expanding directions rather than an attractor. The determinant is a consequence of anonymity and not of fairness, since an arbitrary anonymous throughput f multiplies each eigenvalue by f (j)/f (N − j) and leaves the product at one (§4). (iii) (Propositions, proved) Two invariant structures (§5). The cyclic firing order is a constant of the motion, for any fleet size, any cap and any anonymous rule, so distinct phases stay distinct and synchrony is reached by no trajectory that does not start there. And collision-free configurations, which exist exactly when L ≤ N/(N − 1), are invariant and cannot be entered from outside, the flow being uniquely reversible where the write windows are disjoint. (iv) (Numerical) Off that set, nearby trajectories separate exponentially in most launches, yet a single trajectory does not cluster: no decay of the smallest gap is resolved over 800 cycles, the gap of the pair tightest at launch opens by a factor 6 to 33, and Daido moments up to m = 4 show no coherence that a change of null does not remove. Nor does it keep of its launch anything these observables resolve beyond the invariant order: its distance to the launch, 7–14% of a cycle, 2
is the distance to an independent launch relabelled into the same order sector, cell by cell and within a standard error, and the two distributions agree to a 1-Wasserstein distance below 5% of their own spread, while the 22–26% separating two unrelabelled configurations measures a permutation the dynamics cannot perform. The upper tail of the number of concurrent writers is nonetheless heavier than that of independent phases in every cell (§6). (v) (Numerical, operational) With per-cycle jitter σ on the compute time, each gap of a staggered fleet is a driftless random walk, and a margin m survives 0.11 to 0.24 (m/σ)2 cycles depending on the fleet, verified over two decades (§7). (vi) (Numerical, negative) Heterogeneous jobs behind a binding cap do couple pairwise, with a phase response supported on gaps below the detuning, although both resources remain anonymous. This bounds the reach of (i) (§8). Scope. The results are statements about a model, and the model is deliberately minimal (§2). We prove what follows from three assumptions, we verify each proof against an exact event-driven integrator, and we state explicitly which assumptions, when relaxed, restore a pairwise coupling, in one case by exhibiting it (§8). We do not claim that real fleets never synchronise; we claim that in this model symmetric contention between identical jobs cannot be the cause.
2
Model
Definition 1 (Fleet). A fleet is a set of N jobs. Job i alternates two phases. In the compute phase it must accumulate an amount Ti > 0 of work at the instantaneous rate v(t). In the write phase it must drain a checkpoint volume Vi > 0 at the instantaneous rate ri (t). Time is normalised so that a job computing alone advances at unit rate, and bandwidth is normalised so that a job writing alone drains at unit rate; hence di := Vi is the solo write duration of job i. Job i fires at the instant si at which it enters its write phase, and the observable is the vector of firing times modulo the cycle. Write nw (t) for the number of jobs writing at time t and nc (t) for the number computing. Assumption 1 (Blocking checkpoint). A writing job does not compute: the phases are exclusive. Justification: synchronous checkpointing, in which the training loop halts while state is serialised and flushed, is the baseline that asynchronous mechanisms are designed to improve on [7], and it remains in use wherever that path is unavailable or not enabled. We do not claim it is the majority configuration in production, having no measurement to support such a claim. This is the assumption that matters most (§8). Assumption 2 (Anonymous equal sharing). Each writer receives bandwidth f (nw (t))/nw (t), with f positive, f (1) = 1 and f (n) ≤ n; each computing job advances at v(t) = min{1, C/nc (t)}, where C ∈ (0, ∞] is a power cap expressed as the number of jobs the cap sustains at full speed. Both rates depend on the number of active users only, and are identical across active users. The work-conserving case is f ≡ 1. Justification: processor sharing is the standard idealisation of a parallel file system serving concurrent streams without priorities, and of a power budget divided uniformly across a homogeneous fleet. The condition f (n) ≤ n says only that concurrency makes no single stream faster than it would be alone, and is used solely to order events in Theorem 8. Assumption 3 (Deterministic per-cycle work). Ti and Vi are constants of job i. Heterogeneity across jobs is permitted; stochasticity within a job is not, except where §7 introduces it explicitly. Justification: the compute time between checkpoints is set by a fixed number of steps of a fixed model, and is repeatable to a few percent absent failures.
3
Definition 2 (Anonymous resource, phase-blind coupling). A shared resource is anonymous if the rate it delivers to an active user depends only on how many users are active, and not on which. Let k units be taken in isolation and let the section be the set of states at which every unit is computing and the numbers of writes they have completed differ by a prescribed offset, zero being the case of units launched together. There every unit has the same mode and no residual volume, so the hybrid state is the vector of residual works, which the flow translates uniformly for as long as all units compute; the state modulo that translation is the vector of k − 1 firing differences, and the return map F of those differences is well defined on it. That the section is nonempty and recurrent with a constant offset is a property of the system and not of the definition: it is proved for the pair of Theorem 4 under d < T and assumed in Conjecture 17. Let g(δ) := F (δ) − δ (1) be the increment that map applies. The coupling is phase-blind at order k if g is constant, that is if the increment does not depend on the phase differences themselves. Identical units then have g ≡ 0, since exchange symmetry fixes the synchronous configuration and g is constant; units that differ are transported by a translation set by their detuning alone on a fixed event-order branch with the cap not binding (Proposition 7), and not in general (§8.1). Phase-blindness is the exact vanishing of the coupling, not of the frequency spread. The hierarchy of orders is the hierarchy of subsystems: order k is a statement about k units in isolation, not a term in a series expansion. Blindness stopping at order three therefore means that no dependence on relative phase appears before three jobs are present, and Theorem 8 exhibits it at exactly three. Dimensionless parameters. Set Ti ≡ 1 for the homogeneous case. The two parameters that survive Pare the duty d, the solo write duration in units of the compute time, and the aggregate demand L := i di = N d, the write volume the fleet requests per unit of compute time. Neither is a duty cycle in the usual sense: a job alone occupies the fabric for a fraction d/(1 + d) of its free period P0 = 1 + d. A cyclic schedule giving every write an interval of at least d to itself exists exactly when N d ≤ 1 + d,
equivalently
L ≤ L⋆ :=
N , N −1
(2)
which tends to 1 from above as the fleet grows. We call L⋆ the stagger-feasibility threshold rather than a saturation threshold: it is a geometric condition on the gaps, and it is separately observed (§6.4) that the effective cycle begins to stretch beyond P0 once it is crossed. The same threshold governs Proposition 16.
3
Order two: anonymous resources are phase-blind
The argument rests on a single accounting identity, and the quantity it conserves is not the phase gap itself but the difference in accumulated work. Write v1 := min{1, C} for the compute rate of a job that computes while the other writes, and v2 := min{1, C/2} for the rate when both compute; v1 = v2 = 1 exactly when C ≥ 2. Let Wi , Ki ⊂ [0, t] be the sets of times at which job i writes and computes. Assumption 1 makes them a partition of [0, t], which is the hypothesis that carries everything below: |Ki | = t − |Wi |.
(3)
With two jobs, “job i writes and job j does not” means job i is the only writer, so it drains at f (1) = 1; likewise a job computing while the other writes is the only computer and advances at v1 . An interval on which one job is alone is one on which its rate is known, and anonymity makes that rate the same for whichever job is alone.
4
Lemma 3 (Equal occupancy, then equal solo work). Let N = 2 with identical jobs (Ti ≡ T , di ≡ d) under Assumptions 1–2, and let t be an instant at which both jobs are computing, having completed k1 and k2 writes respectively and started no further one. Then, for any cap C > 0, |W1 | − |W2 | = (k1 − k2 ) d,
|K1 \ K2 | − |K2 \ K1 | = (k2 − k1 ) d.
(4)
Consequently the two jobs have accumulated the same amount of compute work outside the intervals on which both compute, up to a term set by the write offset alone: writing wi (t) for the compute work job i has accumulated since time 0, the difference ∆w(t) := w1 (t) − w2 (t) depends on t only through k1 − k2 . Two counters have to be kept apart. Decompose wi (t) = ki T + xi (t), with xi (t) ∈ [0, T ) the work banked in the current cycle, which is what the next firing time is set by; then along a sequence of section instants of constant offset k1 − k2 the two differences ∆w and ∆x := x1 − x2 are constant together, and |∆x| < T whatever the offset. Proof. Writing. By hypothesis job i has drained exactly ki d by time t, with no partial write outstanding. Split that volume over W1 ∩ W2 and Wi \ Wj . On the common part both writers receive the same rate, whatever it is, so the volumes drained there are equal; call the common value S. On Wi \ Wj job i is the only writer and drains at f (1) = 1, so it drains |Wi \ Wj | there. Hence ki d = S + |Wi \ Wj | for i = 1, 2, so |W1 \ W2 | − |W2 \ W1 | = (k1 − k2 )d and therefore |W1 | − |W2 | = (k1 − k2 )d. Computing. The partition (3) turns that into |K1 | − |K2 | = (k2 − k1 )d, and subtracting the common part |K1 ∩ K2 | gives the second display of (4). On Ki \ Kj job i computes while job j writes, so it is the only computer and advances at v1 ; on K1 ∩ K2 both advance at the same rate and accumulate equally. Every contribution to w1 − w2 after time 0 therefore cancels except the solo term, leaving ∆w(t) = ∆w(0) + v1 (k2 − k1 )d, a function of the offset alone. Theorem 4 (Pair neutrality). Let N = 2 with identical jobs under Assumptions 1–3, with d < T , with any combination of the two shared resources and any cap C > 0. Let δk be the firing gap on cycle k. Then δk+1 = δk for every k and every δ0 : the increment (1) vanishes identically, and no phase gap is either attracted or repelled. Proof. The section instants of Lemma 3 recur, and establishing that needs no control over the effective period, which a binding cap makes several times P0 . Suppose, for contradiction, that the two jobs never compute simultaneously on [0, t]. Assumption 1 makes the phases exclusive, so job 1 computing forces job 2 to be writing: K1 ⊆ W2 , and on K1 job 2 is then the only writer, hence drains at f (1) = 1. The volume it drains on K1 is therefore |K1 |, which cannot exceed the volume it has drained by time t, itself at most (k2 + 1)d with k2 its number of completed writes; and job 1 has accumulated k1 T of work at a rate at most 1, so |K1 | ≥ k1 T . Hence k1 T ≤ (k2 + 1) d, and symmetrically k2 T ≤ (k1 + 1) d. Every rate is bounded below, by min{1, C/2} in the compute phase and by f (2)/2 in the write phase,so k1 and k2 grow without bound with t; summing the two inequalities gives T ≤ d 1 + 2/(k1 + k2 ) and hence T ≤ d in the limit, against the hypothesis d < T . The two jobs therefore compute simultaneously at some instant, and the argument applied from that instant onward makes such instants recur. That consecutive ones are separated by exactly one write of each job, which is what keeps the offset k1 − k2 constant along the sequence and lets Lemma 3 return one and the same ∆x at all of them, is established at the end of the proof by following one cycle through. Let t be such an instant and suppose ∆x := ∆x(t) > 0, so job 1 leads. From t both jobs compute, at the common rate v2 , until job 1 completes its remaining work T − x1 and fires; at that moment job 2 still needs exactly ∆x. It accumulates that residue while job 1 writes, alone and therefore at v1 , for at most the solo write duration d; if it has not fired by then, job 1 resumes computing and the rate returns to v2 . The firing gap is therefore ( ∆x/v1 , ∆x ≤ v1 d, δk+1 = θ(∆x) := (5) d + (∆x − v1 d)/v2 , ∆x > v1 d, 5
a strictly increasing function of ∆x alone. Lemma 3 makes ∆x the same at every section instant, so δk+1 is the same on every cycle, and in particular equals δk . It remains to follow each case to the end of the cycle, which is what shows that neither job gains a write on the other. If ∆x ≤ v1 d, job 2 fires while job 1 is still writing, job 1 having drained ∆x/v1 ≤ d by then. From that instant both write and receive the same rate, so job 1 keeps its lead ∆x/v1 in drained volume and finishes first, leaving job 2 exactly that residue to drain alone at rate 1; job 1 computes meanwhile, and is the only computer, so it accumulates v1 · ∆x/v1 = ∆x by the time job 2 finishes. If ∆x > v1 d, job 1 drains alone and finishes in d, both compute at v2 from then on, and job 2 fires first since it needs (∆x − v1 d)/v2 < T /v2 against the T /v2 job 1 needs from an empty accumulator; job 2 then writes alone for exactly d while job 1 computes alone at v1 , so job 1 holds (∆x − v1 d) + v1 d = ∆x when that write ends. In both cases job 1 has banked ∆x in the new cycle, which is smaller than T by Lemma 3, so it has not fired again; job 2 has banked nothing, and the instant at which job 2’s write ends is a section instant at which each job has completed exactly one further write, so the offset is unchanged and the current-cycle difference is again ∆x. Remark 5 (On the hypothesis d < T , and on the conserved quantity). The restriction is used only to make the section nonempty, and d < T is the regime of interest, a checkpoint being shorter than the compute interval it protects. Whether the conclusion survives d ≥ T we do not know: the integrator returns the same exact neutrality at d = 1.5 and d = 2.5 with T = 1 at every cap tested, which we report as evidence and not as a theorem. The invariant, in either case, is the current-cycle work difference ∆x and not the firing gap, of which it is the preimage under (5). The distinction is invisible when C ≥ 2, where θ is the identity, and it is why a proof treating the isolated firing time as si + T + d fails for C < 1, where no job ever runs at unit rate and that reference schedule is not the free one; the same caution rules out bounding an event separation by a fraction of P0 , the measured period reaching 3.8 P0 at C = 0.5. Remark 6 (What the proof uses). No step used the value of any rate, only that two jobs in the same phase receive the same one and that a job alone in a phase receives a rate independent of its identity. Neutrality therefore holds for a cap so severe that a single job cannot run at nominal speed (C < 1), a regime in which the instantaneous rates of the two jobs do not coincide: equality of occupancy restores over a cycle what the rates break at each instant (Table 1). The proof also survives letting writers draw on the power budget, nc → nc + αnw with α ≥ 0, which changes v1 in (5) and nothing else. What it does not survive is unequal allocation, nor, once the cap binds, unequal jobs (§8). Proposition 7 (Heterogeneous jobs, cap not binding: local translation on a fixed event-order branch). Let two jobs differ in write volume or compute work, with C ≥ 2, and let δ be the firing gap measured on the section at which the leading job fires. On a branch along which the event ordering is unchanged, δ ′ = δ + (d2 − d1 ) + (T2 − T1 ),
(6)
a translation with no dependence on δ: the paired firing times advance by P2 − P1 per cycle, with Pi := Ti + di . That is the difference of the two free periods and not a claim that either job runs at its own: where the branch hypothesis fails, below, neither does. Two conventions have to be fixed for this to mean anything. The labels are the jobs themselves and not their ranks, so they do not exchange when the lagging job overtakes the leading one; and δ is not reduced modulo a period, there being two periods to choose between, so it grows without bound. The hypothesis on the event ordering is not vacuous, and what happens when it fails is instructive: at d1 = 0.1, d2 = 0.3 the writes overlap on 24 of the 106 cycles, and on those the two jobs are delayed by the same amount, exactly as Theorem 4 requires, but not on the same indexed cycle, their periods differing. Pairing firings by index therefore moves the increment to 0.1 or 0.3 on those cycles while its average over the run stays at P2 − P1 .
6
Proof. Rerun the proof of Lemma 3 with d1 = ̸ d2 . The common drained volume S is still the same for both, so di = S + |Wi \ Wj | now gives |W2 | − |W1 | = d2 − d1 : the write phase transports the lag and adds the volume difference. With C ≥ 2 every computing job advances at unit rate whether alone or not, so v1 = v2 = 1, the map (5) is the identity, and the compute phase likewise transports the lag and adds T2 − T1 . Summing the two gives (6). The hypothesis C ≥ 2 is not cosmetic. When the cap binds, v1 ̸= v2 , the two jobs no longer convert a work lead into a time lag at the same rate, and a heterogeneous pair acquires a genuine coupling even though both resources stay anonymous. That is the counterexample of §8, and it is the sharpest limit on how far this section generalises. The interpretation is the point of departure for everything that follows. Both Kuramoto theory and Mirollo–Strogatz theory are built on a pairwise interaction, a coupling function Γ(θi − θj ) whose shape decides whether the synchronous state attracts, and for two identical jobs that object is identically zero: contention charges them equal occupancy, which transports their work difference unchanged, and a preserved work difference translates the pair rather than rotating its relative phase.
4
Order three: the N -writer return map
Theorem 8 (Diagonal return map). Let N identical jobs (Ti ≡ 1, di ≡ d) under a cap C ≥ N , that is under storage contention alone, fire at times 0 = s1 < s2 < · · · < sN < d, so that all N write windows overlap. Write aj := sj+1 − sj for j = 1, . . . , N − 1. Then, with work-conserving sharing, the intervals between consecutive firings on the next cycle are a′j =
N −j aj , j
j = 1, . . . , N − 1.
−1 The map is diagonal, its spectrum is {λj = (N − j)/j}N j=1 , and
Q
(7)
j λj = 1.
Proof. First, the hypothesis sN < d does force all N windows to overlap: by f (n) ≤ n a job drains at rate at most 1, so no job can finish before a time d has elapsed since its own start, and in particular t1 ≥ s1 + d > sN . Hence every job has started, and none has finished, at time sN . Since the drain rate is 1/nw and is common to all active writers, the difference in drained volume between two jobs that are both writing is constant in time. On [sj , sj+1 ) exactly j jobs write, namely 1, . . . , j, none of which has finished by the previous paragraph, so job j drains aj /j while job j + 1 has not started; from sj+1 onward both are active whenever both are unfinished, so job j retains the lead Λj =
aj j
(8)
until it finishes. The leads are strictly positive, so jobs finish in their firing order, t1 < t2 < · · · < tN . At tj , job j has drained d, hence job j + 1 has drained d − Λj and has Λj remaining. On [tj , tj+1 ) the writers are exactly the jobs j + 1, . . . , N , that is nw = N − j, so job j + 1 drains at rate 1/(N − j) and tj+1 − tj = (N − j) Λj . Each job then computes for one unit of time: nc ≤ N ≤ C at every instant, so min{1, C/nc } = 1 and the compute duration is the same for all jobs whatever the order in which they finish writing. This is the only step that uses C ≥ N , and it fails as soon as the cap binds. The next firing times are therefore tj + 1 and the new intervals are a′j = tj+1 − tj , which is (7). Finally QN −1 j=1 (N − j)/j = (N − 1)!/(N − 1)! = 1. Corollary 9 (Reciprocal spectrum; synchrony is a fixed point, not an attractor). The map (7) preserves phase-space volume on its branch and its spectrum is reciprocal, λj λN −j = 1, with ⌈N/2⌉ − 1 expanding directions and as many contracting ones; for odd N these exhaust the spectrum, while for even N the mode j = N/2 has λ = 1. The synchronous configuration a = 0 is the apex of the closed cone {aj ≥ 0} 7
on which (7) is linear and is a fixed point of it, not an interior point of the branch, whose hypothesis is a strict ordering. The statement is therefore directional rather than a linearisation at a smooth point: a perturbation of synchrony falls in one of the N ! cones indexed by the order in which the jobs fire inside the burst, N of which share each of the (N − 1)! cyclic orders that Proposition 14 conserves; relabelling identical jobs exchanges those cones, and on each of them the map is (7). Since λ1 = N − 1 > 1 for N ≥ 3, every perturbation with a1 > 0 is amplified and no neighbourhood of a = 0 is contracted; for N = 2 the map is neutral, consistently with Theorem 4. More generally a cluster state, in which aj = 0 for the ranks inside each cluster, is invariant by exchange symmetry, and the cluster occupying ranks j, j + 1 tightens if j > N/2 and disperses if j < N/2. What det = 1 excludes is a fixed point of this branch map that is asymptotically stable in all of its directions; it does not exclude attraction to a cluster, a manifold of lower dimension being able to attract transversally while expanding along itself, and Corollary 10 exhibits exactly that. Proof. Everything except invariance is read off the spectrum of (7), which is a linear diagonal map on the branch and therefore has a = 0 as a fixed point with those eigenvalues. Invariance of a cluster state is exchange symmetry: identical jobs in identical states receive identical rates at every instant, so their states coincide for all time. Corollary 10 (Recruitment on an invariant cluster manifold). Let M1 := {a1 = 0} be the manifold on which the two leading jobs fire together. It is invariant, and the restriction of (7) to it acts on the surviving intervals with the same eigenvalues. For N = 3 that restriction is a′2 = a2 /2: the third job is recruited by the pair geometrically and the fleet tends to full synchrony, in infinite time and without any two jobs ever coinciding. Synchrony is therefore asymptotically stable relative to M1 while being unstable in the ambient space, which a unit determinant does not forbid. Two things confine the phenomenon and one does not. The transverse direction carries λ1 = N − 1 > 1, so M1 repels on this branch; and by Proposition 14 no trajectory whose jobs fire at distinct instants enters it, so the recruitment lives on a set of measure zero the dynamics cannot reach in finite time. What neither (k) (k) argument excludes is a1 > 0 at every k with a1 → 0, an asymptotic approach to a cluster along the other branches of the global map, on which we compute nothing; §6.4 bounds that by measurement in four cells, not by proof. Symmetrically, the trailing manifolds {aj = 0} with j > N/2 are transversally contracting on this branch, but λ1 > 1 carries any trajectory with a1 > 0 off the branch in finitely many cycles (Remark 13), so that contraction is not asymptotic on this branch either. Proof. The map is diagonal, so a1 = 0 gives a′1 = 0 and leaves the other coordinatesP as in (7); invariance is again exchange symmetry. For N = 3 the surviving eigenvalue is λ = 1/2, so 2 j aj decreases and P the branch hypothesis j aj < d of Remark 13 is preserved at every cycle, which makes the iteration (k)
(0)
legitimate for all time and gives a2 = 2−k a2 . The integrator returns that ratio to machine precision from a1 = 0, a2 = 0.1, d = 0.3. Remark 11 (What “volume preserving” does and does not say). Both det = 1 and the spectrum are properties of the single branch on which all N windows overlap and the cap does not bind. The global map is piecewise affine with many branches; we compute the determinant on no other, and do not claim the pieces glue into a globally measure-preserving map, which where the cap binds they demonstrably do not (§8). Nothing in §6 rests on this corollary. Proposition 12 (Volume preservation is anonymity, not fairness). Let the writers share an anonymous storage rule of total throughput f (nw ), each active writer receiving f (nw )/nw , with f positive and f (n) ≤ n. Under the hypotheses of Theorem 8 the map is still diagonal, with a′j =
(N − j) f (j) aj . j f (N − j)
Its spectrum is again reciprocal, λj λN −j = 1, and det = 1 for every such f . 8
(9)
N = 3 writers: the leading gap is stretched, the trailing gap squeezed a1
entry
job 1
a2
write
compute
job 2
write
compute
job 3
write
compute
exit
a 1 = (N−1) a1 0
0.0
0.2 0.4 0.6 time (units of the nominal cycle)
a 2 = a2/(N−1) 0
0.8
Figure 1: The mechanism of Theorem 8 for N = 3. Jobs enter their write phase at intervals a1 , a2 ; because a job that entered earlier banked a lead at a lower level of concurrency and cashes it out at a higher one, the exit intervals are a′1 = (N − 1)a1 and a′2 = a2 /(N − 1): the leading gap is stretched, the trailing gap squeezed, and their product preserved.
Proof. The condition f (n) ≤ n keeps every individual rate at most 1, so the event ordering established in the first paragraph of the proof of Theorem 8 is unchanged. The rate then enters in exactly two places. On [sj , sj+1 ) there are j writers, so job j drains at f (j)/j and banks Λj = aj f (j)/j; on [tj , tj+1 ) there are N −j, so job j +1 cashes that residue out at f (N −j)/(N −j), taking tj+1 −tj = Λj (N −j)/f (N −j), which is (9). In between, both jobs receive the same rate whenever both are unfinished, whatever that rate Q is, so the lead is unchanged. Anonymity is the only property of f used. Reciprocity is immediate, and j λj = 1 follows by pairing j with N − j. Remark 13 (Domain of validity).PEquation (7) holds for one cycle of a configuration satisfying sN < d, and may be iterated only while j a′j < d. Since λ1 = N − 1 > 1 for N ≥ 3, the leading interval grows geometrically and that condition fails after finitely many cycles: the fully overlapping region is transient, and the spectrum describes one branch of a piecewise-affine map rather than the asymptotics, which §6 measures instead. Two readings of (7) matter physically. The head disperses: a′1 = (N − 1)a1 , so a job firing slightly ahead of a pack is ejected from it. The tail compacts: a′N −1 = aN −1 /(N − 1), so groups formed at the back of the burst tighten. Equation (8) also shows where the pairwise cancellation of §3 fails: the lead Λj is banked at 1/j, set by how many other jobs were already writing, and cashed out at 1/(N − j), set by how many others still are, both third-party counts. With N = 2 they degenerate to j = N − j = 1 and (7) returns a′1 = a1 , recovering Theorem 4.
5
Two invariant structures
Proposition 14 (The firing order is invariant). Let N identical jobs run under Assumptions 1–3, with any cap C > 0 and any anonymous throughput rule. If ski < skj then sk+1 < sk+1 . The cyclic firing i j order is therefore a constant of the motion, all the intervals aj keep their sign, and no two jobs ever exchange rank or coincide. Proof. Two steps, each an application of anonymity. Write phase. Suppose ski < skj . On [ski , skj ) job i drains at a strictly positive rate and job j has not started, so i acquires a strictly positive lead in drained volume. Whenever both are writing they receive the same rate, so that lead is constant; whenever j writes and i does not, i has already finished. In either case eki < ekj . Compute phase. Two cases, since job i may reach T before job j has finished writing. If it does, it fires at sk+1 < ekj < sk+1 , job j having i j to finish write k and then a compute phase before firing again, and there is nothing to prove: this is 9
the only step at which one job laps the other. Otherwise i computes throughout [eki , ekj ), at the rate v(t) = min{1, C/nc (t)} > 0, while job j is still writing and by Assumption 1 accumulates no work; i therefore holds a strictly positive lead in accumulated work at ekj . From then on both receive the same v(t) whenever both compute, so the lead is preserved until i reaches T and fires. Hence sk+1 < sk+1 . i j Corollary 15 (Synchrony is invariant but unreachable). The synchronous set, and more generally each cluster manifold, is invariant (Corollary 9) and of zero Lebesgue measure, and by Proposition 14 it is reached by no trajectory that does not start in it. A fleet whose jobs fire at distinct instants keeps them distinct for all time; clustering, if it occurred, could only be asymptotic. Proposition 16 (Collision-free configurations: invariant, unreachable, of computable measure). Let N identical jobs run under a cap C ≥ N , and let C be the set of configurations whose N cyclic gaps between write starts are all at least d. On C no two write windows overlap, every job runs at unit rate in both phases, every cycle lasts exactly P0 = 1 + d, and every gap is constant for all time. The set is nonempty if and only if N d ≤ 1 + d; it is closed, and its interior is nonempty exactly when the inequality is strict, with Lebesgue measure N d N −1 µ(C) = 1 − (10) P0 relative to starts distributed uniformly over the cycle. Moreover no trajectory enters the interior of C from outside it. Proof. A job whose write starts at least d after the previous one and at least d before the next is the only writer for the whole of its write, so it drains at f (1) = 1 and finishes in exactly d; and nc ≤ N ≤ C makes every computing job advance at unit rate, so its compute lasts exactly 1. Every job therefore has period exactly P0 , the configuration advances rigidly, and the hypothesis is reproduced at the next firing, so F restricted to C is the identity and F (C) = C. For existence, the N cyclic gaps are nonnegative and sum to P0P , so all can be at least d if and only if N d ≤ 1 + d; the set they form is the closed simplex {gi ≥ d, i gi = P0 }, whose relative volume is (10) and whose interior is nonempty exactly when N d < 1 + d. Finally, at an interior state the write windows are pairwise disjoint, so no two events coincide and every rate equals 1; the evolution is locally a translation at unit speed with isolated events, and integrating it backwards from one section to the previous one gives a unique state, with the same disjoint windows. Any x with F (x) ∈ C is therefore that state and is collision-free. On the boundary, where some gap equals d exactly, the backward continuation is the one the tie-breaking convention selects. Three consequences. A stagger in which every write has clearance d is permanent, and exists exactly at and below the threshold (2); the margin available to a fleet spread evenly is the smaller quantity m = P0 /N − d of §7. A fleet that does not start collision-free never reaches the interior of C, in either time direction, so the numerics below is entirely a statement about the contending set, the boundary being left to the tie-breaking convention. And random launches land in C with probability (10), which is 9.2 × 10−2 at N = 8, L = 0.3 and 9.2 × 10−29 at N = 32, L = 0.9: at fleet scale a collision-free schedule has to be constructed, never encountered.
6
Numerical study
6.1
Protocol
All results below come from an exact event-driven integrator (sim/exact.py): between two events all rates are piecewise constant, so stepping from event to event is exact up to floating-point round-off. This is not a convenience, since a fixed-step integrator produces spurious drift of order the step size, which is precisely the quantity under test. It verifies the model as implemented, not the model against a facility; no quantity here is fitted to a measured trace. Ties between simultaneous events are broken by 10
Regime
Parameters
max |δ120 − δ0 |
Storage only Power cap binding Mixed Severe cap, C < 1 Beyond the proof, d > T Unequal volumes
d ∈ {0.05, 0.2, 0.45}, C = ∞ C ∈ {1.5, 1.05}, d = 0.05 C ∈ {1.5, 1.05}, d ∈ {0.2, 0.45} C ∈ {0.9, 0.7, 0.5}, d = 0.1 d ∈ {1.5, 2.5}, C ∈ {∞, 1.5} d1 = 0.1, d2 = 0.3
1.1 × 10−15 8.9 × 10−15 9.8 × 10−15 1.2 × 10−14 1.1 × 10−15 increment 0.2 on 82/106 cycles
Table 1: Pair neutrality, three initial gaps per cell spanning the cycle. Drifts are at the level of floating-point noise, confirming Theorem 4 across all three coupling regimes. The fifth row is outside the hypothesis d < T (Remark 5); the last confirms the translation term of Proposition 7 on the cycles whose writes are isolated, the other 24 carrying a delay shared equally but falling on cycles of different index.
processing write completions before write starts, a convention that selects one continuation on a set of measure zero, which is however exactly where the synchronous and cluster manifolds and the boundary of C sit; the results about those objects (Corollary 9, Propositions 14 and 16) are proved from the rate functions and hold whichever way ties are resolved. Tables 1 and 7 are the only experiments that scan the power cap; every other run is uncapped, the regime C ≥ N in which Theorem 8 holds. The predictions of §3–§4 were written down (PROTOCOL.md in the replication package) after an exploratory fixed-step sweep and before the confirmatory runs were inspected. We call this a written protocol and not a pre-registration: the file carries no third-party timestamp, so it documents the order in which we worked and does not certify it, and two of its predictions were subsequently corrected by the analysis rather than by the data. A random launch draws each job’s initial progress through its compute phase uniformly on [0, 1), so first write starts are uniform on [0, T ] and not on the cycle [0, P0 ]; the probability that such a launch is collision-free is (1 − (N − 1)d)N , which is 0.361 at N = 4, L = 0.3 against 0.375 for the uniform-over-the-cycle law of (10), and it is the former that governs the rejection rates reported here. Launches are conditioned on the complement of C. A launch inside it is rigid for ever (Proposition 16) and returns an exactly zero separation rate, so averaging the two populations together reports neither; the number of draws rejected is given per cell in Table 3. Stochastic quantities are reported as mean ± standard error over the retained launches, 60 per cell for the separation rate, 20 for the configuration and moment tables and 500 for the first-passage times of Table 6, with the spread across seeds given where it matters. Phases are defined as follows. Let ski be the instant of the k-th write start of job i, so that a job falling behind is not re-indexed; let P be the median of sk+1 − ski over the retained window; and set i k k k k ϕi = 2π(si − s̄ )/P mod 2π, with s̄ the mean over jobs. Unless stated otherwise the first half of each run is discarded.
6.2
Pair neutrality and the N -writer spectrum
Table 1 reports the drift of the firing gap over 120 cycles for two identical jobs. Table 2 checks Theorem 8 and Proposition 12 together: the Jacobian of the return map, measured by central differences, reproduces the predicted spectrum, and the determinant stays at 1 while individual eigenvalues move by more than a factor four between throughput rules. For N = 3 the measured Jacobian in firing-time coordinates is 2 0 3/2 1/2 to six digits, and the spectrum is independent of the configuration within the fully overlapping region, as (7) requires. Figure 2 plots the work-conserving case across four fleet sizes.
6.3
Separation of nearby configurations
Away from full overlap the dynamics is piecewise affine, and the combinatorics of who overlaps whom changes as intervals expand and contract. We measure the finite-time separation rate per cycle from the divergence of two trajectories initially 10−9 apart in one job’s phase, fitted over the window in which 11
λj = (N − j)/j (grey), measured (markers)
|λj|
101
N=4 N=6 N=8 N = 12
100
10−1 2
4 6 8 mode index j
10
Figure 2: Measured against predicted spectrum λj = (N − j)/j (grey) of the return map for N fully overlapping writers under the work-conserving rule, at one random configuration per fleet size; markers are the eigenvalue moduli of the Jacobian by central differences. Their reciprocal pairing about the dotted unit line is the unit determinant of Proposition 12, and the ⌈N/2⌉ − 1 markers above it the expanding directions of Theorem 8. throughput rule f (nw )
N
det J
maxj |λmeas − λj | j
1.00000 3.9 × 10−9 1.00000 5.6 × 10−9 1.00000 4.9 × 10−9 1.00000 4.1 × 10−9 1.00000 2.8 × 10−9 Table 2: Measured versus predicted spectrum (N − j)f (j)/ jf (N − j) of the return map for N fully overlapping writers. For N = 6 the leading eigenvalue ranges from 3.37 to 15.0 across these five rules and the determinant is 1 in every case; the last rule is not monotone in nw , so no smoothness or ordering of f is used. The map is affine here, so the residuals are the round-off floor of the difference quotient at ε = 10−7 . 1 (work-conserving) n−0.4 (degrading) 1/(1 + 21 (n − 1)) 1 + 0.3 ln n (improving) 2 if n even, else 1
3, 4, 5, 8, 12 4, 6 4, 6 4, 6 4, 6
the separation stays below 10−3 (Table 3). Separation means maxi |ski − s̃ki |, the largest discrepancy between corresponding write starts of the two trajectories, on absolute firing times: neither reduced modulo the cycle nor quotiented by the rotation that (11) divides out, both of which would bound a quantity we want to see grow. We call this a separation rate and not a Lyapunov exponent: it is a finite-time, finite-separation fit along one pair of trajectories per seed, in one perturbation direction, and no tangent-space calculation with jump matrices at the event surfaces is attempted here. Figure 3 shows one such pair of trajectories per cell. Three controls. The perturbed job does not matter: perturbing each in turn, over a fixed subsample of 20 launches, gives per-job means in [0.184, 0.197] at N = 8, L = 0.6 and [0.297, 0.318] at N = 16, L = 0.9, straddling the 0.188 and 0.305 that subsample returns when job 1 is perturbed. The spread across launches is wide, the 60 retained launches spanning [0.041, 0.284] and [0.196, 0.461] in those two cells, where it is one-signed. At L = 0.3, and at N = 4 throughout, it is not: up to a sixth of the contending launches return a rate that is zero to machine precision and three return a small negative one (Table 3), so nearby trajectories separating is a statement about most launches and not about all of them. Those null launches have a mechanism and not an exception. Every one of them spends at most 4 × 10−5 of its time with three concurrent writers, and exactly none in four of the five cells where such launches occur at all, so every collision along such a trajectory is a two-body collision and Theorem 4 leaves
12
trajectory separation
Divergence from a 10−9 perturbation N = 4, L = 0.6 N = 8, L = 0.6
101 10−1 10−3 10−5 10−7 10−9
N = 16, L = 0.6 N = 8, L = 0.9
fit window ends
0
20
40 cycle
60
Figure 3: Separation of two trajectories launched 10−9 apart in one job’s phase, one contending launch per cell, uncapped. Over the cycles plotted the separation gains 8.7 decades at N = 16, L = 0.6 and 7.3 at N = 8, L = 0.9, against 0.9 at N = 4, so the growth rises with fleet size and with load. The rates fitted on these four traces, 0.035 to 0.299 per cycle in the order of the legend, depart from their cell means (Table 3) by up to a factor 2.3: the spread within a cell is wide, so a single trace carries the trend in (N, L) and not the mean. Each curve ends at the last cycle every job has completed within the horizon, and the fits use the window below the dotted line. N
L = 0.3
L = 0.6
L = 0.9
4 8 16
+0.072 ± 0.007 (28 | 10) +0.110 ± 0.008 (1 | 9) +0.127 ± 0.008 (0 | 5)
+0.081 ± 0.007 (8 | 5) +0.200 ± 0.006 (0 | 0) +0.266 ± 0.006 (0 | 0)
+0.084 ± 0.008 (0 | 5) +0.227 ± 0.008 (0 | 0) +0.305 ± 0.008 (0 | 0)
Table 3: Finite-time separation rate λ per cycle, 80 cycles, uncapped, 60 launches per cell conditioned on the complement of C, mean ± standard error. In parentheses: the number of draws rejected as collision-free, and after the bar the number of retained launches whose rate is null (|λ| < 10−4 ), which §6.3 shows are the launches on which no instant carries three writers. Three launches at N = 4 return a rate below −10−4 , the most negative being −0.007. The rate grows with load and with fleet size and does not saturate between N = 8 and N = 16, so we do not extrapolate it. The control is Proposition 16: started from an evenly staggered collision-free configuration, the same measurement returns a separation that does not grow at all.
nothing to amplify. That condition is necessary and not sufficient, chains of two-body collisions not decomposing into independent pairs: the launches that also stay two-body but do separate return rates confined to [0.037, 0.053] across every cell, a population well below the cell means and distinct from it. The fit window does matter, and is the dominant uncertainty. On the same subsample the reported window gives 0.188 against the 0.200 of the full 60, while [10−6 , 10−3 ] and [10−4 , 10−2 ] give 0.289 and 0.251 at N = 8, L = 0.6, a spread of 50%; those two are reached within 80 cycles by 9 and by 1 of the 20 launches, so they are biased upward and the second is one trajectory. The value of λ should be read to within a factor of about 1.5, and only its sign and its trend in (N, L) are used below. That trend in L carries a further caveat: the rejection rule removes the collision-free draws, whose share falls as L grows (28 rejections at N = 4, L = 0.3 against none at L = 0.9), so the launches compared along a row are not drawn from the same conditional law and part of the increase is a change of population rather than of dynamics.
13
6.4
The order sector, and no memory resolved beyond it
Separation of nearby trajectories says nothing about where a single trajectory goes. Four measurements bound it, in the four cells of Table 4 and on the timescales stated; we have not scanned the cap, the intermediate loads, or launches built to be clustered. The configuration keeps its order, and nothing resolvable beyond it. Table 4 reports the per-cycle displacement of a job’s relative phase, its standard deviation over the run, the distance D to the launch configuration and the configuration autocorrelation ρ(ℓ) = ⟨cos(ϕk+ℓ − ϕki )⟩, with i D(ϕ, ψ) :=
N i h1 X 1 2 1/2 , min ϕi − ψi − α 2π 2π α∈[0,2π) N
(11)
i=1
where | · |2π is the representative in (−π, π] and the rotation α is divided out, that rotation being the time-translation symmetry of an autonomous system and not a reorganisation. The minimum is taken exactly, over the N Pcandidates that make the piecewise-quadratic objective stationary, and not at the circular mean arg i ei(ϕi −ψi ) , which minimises the chordal sum instead and overstates D by up to 0.04 of a cycle at these fleet sizes. Phases move by 1–6% of a cycle per cycle and fluctuate within 5–10% of a cycle of their own mean. The distance to the launch settles at 7–14% of a cycle, against the 22–26% separating two independent configurations, which reads as a fleet staying nearer its launch than chance would put it. That reading is wrong, and the reason is Proposition 14. Two independent configurations generally differ by a permutation of the jobs, and the cyclic firing order is a constant of the motion, so no trajectory can ever produce one from the other: the 22–26% measures a rearrangement the dynamics is forbidden to perform. The reference has to be drawn in the order sector the launch fell in. Taking it as an independent launch of the same law, relabelled into that sector, which is a benchmark for the complete loss of detectable memory inside the sector and not the farthest a reorganisation could carry the fleet, gives 0.124 ± 0.007, 0.138 ± 0.004, 0.099 ± 0.004 and 0.074 ± 0.003 in the four cells of Table 4, against 0.122, 0.136, 0.098 and 0.074 measured (sim/nullorder.py). Equality of two means is not equality of two laws, so that comparison is run three ways. The paired difference is −0.0013, −0.0018, −0.0003 and +0.0001 of a cycle, and against a declared equivalence margin of 0.01, a tenth of the separation at issue, 20 launches establish equivalence in the two larger fleets and bound the difference by 0.017 at N = 8. The two pooled distributions of D differ by a 1-Wasserstein distance of 0.0005 to 0.0029 of a cycle against an interquartile spread of 0.034 to 0.075, with p ≥ 0.74 under a permutation that swaps the two references launch by launch. The autocorrelation behaves the same way: the 0.64–0.89 plateau of Table 4 is 0.63–0.89 for the relabelled reference. What the fleet retains of its launch is its firing order, which is a theorem; beyond it these observables resolve nothing, at the power just stated and not as a proof of independence. The dynamics is therefore neither frozen, which would give ρ ≡ 1 and zero displacement, as the collision-free control does exactly, nor free to mix: the autocorrelation is already at its sector value by a lag of 50 cycles and is still there at 150. That also bounds what the separation rate of §6.3 can mean, the sector being invariant, so exponential separation saturates at the scale of that sector and not of the cycle. Ranks never change, and no gap is seen to close. In every cell of Table 4, over 20 launches and 150 cycles, the cyclic firing order is the same at every cycle, as Proposition 14 requires. Invariance of the order is compatible with all the gaps contracting towards zero without ever crossing, which is what asymptotic clustering would look like, so we follow the gaps by rank, which that invariance is what makes possible: rank j separates the same pair of jobs at every cycle. The minimum over ranks is stationary. Its mean over launches changes by +21%, +4%, −24% and +44% over 800 cycles, from 0.0137 ± 0.0025 to 0.0166 ± 0.0022 cycles at N = 8, L = 0.3 and from 0.0009 ± 0.0002 to 0.0013 ± 0.0003 at N = 32, L = 0.9, and paired launch by launch its ratio has a 95% interval inside [0.37, 2.9] in all 14
four cells: a contraction by more than a factor 2.7 is excluded over that span, a slower one is not. No particular pair tightens, though. The gap of the rank smallest at launch has a paired ratio confined to [2.2, 68], above 1 in every cell, and the rank realising the minimum has moved by the end of the run in 16 to 20 of the 20 launches: the small gaps are a rotating population and not a nucleating cluster. No coherence in the first four Daido moments. Assign Rm = |⟨eimϕ ⟩| [2] for m = 1, . . . , 4: the first moment alone would not settle the question, since two antipodal clusters, three equidistant ones and a splay state all give R1 ≃ 0 and are distinguished by the higher moments, with Rm = 1 at m equal to the number of clusters. The reference is not 0 but the finite-N value for independent √ uniform phases, √ which we resample at the same N rather than take from the Rayleigh limit π/(2 N ); the two differ by 0.5%, which is not negligible against the departures under test. The largest departure anywhere in the 36 cells of Table 5 is +3.2 standard errors (N = 8, L = 0.9, m = 3), and multiplicity alone does not dispose of it. The 36 statistics are not independent, four moments being read off the same configurations and three loads sharing their launch draws at fixed N ; sign-flipping whole launches, which resamples the joint null with that dependence in place (median absolute correlation 0.20), puts the 95th percentile of the maximum at 2.80 rather than the 3.72 that 36 independent Student statistics on 19 degrees of freedom give. What that departure does not survive is the choice of null. Drawing phases under the launch law of the runs, uniform on the compute phase and conditioned outside C, raises the floor at that cell from 0.315 to 0.328 and brings the departure down to +1.7 standard errors; and the paired test, which differences Rm over the run against Rm at the launch of the same run and so cancels the launch law exactly, returns −0.002 ± 0.042 there, its largest departure anywhere being +2.3 standard errors (N = 16, L = 0.3, m = 3) against a jointly resampled threshold of 2.84. Retaining the third and fourth quarters of the run separately changes no conclusion. What these moments exclude is a low-order cluster state or a splay; they are not precise enough to exclude a weak one, and a cell 3.2 standard errors above the uniform floor is the size of departure they leave open. The burst distribution is nonetheless not that of independent phases, and its tail is the heavier one. The quantity an operator sees is the number of jobs writing at the same instant. Conditioned on the effective write duty q the fleet realises, independent phases would give Binomial(N, q), and the measured distribution departs from it in both directions. At N = 8, L = 0.3 (q = 0.046) the fleet spends 7.0 ± 0.8% of its time with two or more writers against 5.0% for independent phases; at N = 32, L = 0.9 (q = 0.100) it spends 63.9 ± 1.3% against 84.5%, so ordinary crowding is rarer there. The upper tail goes the other way in every cell: the 99th percentile of the number of concurrent writers is 3, 5, 5 and 12 against 2, 4, 4 and 8 for independent phases, and the 99.9th is 4, 6, 6 and 15 against 3, 6, 5 and 9. A fleet that does not lock can still spend a thousandth of its time above the worst concurrency independent phases would ever produce, which is why §7 keeps the two claims apart. We compare time-weighted quantiles of the two laws and not sample maxima, which depend on how long each is observed. The departure is not an artefact of the launch: the same statistic on the launch configuration itself, before any dynamics, sits within 0.006 of its binomial value in all four cells, against departures of +0.020 to −0.206 after 300 cycles. Part of what remains is mechanical rather than dynamical, a write window lengthening precisely when it overlaps others, which inflates the time at high nw at fixed duty. These four say what the fleet does in these cells: it wanders over the sector its launch fixed, without freezing and without leaving it, and it does not cluster. That is why the moments sit at the incoherent floor. A random launch produces an incoherent configuration and the dynamics does not transform it, which is a statement about the map and not evidence that the invariant measure is uniform; the writer distribution shows directly that “no synchronisation” is not the same as “as if independent”. In particular the moments do not license calling the long-run phase distribution independent, and we do not. What the load parameter does not show. Sweeping the duty d upward without bounding the demand makes R1 rise steeply, which invites reading a synchronisation transition. It is queueing. Past 15
N
L
step/cycle
spread
D̄
D̄sector
ρ(1)
ρ(50)
ρ(150)
rank ch.
8 8 16 32
0.3 0.9 0.6 0.9
0.011 0.064 0.029 0.045
0.085 0.104 0.073 0.055
0.122 0.136 0.098 0.074
0.124 0.138 0.099 0.074
0.996 0.887 0.975 0.942
0.689 0.645 0.800 0.888
0.709 0.645 0.799 0.886
0 0 0 0
8 × 10−16
0
2 × 10−13
—
1
1
1
0
collision-free control
Table 4: Motion of the configuration, uncapped, 800 cycles, second half retained, 20 launches. “step/cycle” is the mean per-cycle change of a job’s relative phase and “spread” its standard deviation over the window; D̄ is the mean distance (11) to the launch and D̄sector the same distance to an independent launch relabelled into the trajectory’s order sector (sim/nullorder.py), both in units of the cycle. Standard errors over launches are at most 0.007 on the distances, 0.005 on the first two columns and 0.031 on the correlations. Two independent configurations whose relative order is left free are 0.220, 0.241 and 0.256 apart at N = 8, 16 and 32: the reference the fleet appears to beat, and does not. The largest distance reached within a run averages 0.206, 0.257, 0.194 and 0.175, so a trajectory does visit configurations farther from its launch than a fresh draw in its own sector, one reason that draw is a benchmark and not a bound. “rank ch.” counts cycles whose cyclic firing order differs from the previous one, over all launches, on a separate 300-cycle run. N
L
R1
R2
R3
R4
floor
8 8 8
0.3 0.6 0.9
0.306 ± 0.017 0.314 ± 0.014 0.319 ± 0.018
0.313 ± 0.014 0.332 ± 0.013 0.332 ± 0.014
0.312 ± 0.015 0.324 ± 0.009 0.343 ± 0.009
0.317 ± 0.016 0.322 ± 0.010 0.331 ± 0.007
0.315 0.315 0.315
16 16 16
0.3 0.6 0.9
0.216 ± 0.007 0.225 ± 0.006 0.233 ± 0.012
0.244 ± 0.012 0.217 ± 0.005 0.236 ± 0.011
0.218 ± 0.007 0.221 ± 0.005 0.232 ± 0.008
0.208 ± 0.006 0.220 ± 0.006 0.228 ± 0.006
0.222 0.222 0.222
32 32 32
0.3 0.6 0.9
0.154 ± 0.011 0.170 ± 0.008 0.176 ± 0.010
0.162 ± 0.009 0.166 ± 0.006 0.168 ± 0.007
0.171 ± 0.007 0.163 ± 0.006 0.163 ± 0.004
0.157 ± 0.007 0.163 ± 0.005 0.160 ± 0.004
0.157 0.157 0.157
Table 5: Daido moments below the stagger-feasibility threshold, uncapped, 300 cycles, second half retained; mean ± standard error over 20 launches, no draw having been rejected at these fleet sizes. The √ floor√is E[Rm ] for N independent uniform phases, resampled over 2 × 105 draws (the Rayleigh approximation π/(2 N ) gives 0.313, 0.222, 0.157). Under the launch law of the runs it rises to 0.324 and 0.328 in the two N = 8 cells discussed in the text, which brings the largest departure in the table down to +1.7 standard errors. A k-cluster state would show Rk → 1, a splay state every Rm well below the floor; neither appears at any (N, L, m). Moments up to m = 4 do not resolve five or more clusters, unequal or unequally spaced ones, or associations between particular jobs, so this is a test against low-order coherence and not against all of it.
L⋆ , which is 1.03 at N = 32, the fabric cannot serve the fleet within its free period and the effective cycle stretches to ≈ N d: the integrator gives Peff = 1.22, 3.46, 6.78, 11.31 at L = 1.05, 3.2, 6.4, 11.2, while R1 climbs from 0.28 to 0.61, 0.67 and back to 0.48. Every job now spends most of its time writing, so a high R1 records queue occupancy, and its non-monotonicity in L confirms that it is not measuring coherence. That leaves the synchronisation question meaningful above L⋆ and makes R1 the wrong observable there: any empirical claim of fleet synchronisation must control for L.
7
What this says operationally
In the uncapped homogeneous regime tested here, mutual contention does not phase-lock a fleet, which is not the same as leaving it quiet. Symmetric contention between identical jobs has exactly zero pairwise coupling (Theorem 4), cannot permute the firing order (Proposition 14), and has a three-body term that is volume preserving on the branch we can compute, with synchrony an unstable fixed point of it (Corollary 9). Exact synchrony is therefore unreachable in finite time, but 16
that leaves open an asymptotic approach to it, gaps shrinking without ever crossing, which happens on the cluster manifolds (Corollary 10) that no such fleet can enter, and which we exclude elsewhere by measurement and not by proof: no cell of §6.4 resolves a decay of the smallest gap over 800 cycles, the Daido moments stay within a change of null of their floor, and the firing order never changes. What those measurements do not support, once the invariance of that order is taken into account, is any memory of the launch beyond the order itself. Two restrictions travel with all of this. It concerns phase locking and not concurrency: the same runs put the upper tail of the number of concurrent writers above its independent-phase value in every cell (§6.4), so a fleet with no aligning mechanism can still burst harder than uncorrelated jobs would, and how to size for that is not a question settled here. And it concerns this model on the timescales simulated with the cap not binding: correlated bursts observed in production would then have to come from mechanisms it excludes. Those are external common causes, structural rather than dynamical (jobs launched together, checkpoint intervals set to the same round number of steps, wall-clock-aligned policies, correlated restarts after a shared failure); endogenous mechanisms excluded by assumption (asynchronous checkpointing, unequal allocation, heterogeneity behind a binding cap, delayed power control [5]); and the within-job alignment of thousands of ranks. The falsifiable content is a redirection: storm incidence should track launch-time and interval-setting statistics rather than the fleet’s history of past collisions. A stagger is permanent while the cap does not bind, and jitter is what ends it. An offset schedule in which every pair of consecutive write starts is at least d apart never degrades, and exists whenever L ≤ N/(N − 1). That is Proposition 16, and it assumes C ≥ N : what a binding cap does to a staggered fleet is open (§8), and it is the assumption a production fleet is least likely to satisfy. Under Assumption 3 nothing erodes the stagger. Relaxing that assumption is what gives the intervention a finite lifetime, and the mechanism is diffusive rather than dynamical: with per-cycle jitter Tik = T (1 + σzik ), the zik standard normal and independent across jobs and cycles, the work floored at zero, a fleet inside C has every job at its own free period, so each gap performs a driftless random walk of per-cycle variance 2σ 2 . The N walks are not independent: adjacent gaps share a jitter draw, so their increments have covariance −σ 2 , and the sum of the gaps is pinned to the cycle. The lifetime of a stagger of margin m is then a first-passage time of that correlated family, which still scales as (m/σ)2 in the margin. Table 6 measures it over 500 launches per cell: the median number of cycles to the first overlap is 0.11 to 0.24 times (m/σ)2 , over two decades in (m/σ)2 and a factor four in N : thirteen of the sixteen cells of Figure 4 fall in a band of unit slope on log axes although m and σ vary separately by factors of 4 and 8, which is the content of the scaling, and the three above it sit on the one-cycle floor of the measurement. That coefficient is not a constant of the model: it falls with N and depends on the jitter law through more than its variance, so it summarises the cells tested and is not a formula. What the quadratic scaling tests is narrower than it appears. Before the first overlap the fleet is inside C, where no two writes contend, so Theorem 4 is not what keeps the walk driftless there: each gap is a difference of two independent jitters and would diffuse the same way whatever the dynamics does after a collision. What the exponent does exclude is a drift in the gaps under jitter alone, which Proposition 16 does not cover, being a statement about the deterministic flow. A systematic drift µ would give a passage time of order m/|µ| rather than (m/σ)2 , and none is seen over two decades. This replaces the natural but incorrect recipe of re-staggering on the e-folding time ln(1/ε)/λ of the separation rate, which is the time for two nearby trajectories to diverge and not, as §6.4 shows, the time for one trajectory to lose its structure. The operator’s quantity is the first passage of the margin to zero, set by the jitter budget and not by λ. A median is not a guarantee, so the useful form is a survival probability. At N = 8, L = 0.3, where an even stagger affords m = 0.092, a jitter of 1% leaves the schedule intact for 10 cycles with probability 0.78 and for 50 with probability 0.01; halving the jitter raises those to 1.00 and 0.60. These are proportions over 500 launches, so their binomial standard error is at most 0.022. Buying a hundred cycles at 1% jitter would take a margin of 0.24 of a 17
N
L
margin m
σ = 0.020
0.010
0.005
0.0025
8 8 16 32
0.3 0.6 0.3 0.3
0.0922 0.0594 0.0449 0.0222
5 (0.24) 2 (0.23) 1 (∗) 1 (∗)
16 (0.19) 7 (0.20) 4 (0.20) 1 (∗)
56 (0.17) 25 (0.18) 11 (0.14) 3 (0.15)
226 (0.17) 92 (0.16) 42 (0.13) 9 (0.11)
median cycles to first overlap
Table 6: Lifetime of an evenly staggered collision-free schedule under per-cycle jitter of the compute time: median cycles to the first overlap over 500 launches, and in parentheses the ratio to (m/σ)2 . Uncapped, margin m = P0 /N − d, jitter standard normal and independent across jobs and cycles. Cells marked ∗ sit at the one-cycle floor of the measurement, where the median is not resolved and the ratio is meaningless. Elsewhere it lies in [0.11, 0.24] and decreases with N , the expected effect of taking a minimum over N walks, correlated through the jitter draws they share. The distribution is wide, as a first-passage time is: at N = 8, L = 0.3, σ = 0.005 the 10th and 90th percentiles are 30 and 110 cycles. No run was censored at the simulation budget.
Lifetime is set by (m/σ)2, not by λ N = 8, L = 0.3 N = 8, L = 0.6 N = 16, L = 0.3 N = 32, L = 0.3
102
ratio 0.11–0.24
101 100 100
101 102 jitter budget (m/σ)2
103
Figure 4: The same 500 launches per cell as Table 6, plotted against the jitter budget (m/σ)2 ; the band is the reported range [0.11, 0.24] of the ratio. Hollow markers are the three cells whose median sits at the one-cycle floor, above the band because a median cannot fall below one cycle and not because the scaling fails there. The residual ordering within the band is the decrease of the coefficient with N .
cycle, which no fleet of more than four jobs can offer, since m = P0 /N − d falls as 1/N and is already 0.045 at N = 16 and 0.022 at N = 32. At scale it is the refresh interval and not the margin that is the adjustable quantity. Fleet size cuts both ways. The leading expansion rate on the fully overlapping branch is N − 1, and the measured separation rate rises with N throughout Table 3. Working against that, the collision-free set survives to arbitrary N but on ever tighter terms: its threshold falls to 1 and the margin per job as 1/N , so its lifetime under jitter falls with the m2 ∼ 1/N 2 of the margin, and the measured coefficient decreases with N on top of that. Three fleet sizes, three of whose cells sit on the resolution floor, do not establish an exponent, so we report the direction and not a scaling.
8
Limitations
8.1
A boundary made explicit: heterogeneity behind a binding cap
Theorem 4 assumes identical jobs and Proposition 7 assumes a cap that does not bind. Neither hypothesis can be dropped. Take T1 = T2 = 1, d1 = 0.2, d2 = 0.4, a detuning ∆ = 0.2, and a cap 18
cap C
v1 = min{1, C}
v2 = min{1, C/2}
∞, 2.0 1.99 1.8 1.5 1.2 1.0 0.7
1 1 1 1 1 1 0.7
1 0.995 0.9 0.75 0.6 0.5 0.35
κ measured
1/v2 − 1/v1
0.0000 0.0050 0.1111 0.3333 0.6667 1.0000 1.4286
0.0000 0.0050 0.1111 0.3333 0.6667 1.0000 1.4286
Table 7: Pairwise coupling for two heterogeneous jobs (d1 = 0.2, d2 = 0.4) behind a binding cap, measured at δ = 0.1 by central differences on the exact integrator (sim/hetero.py). The coupling is exactly zero for C ≥ 2, where the cap never binds with two jobs, and grows as the two compute rates separate. Anonymity holds throughout: it is not sufficient for phase-blindness once the units differ.
C = 1.2. Both resources remain anonymous, neither rule being able to name a job, yet the gap map is not a translation: the integrator gives δ = 0.05 7→ 0.283, 0.10 7→ 0.367, 0.15 7→ 0.450, against the pure translation δ + ∆ that Proposition 7 gives at C ≥ 2. The increment depends on δ, which is a coupling in the sense of Definition 2. Two features make it interpretable. The coupling is supported on gaps below the detuning: for δ ≥ ∆ the increment returns to a constant, 0.333 in this cell, since the faster writer has then finished and resumed computing before the slower one is delayed by it. The phase response is thus carried by an ¯ and it disappears as the fleet becomes homogeneous, interval of width ∆ in a cycle of length ≈ 1 + d, consistently with Theorem 4. And its strength is set by the gap between the two compute rates: writing κ := dg/dδ for the slope of the increment on that support, Table 7 finds κ=
1 1 − v2 v1
(12)
to five digits at every cap tested, vanishing exactly at C = 2 where v1 = v2 . We report (12) as a measured law for this two-parameter family and not as a theorem: the mechanism is clear, the two jobs converting a work lead into a time lag through different rate sequences once their write durations differ, but the coefficient is not derived in general, and neither a fixed point of the resulting map nor its stability is established. What this bounds is the reach of Theorem 4: order-two phase-blindness is a statement about identical units, and a real fleet is never exactly homogeneous.
8.2
The remaining assumptions, and what is not covered
Blocking checkpoints (Assumption 1) is the fragile one. Under asynchronous checkpointing the job keeps computing while its state is flushed, and the cycle length becomes max{T, write duration}. Two things break at once: the phases stop partitioning time, so (3) fails and with it the step that turns equal writing time into equal computing time; and the maximum is a threshold nonlinearity, so a job whose write is stretched past T is delayed while one whose write finishes early is not. We expect a pairwise coupling to reappear where contention pushes the write past T , but we have not computed it and the expectation is not a result: a first attempt produced drifts that are integer multiples of the period, that is skipped checkpoints rather than phase slip, and settling whether a checkpoint that misses its slot is dropped or queued changes the model before it changes the answer. This is the first extension to compute, and the regime of modern stacks. A binding power cap is outside the N -body result. Theorem 4 and Proposition 14 hold for identical jobs at every C > 0, but Theorem 8 needs C ≥ N : when the cap binds, jobs that finish writing at different instants compute at rates that depend on how many others have already finished, and the compute durations cease to be equal. Measuring the Jacobian numerically at N = 4, d = 0.45 gives det J = 1.50 at C = 2 and det J = 24 at C = 1, against 1 for C ≥ N : on that branch the map is 19
strongly expanding rather than volume preserving, and the unit determinant of §4 is a property of the storage channel alone. Every run in §6 is uncapped for this reason, and what a homogeneous fleet does under a binding cap is open. It is the most consequential gap here, since production fleets are capped. Unequal allocation restores coupling; degrading throughput does not. Remark 6 and Proposition 12 cover any anonymous rule, monotone or not. Priorities, weighted classes and per-job bandwidth caps break anonymity and are expected to produce a coupling of first order in the weight difference; we have not derived it. Determinism matters because the map is marginally stable in the pairwise sector, which is what §7 quantifies: jitter turns the neutral direction into a random walk, phase gaps diffuse rather than lock, and collisions recur on the diffusive timescale. For small fleets this competes with, and may dominate, the separation measured in §6.3. Four further items are outside scope. The intra-job burst: the model collapses each job to one scalar writer, so a fleet that never synchronises across jobs can still produce correlated I/O for that reason alone, and separating the two experimentally means conditioning on the number of concurrently checkpointing jobs, not on aggregate bandwidth. Network collectives, which couple ranks within a job and could set the effective Ti , are absent. Electrical response: the facility is not modelled, so the megawatt-scale figures of §1 are context and not calibration. Scale and ergodicity: the fleets reach N = 32, well below production scale, and no mixing is established in the ergodic-theoretic sense, which would need a full Lyapunov spectrum, a decay-of-correlations estimate and a global invariant measure.
9
Relation to the companion paper, and an open question
The companion paper [5] asks whether co-located training jobs phase-lock behind a shared power envelope, and answers that they can: load-dependent throttling is a coupling channel whose sign is set by the lag of the control loop. Its throttle is a controller with memory, acting on aggregate demand after a lag τ , so the rate a job receives at t depends on the state of the fleet at t − τ . The cap here is v = min{1, C/nc (t)}, an instantaneous function of current occupancy and so the τ → 0 limit of that family; in that limit the companion’s coupling coefficient, proportional to − sin(mω̄τ ) in the m-th harmonic, vanishes, and the two statements agree where their domains meet. What is added here is that the vanishing is exact rather than leading-order, and that its cause is not the smallness of a parameter but anonymity together with exclusive phases, through Lemma 3: the coupling in [5] is a property of the controller, not of the sharing. Everything above concerns one storage fabric, one power cap and two phases per cycle, and the mechanism used little of that: Theorem 4 used that the resource cannot tell two colliding users apart, that its rate depends on the present occupancy alone, and that a unit is active on exactly one resource at a time. Two hypotheses are not decoration and travel with the statement: the phases must partition time, and the section of Definition 2 must recur with a constant offset, which is what d < T buys here. Conjecture 17 (Anonymity forbids pairwise coupling). Let a population of identical integrate-and-fire units interact only through anonymous resources that are memoryless, that is whose rates depend on the present occupancy alone, in any number and with arbitrary rate functions, through any number of phases per cycle, with each unit active on exactly one resource at any instant, and suppose the section of Definition 2 recurs along every trajectory with a constant offset. Then the coupling is phase-blind at order two. The conjecture is deliberately confined to order two. Whether the leading nonvanishing interaction is then of order three, and whether the three-body map preserves volume on the fully overlapping branch, are separate questions that do not follow from it and that we do not conjecture. Two of the three structural hypotheses are known to be necessary rather than convenient, the recurrence of the section being a regularity condition. Identical units: §8.1 exhibits an anonymous, memoryless, exclusive-phase system with a nonzero pairwise coupling as soon as two units differ and 20
the cap binds. Memoryless resources: [5] exhibits an anonymous resource with a lag whose coupling is nonzero and whose sign is tunable. Exclusive phases we believe necessary, on the strength of the broken proof step of §8, but we have exhibited no counterexample, and a broken proof is not a false conclusion. The obstruction is identifiable, which is the reason for stating the conjecture. The occupancy identity (4) closes because Assumption 1 makes the two phases partition time: equal total writing time is read off the write phase and transported to the compute phase, where it becomes equal solo work and hence an invariant work difference. With M resources and p mutually exclusive phases the same bookkeeping gives p occupancy identities and one partition identity, and nothing forces the solo intervals on resource r to match those on resource r′ unit by unit: the single scalar ∆x must be replaced by an object that survives the combinatorics of which unit is where. The conjecture also has a reading testable long before it is settled: if it holds, a memoryless anonymous resource cannot steer relative phase. Such a resource still fixes everything else, the absolute rate, the mean period, the duty cycle and the queue, so uncontrollable is the wrong word for it; what it cannot do is act on the gap between two identical users as a function of that gap. An operator who wants to steer relative phases through the resource a fleet shares must therefore break anonymity, add memory, or exploit the heterogeneity that is already there.
10
Conclusion
Treating checkpointing jobs as pulse-coupled oscillators makes a counter-intuitive prediction that we prove and verify within the model: between identical jobs, symmetric contention has no pairwise coupling at all, it cannot permute the firing order, and its three-body coupling preserves volume on the branch where it can be computed, which makes synchrony a fixed point with ⌈N/2⌉ − 1 expanding directions rather than an attractor. A cluster does recruit, but only inside the manifold on which it already exists (Corollary 10), which a fleet firing at distinct instants cannot enter. Sharing one storage fabric therefore gives such a fleet no mechanism, within this model, that would drive it into phase locking of its own accord, which is the sense in which its checkpoint storms are not self-reinforcing; its bursts are not thereby milder than uncorrelated ones, the upper tail of concurrency being the heavier in every cell measured. What it does instead depends on how it was started and on very little else: spread widely enough that no two writes overlap it is rigid, a regime available exactly below the stagger-feasibility threshold, reachable from nowhere else and conditional on the cap not binding; started outside it, in the cells we measured, it neither locks nor clusters over hundreds of cycles, and the only memory of its launch these observables resolve is the firing order that anonymity freezes. Whether any of this depends on checkpointing is the open question of §9 rather than a result: the proof uses two exclusive phases, a deterministic amount of work per cycle and d < T , and Conjecture 17 is precisely the claim that the first of those is not essential. The boundaries are as informative as the statement, and two of the three hypotheses can be removed with the coupling returning: with memory it returns with a tunable sign [5], and with heterogeneity behind a binding cap at O(1) on a support set by the detuning. What follows for practice is that a stagger is worth constructing, that it is permanent in the deterministic uncapped model, and that the quantity governing its refresh is the jitter budget through (m/σ)2 rather than any exponent of the free dynamics. Reproducibility. The replication package, submitted as ancillary files with this preprint, contains the integrator including its tie-breaking convention, its launch law, its rejection rule and the distance (11) with the grid check of its rotation minimisation (sim/exact.py), the written protocol (PROTOCOL.md), and one script per table: test_p1 (Table 1), spectrum (2), dynamics (3), checks (4), nullorder (its order-sector reference and the equivalence tests), order (5), burst (the writer distribution of §6.4 and the lifetimes of Table 6), hetero (7), with frozen covering §5 and §6.4. All runs are seeded, require only numpy, and python sim/all.py regenerates every number quoted here, under Python 3.14.0 and numpy 2.4.6. The four figures are produced separately by sim/figures.py, which additionally requires 21
matplotlib.
References [1] Carmen C. Canavier and Ruben A. Tikidji-Hamburyan. Globally attracting synchrony in a network of oscillators with all-to-all inhibitory pulse coupling. Physical Review E, 95(3):032215, 2017. doi: 10.1103/PhysRevE.95.032215. [2] Hiroaki Daido. Onset of cooperative entrainment in limit-cycle oscillators with uniform all-to-all interactions: bifurcation of the order function. Physica D: Nonlinear Phenomena, 91(1–2):24–66, 1996. doi: 10.1016/0167-2789(95) 00260-X. [3] Aaron Grattafiori et al. The Llama 3 herd of models, 2024. arXiv:2407.21783. [4] Dillon Jensen, Obi Nnorom Jr., Grant Wilkins, Hugo Budd, Ram Rajagopal, Juan Rivas-Davila, and Phil Levis. EasyRider: Mitigating power transients in datacenter-scale training workloads, 2026. arXiv:2604.15522. [5] Brieuc Le Roux Tardif. Do co-located AI training jobs synchronize? load-dependent throttling as a coupling mechanism for phase-locking behind a shared power cap, 2026. arXiv:2607.19638, submitted 22 July 2026. [6] Renato E. Mirollo and Steven H. Strogatz. Synchronization of pulse-coupled biological oscillators. SIAM Journal on Applied Mathematics, 50(6):1645–1662, 1990. doi: 10.1137/0150098. [7] Saurabh Mishra, Meet Vadakkanchery, Pradeep Fernando, Saiteja Samudrala, Gerson Kroiz, Jingxin Ye, and Viacheslav Kovalevskyi. Distributed checkpoint: Efficient checkpointing in large-scale jobs. https://pytorch. org/blog/distributed-checkpoint-efficient-checkpointing-in-large-scale-jobs/, 2025. PyTorch blog, 11 September 2025; consulted 2026-07-28. [8] OPAL-RT. AI workload variability and its impact on data center power stability. https://www.opal-rt.com/blog/ ai-workload-variability-and-its-impact-on-data-center-power-stability/, 2026. Vendor technical blog, 22 March 2026. Cited for the sub-second timescale, not for magnitudes. [9] SemiAnalysis. AI training load fluctuations at gigawatt-scale: Risk of power grid blackout? https://newsletter. semianalysis.com/p/ai-training-load-fluctuations-at-gigawatt-scale-risk-of-power-grid-blackout, 2025. Industry analysis, 25 June 2025. [10] Maxwell Twelftree, David Lemphers, An-chi He, and Yue Yang. Not every sync is safe: Calibrated DiLoCo scheduling for shared AI infrastructure, 2026. arXiv:2607.02544. [11] Geoff Werner-Allen, Geetika Tewari, Ankit Patel, Matt Welsh, and Radhika Nagpal. Firefly-inspired sensor network synchronicity with realistic radio effects. In Proceedings of the 3rd ACM International Conference on Embedded Networked Sensor Systems (SenSys ’05), pages 142–153, San Diego, CA, 2005. doi: 10.1145/1098918.1098934.
22