Conceptio › Archive › arXiv CS
arXiv CSopen access

LoopSpec: Pipelined Self-Speculative Decoding for Looped Transformers

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Preprint.

L OOP S PEC : P IPELINED S ELF -S PECULATIVE D ECOD ING FOR L OOPED T RANSFORMERS SangLyul Cho1∗ Langqing Cui2∗ Sehoon Kim2 2 Seoul National University KAIST

Dongsu Han2

Insu Han2†

1

arXiv:2609.17184v1 [cs.LG] 15 Sep 2026

A BSTRACT Looped Transformers achieve strong performance with compact parameter sizes by repeatedly applying a shared stack of Transformer blocks across recurrent depths. However, they incur higher decoding latency than standard Transformer models of comparable parameter size because shared weights are accessed at every recurrent depth. To improve decoding efficiency, self-speculative decoding is particularly well suited to Looped Transformers, as their intermediate recurrent states can directly provide draft predictions without an auxiliary draft model. We therefore propose L OOP S PEC, a training-free self-speculative decoding framework tailored for Looped Transformers. L OOP S PEC extracts draft tokens from early recurrent states and operates in a pipelined manner, overlapping draft generation of future tokens with target verification of the current token. To improve draft accuracy without excessive compute overhead, we introduce a selective second proposal from deeper recurrent depth while ensuring lossless decoding under both greedy and sampling regimes. Furthermore, we derive the optimal proposal depths in closed form and show the prediction matches measurement. Across reasoning and coding benchmarks, L OOP S PEC achieves up to 6.83× inference speedup across diverse Looped Transformers. € Project Page : https://langq1225.github.io/loopspec/ § Code : https://github.com/kaist-flexml-lab/loopspec

1

I NTRODUCTION

Looped Transformers (Geiping et al., 2025; Zhu et al., 2025; McLeish et al., 2026; Nanbeige Lab et al., 2026) have recently emerged as an efficient architectural paradigm that iteratively applies shared Transformer blocks across recurrent depths, drastically reducing the total parameter count. For instance, a model of four blocks applied 8 times achieves the effective depth of a 32-layer model while storing only four layers of parameters. By expanding computational depth through this recurrence, Looped Transformers match or outperform standard Transformers of comparable parameter count on reasoning tasks. However, executing shared Transformer blocks across multiple recurrent depths repeatedly fetches identical parameters within each token generation step, leaving inference memory-bandwidth bound thus increasing decoding latency (Hooper et al., 2023). While speculative decoding with a separately trained draft model (Leviathan et al., 2023; Chen et al., 2023; Li et al., 2025; Chen et al., 2026) has emerged as a promising approach for lossless inference acceleration, it introduces additional training costs, potential distribution mismatch between draft and target models, and extra memory. Self-speculative decoding (Zhang et al., 2024; Elhoushi et al., 2024; Cha et al., 2026) avoids these limitations by using intermediate computations of the target model itself to generate draft tokens, eliminating the need for a separate draft model. Looped Transformers are naturally suited to this paradigm, since the intermediate representation in each recurrent depth forms a sequence of increasingly refined proposals toward the final prediction. To this end, we propose L OOP S PEC, a training-free self-speculative pipelining framework specifically designed for Looped Transformers. The key idea is to exploit the token predictions from the ∗ †

Equal contribution. Corresponding author.

1

Preprint.

Prompt : The Capital of France is Ground Truth Generation : Paris and it is also the (a) Vanilla Looped Transformer 1

2

3

3 tokens in 12 recurrents

4

5

6

7

8

9

10

11

12

Paris

(c) LoopSpec w/ 1st & 2nd proposals 3 tokens in 7 recurrents 1

2

France

Paris the

and

3

1

2

3

France

4

5

Paris

✗

6

7

8

9

PRUNED!

city

r=1

r=2

r=3

r=4

Recurrent Depth r

of

the

PRUNED!

and

and

and it

✔ it

it

the

accepted token

✔ …

is

…

also the

it city

…

the

active branch

is in

is

also

✔ PRUNED! …

was

pruned branch also

✔

gate closed, no second proposal

Paris and

7

of

Notations

the

6

city city

3 tokens in 9 recurrents

5

✗ ✔

a

it

(b) LoopSpec w/ 1st proposal only

4

Paris

the

…

the

…

the

…

Figure 1: L OOP S PEC pipelines decoding over R = 4 recurrent depth. (a) A vanilla Looped Transformer finishes all R recurrences of one token before starting the next. (b) L OOP S PEC with a first proposal only at d1 = 1: an early-depth draft starts the next token immediately, but a rejection stalls the pipeline (Section 4.2). (c) L OOP S PEC with first proposal at d1 = 1 and a residual second proposal with gating mechanism at d2 = 2: the second proposal creates a fallback branch, causing fewer stalls in the pipeline (Section 4.3).

intermediate recurrent depths. Rather than waiting for all recurrent steps to finish, L OOP S PEC drafts a token from an early recurrent state and immediately starts computing the subsequent token conditioned on this draft (Figure 1(b)). The original computation and this speculative continuation use the same Transformer blocks, so they can be processed together in one batch even at different recurrent depths. Once the original computation reaches its final depth, it verifies the early draft. If the draft is accepted, some recurrent steps for the subsequent token have already been completed, so fewer steps are needed before it can reach the final depth. Otherwise, we prune the speculative draft and restart from the verified prefix. Since deeper recurrent computations progressively refine the model’s prediction, L OOP S PEC uses a deeper recurrent state to produce a second proposal and starts an alternative continuation before verification. Accepting this second proposal preserves its progress when the first proposal is rejected (Figure 1(c)). The resulting computation forms a pipeline in which the target computation and multiple speculative continuations advance in parallel, substantially increasing hardware utilization while preserving the exact target-model distribution. Figure 1 provides an overview of the complete L OOP S PEC pipeline. Our primary contributions can be summarized as follows: • We propose L OOP S PEC, a training-free self-speculative pipelining framework that accelerates Looped Transformers without external drafters or structural modifications. To improve draft acceptance while limiting computational overhead, L OOP S PEC introduces (i) a residual second proposal from a deeper recurrent state to recover from first proposal rejection, and (ii) a gating mechanism that creates the fallback speculative continuation only when the deeper state disagrees with the first draft (Sections 4.2 and 4.3). • We prove that the restart probability depends only on the second proposal depth, not the first. Using the empirically observed power-law decay of the intermediate-to-target TV distance, we derive the optimal proposal depths in closed form matching the measured optima (Section 5). • We empirically demonstrate up to 6.83× lossless speedup across seven checkpoints from two Looped Transformer families through an SGLang (Zheng et al., 2024) implementation with branch-wise KV management and CUDA Graph optimization, outperforming the training-based DFlash (Chen et al., 2026) by up to 1.7× (Section 6 and B). 2

Preprint.

Layer 5

LM Head

Post Layer 4

Layer 3

Layer 2

r = 4 Layer 4

Layer 2

Layer 3

Layer 3

r = 3 Layer 4

Layer 2

r = 2 Layer 4

Layer 2

Layer 1

Embedding

Layer 3

r = 1

Pre

optional: present in Raven, absent in Ouro

Figure 2: Overview of a Looped Transformer with R = 4 recurrent steps. In Ouro models (Zhu et al., 2025), the pre-layers P and post-layers C contain only the Embedding layer and LM Head, respectively, whereas in Raven models (McLeish et al., 2026) both P and C contain several Transformer layers.

2

R ELATED W ORK

Speculative Decoding. Speculative decoding (Leviathan et al., 2023; Chen et al., 2023) is a lossless inference acceleration technique that uses a lightweight drafting mechanism to generate multiple candidate tokens, which are then verified in parallel by the target model. Drafting mechanisms typically employ either a smaller pretrained model from the same architecture family (Leviathan et al., 2023) or a specially trained external drafter. Modern state-of-the-art drafters, such as EAGLE-3 (Li et al., 2025) and DFlash (Chen et al., 2026), achieve high acceptance rates by conditioning proposals on intermediate target representations. Self-speculative decoding instead generates draft tokens using the target model itself without a separate external drafter. Draft & Verify (Zhang et al., 2024) obtains drafts by skipping intermediate layers, and LayerSkip (Elhoushi et al., 2024) utilizes early exits. Looped Transformers. Looped Transformers (Geiping et al., 2025; Zhu et al., 2025; McLeish et al., 2026; Park et al., 2026; Nanbeige Lab et al., 2026) repeatedly apply shared Transformer blocks across recurrent depths, increasing effective depth without proportionally increasing the number of parameters. While they achieve higher task accuracy than standard Transformers of comparable size, their recurrent computation incurs substantially greater memory traffic during inference, limiting decoding throughput. Recent work reduces recurrent inference cost by modifying the model architecture or adapting the number of recurrent steps. LT2 (Deng et al., 2026) substitutes full attention with linear or sparse attention variants and distills hybrid checkpoints, while Think-at-Hard (Fu et al., 2026) incorporates a routing controller with depth-aware adapters to adjust recurrent passes dynamically. Rather than altering attention or adding routing heads, SPEED (Hooper et al., 2023) explores speculative pipelining across cyclically shared decoder groups. However, SPEED requires a custom training process and only supports greedy decoding.

3

P RELIMINARIES

3.1

L OOPED T RANSFORMERS

For a sequence prefix X = x1:n of length n, the Looped Transformer computation can be abstracted into (1) pre-layers P, (2) recurrent-layers B, and (3) post-layers C: (0)

HX = P(X),   (r) (r−1) HX = B HX ,   (R) pR (· | X) = C HX . (0)

(r)

r = 1, . . . , R,

(1)

Here HX and HX denote hidden states before recurrence and at recurrent depth r, respectively; r indexes recurrent depth, and R denotes the total recurrent depth of the model. xn+1 ∼ pR (· | X) is the target next-token probability distribution. The pre-layers P include the token embedding, and the post-layers C incorporate the LM head as well as temperature scaling, optional top-k/top-p filtering, and softmax normalization to output a probability distribution over the vocabulary V. 3

Preprint.

More generally, applying the post-layers at any recurrent depth r yields an intermediate next-token (r) distribution pr (· | X) = C(HX ), of which the target distribution pR is the special case r = R. Existing Looped Transformers instantiate this abstraction in different ways. Ouro (Zhu et al., 2025) places all Transformer blocks inside B, whereas Raven (McLeish et al., 2026) places only a subset of intermediate blocks in B and the remaining ones in P and C (Figure 2). 3.2

R EJECTION S AMPLING AND S PECULATIVE D ECODING

For any two distinct probability distributions µ and ν over a vocabulary V, we define their elementwise positive difference as [µ − ν]+ (v) := max(µ(v) − ν(v), 0) for each v ∈ V, and the residual distribution operator (µ ⊖ ν) as: [µ − ν]+ (v) (µ ⊖ ν)(v) := P . (2) u∈V [µ − ν]+ (u) The core idea of speculative decoding (Leviathan et al., 2023; Chen et al., 2023) is to draft fast and then verify via rejection sampling. We state the rule for a single token position, which is the form L OOP S PEC builds on. Formally, let X be the current prefix, and let p(· | X) and q(· | X) denote the target and the proposal distribution of the next token, respectively. A candidate x e is drawn from the proposal distribution q(· | X) and accepted with probability   p(e x | X) min 1, . (3) q(e x | X) If the candidate x e is rejected, the next token is instead resampled from the residual distribution:   x ∼ p(· | X) ⊖ q(· | X) . (4) Leviathan et al. (2023) show that the token produced by Equations (3) and (4) is distributed exactly as a token sampled from the target p, so the procedure is lossless. Crucially, losslessness holds for any valid proposal distribution q.

4

L OOP S PEC : P IPELINED S ELF -S PECULATIVE D ECODING FOR L OOPED T RANSFORMERS

We present L OOP S PEC, a training-free self-speculative decoding framework for Looped Transformers. We first show that early recurrent states can provide effective draft proposals (Section 4.1). Building on this observation, we introduce a pipeline with a single proposal depth d1 < R, where each draft starts computation for the next token before final-depth verification (Section 4.2). As decoding is memory-bandwidth bound, batching does not introduce noticeable overhead. We then extend the pipeline with a second proposal depth d2 , using the residual distribution and the gating mechanism to further improve the proposal quality and acceptance (Section 4.3). 4.1

O BSERVATION : E ARLY R ECURRENT S TATES ARE E FFECTIVE D RAFTERS

In Looped Transformers, proposals derived from intermediate recurrent states can tightly approximate the target distribution. In Figure 3, we observe this behavior in the Raven-Llama-3.2 model (McLeish et al., 2026). For instance, on GSM8K (Cobbe et al., 2021), the depth-1 greedy agreement exceeds 86% and surpasses 99% by depth 8, with a similar trend on MATH-500 (Hendrycks et al., 2021; Lewkowycz et al., 2022; Kydlicek et al., 2025; Lightman et al., 2024). We further estimate the empirical total variation (TV) distance between the intermediate readout and the target distributions under the same setting, and observe that it decreases sharply as recurrence depth r increases. To characterize this, we bound the empirical TV distance with a power-law upper envelope ϕ(r) = βr−α . Specifically, we choose parameters α, β > 0 to minimize the maximum ratio between ϕ(r) and the empirical TV distance over all depths r. This power-law upper envelope allows us to analyze the optimal depth configuration studied in Section 5. These observations motivate using early recurrent states to draft future tokens, while continuing recurrence to verify the drafts at depth R. Deeper intermediate states can further serve as fallback proposals when earlier predictions diverge from the target. 4

100 99.13%

92.56%

90 85

99.69%

97.38%

95

Total variation (TV)

Greedy agreement (%)

Preprint.

GSM8K MATH-500

1

8

16

24

Recurrent depth r (a)

32

φ(r) = 0.175r−1.03 0.1 0.02 GSM8K MATH-500

0.005 1

3

10

Recurrent depth r (b)

30

Figure 3: Comparison between intermediate proposals at depth r and the target distribution at depth R = 32 for Raven-Llama-3.2. (a) Greedy top-1 agreement on GSM8K and MATH-500. (b) Sampling (T = 1.0, top-p = 0.7) total variation (TV) distance to the target on the same benchmarks. TV is obtained on 10 depths over GSM8K and MATH-500, averaged per token position on top-p distribution. ϕ(r) is the empirical power-law upper envelope for the TV distance, constructed as described in Section 4.1.

4.2

P IPELINED S ELF -S PECULATIVE D ECODING WITH A S INGLE P ROPOSAL

We first describe the single-proposal pipeline in Figure 1(b), where each token is drafted at only one intermediate depth d1 where d1 < R, and verified at the final depth R. Unlike vanilla decoding, which completes all R recurrences before moving to the next token (Figure 1(a)), this schedule overlaps computation across token positions (Figure 1(b)). We call the recurrent computation for one prefix a branch, represented as a horizontal sequence of recurrent states in the figure. A draft starts a child branch, while its parent continues toward verification. Multiple branches can remain active at different recurrent depths and crucially they are advanced by a single batched call to the shared block. (r)

For a branch b, we denote its prefix by X (b) and use the hidden states HX (b) and readouts pr (· | X (b) ) from Section 3.1. We write pr when the prefix is clear. After prefilling the prompt and committing the first decoded token, L OOP S PEC initializes a branch on the resulting prefix at depth 0 and repeats the following steps: 1. Batched Recurrence. All active branches advance by one recurrent depth through a single batched call to B. In Figure 1(b), each column represents one such step across the active branches. (d )

1 2. Early Drafting. When branch b reaches d1 , it reads out q1 := pd1 = C(HX (b) ) and samples

(1)

(1)

x eb ∼ q1 . A child branch is then initialized at depth 0 with the extended prefix X (b) ∥ x eb . For example, with d1 = 1 in Figure 1(b), the draft “France” at timestep 1 starts a child at timestep 2 before “France” is verified at timestep 4. (1)

3. Verification and Continuation. When the branch reaches R, it verifies x eb against pR using rejection sampling as Equation (3). If accepted, the draft is committed and decoding continues from its child, which has already completed R − d1 recurrent steps. In Figure 1(b), accepting “and” at timestep 8 preserves its child branch that started at timestep 6. If rejected, a replacement xb ∼ pR ⊖ q1 is committed, all speculative descendants of b are pruned, and a new branch starts on X (b) ∥ xb at depth 0. The “France” draft at timestep 1 in Figure 1(b) illustrates this case: depth-R commits “Paris”, which differs from the draft “France”, so we prune the descendants of “France” and restart computation for the next token. With sustained draft acceptance, tokens are committed every d1 recurrent steps rather than every R steps. However, each rejection discards the speculative progress and requires another R recurrent steps before the next commitment. This motivates a more accurate second proposal that can provide a fallback draft when the first draft is rejected. 5

Preprint.

4.3

R ESIDUAL S ECOND P ROPOSAL WITH G ATING M ECHANISM

We now extend the single-proposal pipeline in Section 4.2 by adding a second proposal depth d2 , with d1 < d2 < R. At depth d2 , a branch can draft an alternative token for the same position, starting a fallback child alongside the first child. Accepting this fallback preserves its progress instead of restarting the pipeline. We next describe how to construct the fallback proposal, when to create it, and how to verify the two candidates. Residual Second Proposal. The second proposal is used only after the first proposal is rejected, so the required target distribution is pR ⊖ q1 . We therefore draft from an estimate of this residual (d2 ) target. When branch b reaches d2 , it obtains a deeper readout qe2 := pd2 = C(HX (b) ) and forms (2)

q2 := qe2 ⊖ q1 = pd2 ⊖ pd1 ,

x eb ∼ q2 .

(5)

Here we use qe2 as an estimator for pR and sample the second proposal from the residual distribution (2) q2 . A second child then starts from depth 0 with the prefix X (b) ∥ x eb . The two proposals represent alternative branches for the same token position, as illustrated by the “France” and “Paris” drafts at timesteps 1 and 2 in the top row of Figure 1(c). Our ablation study in Section 6.3 shows that the residual second proposal q2 consistently improves the decoding speedup over the non-residual second proposal qe2 across various benchmarks. Gating Mechanism. Creating a fallback at every branch would rapidly increase the number of active branches, since each child can itself issue two proposals. Specifically, if every active branch forks at both d1 and d2 , the number of new branches sk spawned in the k-th interval of d1 recurrent steps obeys the delayed recurrence sk = sk−1 + sk−d2 /d1 , where sk = 1 for 0 ≤ k < d2 /d1 1 . For PR/d −1 R = 32, d1 = 2, and d2 = 8, this gives maximum branch count B = k=01 sk = 249. To limit this growth of B, L OOP S PEC opens the gate of second proposal and creates a second proposal branch only if (1) (1) qe2 (e xb ) < q1 (e xb ). (6) Intuitively, a fallback branch is created when the deeper recurrent state withdraws confidence from (1) the primary candidate. Under greedy decoding, the rule reduces to arg max qe2 ̸= x eb . In Figure 1(c), the gate stays closed for the “and” branch where no second proposal is triggered at timestep 4. Section 6.3 evaluates the benefit of the gating mechanism. Moreover, Section D examines the gate opening frequency and its effect on the effective batch size. Cascade Verification. When a branch b reaches depth R, it produces the exact target distribution pR and a cascade verification attempts to accept the primary candidate, then the fallback candidate, and restarts the pipeline only if neither candidate is accepted. n o (1) (1) (1) 1. First Proposal. Accept x eb against pR with probability min 1, pR (e xb )/q1 (e xb ) . If accepted, commit it and keep its descendants, and prune the second child’s subtree if present. In Figure 1(c), at timestep 7, accepting the first draft “it” from timestep 4 preserves its continuation while pruning the second draft “the” from timestep 5 and its subtree. (2)

2. Second Proposal. If the first proposal eb against n is rejected and a second o proposal exists, verify x (2)

(2)

ρ1 := pR ⊖ q1 with probability min 1, ρ1 (e xb )/q2 (e xb ) . If accepted, commit it and keep its descendants, and prune the first child’s subtree. This child has already completed R−d2 recurrent steps, so the next commitment requires only d2 more steps instead of R. This is the rejection of “France” and acceptance of the second draft “Paris” in the top 5 rows of Figure 1(c). 3. Residual Resampling. If both proposals are rejected, sample xb ∼ ρ1 ⊖ q2 . If the first is rejected and the gate was closed, sample xb ∼ ρ1 instead. Commit xb , prune the speculative descendants of b, and start a fresh branch on X (b) ∥ xb at depth 0. (1)

(2)

Under greedy decoding, this cascade checks whether x eb matches arg max pR , falling back to x eb if available, and committing arg max pR otherwise. The algorithm is provided in Section H. 1

Throughout the paper, we consider that d2 is divisible by d1 .

6

Preprint.

5

T HEORETICAL A NALYSIS

In this section, we provide theoretical analysis of L OOP S PEC from two perspectives. First, we show that the restart probability depends only on the approximation error of the second proposal and is independent of that of the first proposal. Second, we use the power-law decay of approximation error across recurrent depths, as observed in Figure 3(b), to derive the optimal proposal depths that minimize the expected number of recurrent steps per generated token. The optimal proposal depths derived from our analysis closely match the empirically optimal depth configurations. Theorem 1 (Rejection Rate of L OOP S PEC). Fix proposal depths 1 ≤ d1 < d2 < R and let ε2 be the total variation distance between pd2 and pR . Under the residual second proposal q2 = pd2 ⊖ pd1 and cascade verification, the probability that both candidates are rejected at a given position is at most ε2 , independently of d1 . Moreover, with the gating mechanism, the bound is at most 2ε2 . Proofs of all theorems are provided in Section G. Theorem 1 implies that the residual proposal pd2 ⊖ pd1 achieves the same rejection bound ε2 as a proposal drawn directly from pd2 , while avoiding any additional dependence on d1 . Thus, introducing the residual proposal does not worsen the rejection bound compared with using pd2 directly. With the gating mechanism described in Section 4.3, the bound increases only by a constant factor, from ε2 to 2ε2 . This result allows us to characterize the cost of a rejection using only the approximation quality at the second proposal depth d2 . Hence, we can focus on how the approximation error decreases with recurrent depth to derive the proposal depths that minimize the expected decoding cost. Motivated by the observations in Figure 3(b), we assume a non-increasing power-law envelope ϕ(r) = βr−α that upper bounds the expected total variation distance between the intermediate readout distribution pr and the target distribution pR . Theorem 2 (Proposal Depth Selection). Let X be a random prefix drawn from a fixed distribution over prefixes. Assume that there exist constants α, β > 0 such that EX [TV (pr (· | X), pR (· | X))] ≤ βr−α .

(7) α2 +α+1

1 for every r = 1, ..., R and the total depth R > 2α max{(αβ)1/α , (αβ)−α−1 , (αβ)1/α (2α) α2 }. Then, the optimal proposal depths to minimize expected number of recurrent steps to commit a single token are α+1

1

α

d⋆1 = (αβ) α2 +α+1 (2αR) α2 +α+1 ,

α+1

d⋆2 = (αβ) α2 +α+1 (2αR) α2 +α+1 .

(8)

As shown in Figure 3(b), we obtain an empirical upper bound with β = 0.175 and α = 1.03 for Raven-Llama-3.2 model with R = 32. Substituting these values into Equation (8) gives us (d⋆1 , d⋆2 ) = (1.26, 8.85). In Section 6, we explore various (d1 , d2 ) configurations and find that (d1 , d2 ) = (2, 8) performs best for the same model. This demonstrates that our theoretical analysis closely captures the empirically optimal depth configuration.

6

E XPERIMENTS

Models and Evaluations. We evaluate L OOP S PEC across two representative Looped Transformer families: the Ouro family with R = 4 (Zhu et al., 2025) and the Raven family with R = 32 (McLeish et al., 2026). Evaluations cover mathematical reasoning, general reasoning, and code generation: GSM8K (Cobbe et al., 2021), MATH-500 (Hendrycks et al., 2021; Lewkowycz et al., 2022; Kydlicek et al., 2025; Lightman et al., 2024), BBH (Suzgun et al., 2023), HumanEval+ and MBPP+ (Chen et al., 2021; Austin et al., 2021; Liu et al., 2023) for base models, alongside GSM8K-CoT, MATH-500, AIME 2024, and AIME 2025 (Mathematical Association of America, 2024; 2025) for reasoning models (i.e., with thinking mode enabled). Base models are evaluated under greedy decoding (T = 0.0), whereas reasoning models are evaluated using chat templates with thinking modes enabled, under both greedy decoding (T = 0.0) and sampling (T = 1.0, top-p = 0.7). All evaluations are conducted in a single-batch setting integrated with lm-evaluation-harness (Gao et al., 2023). Detailed checkpoint identifiers and benchmark configurations are provided in Section A. 7

Preprint.

Table 1: L OOP S PEC decoding speedup over the standard autoregressive decoding and mean accepted length γ on general reasoning, math, and code benchmarks across different proposal depths. Bold values mark per-model, per-benchmark maxima, with speedup and γ selected independently. Models

Proposal Depths

GSM8K Speedup

γ

MATH-500 γ

Speedup

BBH Speedup

HumanEval+ γ

MBPP+

Speedup

γ

Speedup

γ

Greedy Setting: Temperature=0.0 Ouro-1.4B

1 1, 2

2.64× 2.85×

3.13 3.50

2.64× 2.78×

3.25 3.55

2.66× 2.79×

3.30 3.59

3.17× 3.16×

3.73 3.86

2.79× 2.92×

3.36 3.64

Ouro-2.6B

1 1, 2

3.03× 3.12×

3.43 3.68

2.93× 2.99×

3.50 3.72

2.89× 2.98×

3.49 3.72

3.36× 3.33×

3.78 3.89

3.00× 3.09×

3.44 3.69

RavenLlama-3.2

2 4 4, 8 2, 10 2, 8

3.04× 7.62 4.77× 11.56 4.19× 10.35 4.93× 11.73 3.96× 9.70 4.06× 6.55 4.76× 7.43 4.57× 7.21 4.89× 7.52 4.55× 7.14 4.33× 7.51 4.90× 7.83 4.76× 7.74 5.01× 7.84 4.72× 7.70 4.38× 12.02 5.70× 14.40 5.32× 13.72 5.83× 14.36 5.07× 13.08 4.49× 12.50 5.82× 14.60 5.39× 13.98 5.88× 14.55 5.14× 13.44

RavenOLMo-2-0425

2 4 4, 8 2, 10 2, 8

3.75× 8.95 3.16× 8.34 3.52× 9.24 5.05× 12.12 5.04× 12.06 4.38× 6.97 4.04× 6.94 4.17× 7.07 4.78× 7.59 4.81× 7.62 4.60× 7.67 4.26× 7.67 4.38× 7.73 4.86× 7.87 4.89× 7.88 5.06× 12.87 4.29× 12.65 4.57× 13.17 5.75× 14.37 5.86× 14.56 5.20× 13.35 4.29× 13.11 4.66× 13.60 5.86× 14.74 5.97× 14.86

2 4 Raven4, 8 TinyLlama-3T 2, 10 2, 8

3.86× 8.18 4.69× 6.72 5.01× 7.58 5.38× 12.43 5.53× 12.89

prompts exceed max context length

5.85× 11.94 4.60× 9.69 5.40× 7.46 4.98× 7.08 5.58× 7.85 5.18× 7.71 6.83× 14.49 5.96× 13.37 6.81× 14.70 5.76× 13.55

Implementation and Metrics. We implement L OOP S PEC by extending the SGLang serving engine (Zheng et al., 2024). All experiments are executed on an NVIDIA RTX PRO 6000 Blackwell GPU in BF16 precision using the Triton attention backend and the TinyGEMM backend provided by FlashInfer (Ye et al., 2025). Further SGLang implementation details are in Section B. We report speedup on wall-clock time and mean accepted length γ. Wall-clock time speedup is measured over the decoding phase, after the completion of prompt prefilling until the end of generation. We define γ as the average number of generated tokens produced per full model execution of R recurrent steps: γ :=

Ndecode · R , Naccept,1 · d1 + Naccept,2 · d2 + (Ndecode − Naccept,1 − Naccept,2 ) · R

(9)

where Ndecode is the number of generated tokens, and Naccept,1 and Naccept,2 denote the number of tokens accepted from the first and second proposals, respectively. By construction, γ ≤ R/d1 , with the upper bound attained when all tokens are accepted at the first proposal depth d1 . 6.1

BASE M ODELS

Table 1 presents the decoding speedup and mean accepted length γ across base model checkpoints. Across various tasks, L OOP S PEC provides consistent speedups over standard autoregressive decoding. For the Ouro family (R = 4), L OOP S PEC achieves speedups up to 3.36×, with γ exceeding 3.5 out of R/d1 = 4. For the deeper Raven family (R = 32), L OOP S PEC yields peak speedups ranging from 5.88× to 6.83× across models, with γ exceeding 14.5. Here, while a single proposal at depth d1 = 2 suffers from lower acceptance rates, adding a second proposal (e.g., d1 = 2, d2 = 8) compensates for early errors, allowing the first proposal depth to be pushed down to d1 = 2 to enlarge the pipeline depth and maximize speedup. In Section E, we further show that the first proposal accounts for over 86% of committed tokens, while the second proposal recovers the majority of the remainder, leaving at most 2.95% of all committed tokens to full pipeline restart. 8

Preprint.

Table 2: L OOP S PEC decoding speedup over the standard autoregressive decoding and mean accepted length γ with thinking enabled across different proposal depths. Results under temperature 0.0 and 1.0 (with top-p = 0.7) are reported separately. Bold values mark per-model, per-benchmark maxima, with speedup and γ selected independently. Proposal Depths

Models

GSM8K Speedup

MATH-500 γ

Speedup

γ

AIME 2024

AIME 2025

Speedup

γ

Speedup

γ

Greedy Setting: Temperature=0.0 Ouro-1.4B-Thinking

1 1, 2

2.47× 2.70×

2.95 3.36

2.49× 2.65×

3.12 3.50

2.56× 2.52×

3.26 3.54

2.45× 2.59×

3.23 3.55

Ouro-2.6B-Thinking

1 1, 2

2.76× 2.95×

3.21 3.55

2.63× 2.78×

3.33 3.64

2.53× 2.58×

3.31 3.63

2.58× 2.66×

3.22 3.64

Sampling Setting: Temperature=1.0 Ouro-1.4B-Thinking

1 1, 2

2.11× 2.42×

2.64 3.21

2.17× 2.41×

2.88 3.35

2.07× 2.34×

2.77 3.31

2.05× 2.27×

2.76 3.28

Ouro-2.6B-Thinking

1 1, 2

2.62× 2.83×

3.11 3.54

2.52× 2.70×

3.21 3.61

2.37× 2.51×

3.14 3.54

2.39× 2.65×

3.10 3.53

Table 3: Ablation study of different second proposal methods in L OOP S PEC. Model: Raven-Llama-3.2

GSM8K

MATH-500

HumanEval+

MBPP+

Standard Autoregressive Decoding + First & Non-residual Second Proposal + Residual Second Proposal + Gating Mechanism

1.00× 3.24× 3.52× 3.74×

1.00× 4.49× 4.73× 4.85×

1.00× 4.59× 4.76× 5.00×

1.00× 4.07× 4.57× 4.71×

Notably, γ measures the reduction in serial recurrent steps but excludes the cost of P and C computation. A smaller d1 triggers more frequent drafting and branch initialization, increasing these costs, particularly in Raven models with heavy pre- and post-layers. This overhead can outweigh the recurrent step savings, explaining why some configurations on Raven models achieve higher γ values but lower wall-clock time speedup. 6.2

R EASONING M ODELS

Table 2 reports the performance of L OOP S PEC in reasoning models with the thinking mode enabled. Across long-context chain-of-thought generation on GSM8K, MATH-500, and AIME 2024/2025 benchmarks, L OOP S PEC delivers speedups up to 2.95× with high γ values even on complex, longreasoning benchmarks. For stochastic sampling (T = 1.0, top-p = 0.7), L OOP S PEC maintains competitive speedups up to 2.83×. Here, the second proposal proves particularly effective: drafting from the residual distribution upon rejection increases γ from 3.11 to 3.54 on Ouro-2.6B-Thinking under the GSM8K benchmark, yielding substantial speedup gains over using only the first proposal. 6.3

A BLATION S TUDY

Residual Second Proposal and Gating Mechanism. Table 3 presents the results of an ablation study on the second proposal in L OOP S PEC, which are the residual second proposal and the gating mechanism. The ablation study is conducted on Raven-Llama-3.2 model under sampling (T = 1.0, top-p = 0.7) on GSM8K, MATH-500, HumanEval+, and MBPP+ benchmarks. We mask out (1) the first proposal token x eb from qe2 for the non-residual second proposal and from q2 for the residual second proposal, and renormalize the corresponding distribution before sampling. This ensures that the second proposal differs from the first proposal. With gating, this masking can be omitted for 9

2.27

2.65

1.57

+71%

1.33

2.51 1.57

1.45

1.66

1.78

1.67

1.91

2.34

2.07 2.41 2.27 2.70

MATH-500 AIME 2024 AIME 2025

(b) Stochastic sampling (T = 1.0, top-p = 0.7)

2.66

GSM8K

(a) Greedy decoding

2.59

MATH-500 AIME 2024 AIME 2025

2

2.58

GSM8K

2.52

2.31 2.42 2.41 2.83

+60% 2.43 2.65 2.45 2.78

3

2.52 2.70 2.58 2.95

Speedup over baseline (×)

Preprint.

1 0

DFlash (Ouro-1.4B-Thinking) LoopSpec (Ouro-1.4B-Thinking)

DFlash (Ouro-2.6B-Thinking) LoopSpec (Ouro-2.6B-Thinking)

Figure 4: Decoding speedup over the standard autoregressive decoding for L OOP S PEC versus DFlash on the Ouro thinking models, under greedy decoding and stochastic sampling. (1)

(1)

the residual second proposal because the gate opens only when qe2 (e xb ) < q1 (e xb ), in which case (1) q2 (e xb ) is already zero. Adding the first and non-residual second proposal at d1 = 2, d2 = 8 yields speedups ranging from 3.24× to 4.59×. Replacing the non-residual second proposal with the residual second proposal and adding the gating mechanism consistently improve the speedups to 3.74× to 5.00×, showing the effectiveness of the residual second proposal and gating mechanism. Section D further examines how often the gate opens and how gating reduces the effective batch size during decoding. Comparison between L OOP S PEC and DFlash. We additionally compare our training-free L OOP S PEC to DFlash (Chen et al., 2026) on various reasoning tasks, where DFlash is an effective training-based speculative decoding method that has been largely adopted in standard autoregressive models, e.g., Xiaomi MiMo Team (2026). As shown in Figure 4, L OOP S PEC consistently outperforms DFlash by showing up to 1.7× speedup on decoding time under both greedy and stochastic sampling settings. More details on implementation and settings are provided in Section C.

7

C ONCLUSION

We introduced L OOP S PEC, a training-free self-speculative decoding framework for Looped Transformers that converts recurrent depth into token-position parallelism. An early recurrent state drafts the next token and immediately begins computing it while the current token continues toward finaldepth verification. Furthermore, we proposed a residual second proposal with a gating mechanism from a deeper recurrent depth alongside a cascade verification that keeps decoding lossless. Across the Ouro (R = 4) and Raven (R = 32) families, L OOP S PEC achieves speedups of up to 6.83× on general reasoning, math, and coding benchmarks.

R EFERENCES Jacob Austin, Augustus Odena, Maxwell I. Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J. Cai, Michael Terry, Quoc V. Le, and Charles Sutton. Program synthesis with large language models. CoRR, abs/2108.07732, 2021. URL https: //arxiv.org/abs/2108.07732. 6 Seongjin Cha, Gyuwan Kim, Dongsu Han, Tao Yang, and Insu Han. KnapSpec: Self-speculative decoding via adaptive layer selection as a knapsack problem. In Forty-third International Conference on Machine Learning, 2026. URL https://openreview.net/forum?id= k5nKHWp9VC. 1 Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre, and John Jumper. Accelerating large language model decoding with speculative sampling, 2023. URL https://arxiv.org/abs/2302.01318. 1, 2, 3.2 10

Preprint.

Jian Chen, Yesheng Liang, and Zhijian Liu. DFlash: Block diffusion for flash speculative decoding. In Forty-third International Conference on Machine Learning, 2026. URL https: //openreview.net/forum?id=Oz335dV48X. 1, 1, 2, 6.3 Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Pondé de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Joshua Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. Evaluating large language models trained on code. CoRR, abs/2107.03374, 2021. URL https://arxiv.org/abs/2107.03374. 6 Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems. CoRR, abs/2110.14168, 2021. URL https://arxiv.org/abs/2110.14168. 4.1, 6 Chunyuan Deng, Yizhe Zhang, Rui-Jie Zhu, Yuanyuan Xu, Jiarui Liu, T. S. Eugene Ng, and Hanjie Chen. LT2: Linear-time looped transformers, 2026. URL https://arxiv.org/abs/ 2605.20670. 2 Mostafa Elhoushi, Akshat Shrivastava, Diana Liskovich, Basil Hosmer, Bram Wasti, Liangzhen Lai, Anas Mahmoud, Bilge Acun, Saurabh Agarwal, Ahmed Roman, Ahmed Aly, Beidi Chen, and Carole-Jean Wu. LayerSkip: Enabling early exit inference and self-speculative decoding. In LunWei Ku, Andre Martins, and Vivek Srikumar (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 12622–12642, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/ 2024.acl-long.681. URL https://aclanthology.org/2024.acl-long.681/. 1, 2 Tianyu Fu, Yichen You, Zekai Chen, Guohao Dai, Huazhong Yang, and Yu Wang. Think-at-Hard: Selective latent iterations to improve reasoning language models. In Forty-third International Conference on Machine Learning, 2026. URL https://openreview.net/forum?id= eQaJSRZiGn. 2 Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. A framework for few-shot language model evaluation, 12 2023. URL https://zenodo.org/records/ 10256836. 6 Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer, Siddharth Singh, Brian Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, and Tom Goldstein. Scaling up testtime compute with latent reasoning: A recurrent depth approach. In D. Belgrave, C. Zhang, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, and N. Chen (eds.), Advances in Neural Information Processing Systems, volume 38, Main Conference, pp. 41340–41391. Curran Associates, Inc., 2025. doi: 10.52202/085713-1380. URL https://proceedings.neurips.cc/paper_files/paper/2025/file/ 3b01972cf31e6fa0fe29e4b8b5c2a0a1-Paper-Conference.pdf. 1, 2 Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the MATH dataset. In J. Vanschoren and S. Yeung (eds.), Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, volume 1, 2021. URL https: //datasets-benchmarks-proceedings.neurips.cc/paper_files/paper/ 2021/file/be83ab3ecd0db773eb2dc1b0a17836a1-Paper-round2.pdf. 4.1, 6 11

Preprint.

Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Hasan Genc, Kurt Keutzer, Amir Gholami, and Yakun Sophia Shao. SPEED: Speculative pipelined execution for efficient decoding. In Third Workshop on Efficient Natural Language and Speech Processing (ENLSP-III): Towards the Future of Large Language Models and Their Emerging Descendants, New Orleans, Louisiana, USA, 2023. URL https://neurips2023-enlsp.github.io/papers/paper_17.pdf. 1, 2 Hynek Kydlicek, Alina Lozovskaya, Nathan Habib, and Clémentine Fourrier. Fixing open llm leaderboard with math-verify, 2025. 4.1, 6 Yaniv Leviathan, Matan Kalman, and Yossi Matias. Fast inference from transformers via speculative decoding. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett (eds.), Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pp. 19274–19286. PMLR, 23–29 Jul 2023. URL https://proceedings.mlr.press/ v202/leviathan23a.html. 1, 2, 3.2, 3.2 Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra. Solving quantitative reasoning problems with language models. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (eds.), Advances in Neural Information Processing Systems, volume 35, pp. 3843–3857. Curran Associates, Inc., 2022. doi: 10.52202/068431-0278. URL https://proceedings.neurips.cc/paper_files/paper/2022/file/ 18abbeef8cfe9203fdf9053c9c4fe191-Paper-Conference.pdf. 4.1, 6 Shenggui Li, Chao Wang, YIKAI ZHU, Yubo Wang, Fan Yin, Shuai Shi, Yefei Chen, Xiaomin Dong, Qiaoling Chen, Jin Pan, Ji Li, Yineng Zhang, Lei Yu, Yonggang Wen, Ivor Tsang, and Tianwei Zhang. SpecForge: A flexible and efficient open-source training framework for speculative decoding. In Forty-third International Conference on Machine Learning, 2026. URL https://openreview.net/forum?id=CQOEbxy0tE. C Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang. EAGLE-3: Scaling up inference acceleration of large language models via training-time test. In D. Belgrave, C. Zhang, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, and N. Chen (eds.), Advances in Neural Information Processing Systems, volume 38, Main Conference, pp. 136737–136756. Curran Associates, Inc., 2025. doi: 10.52202/085713-4562. URL https://proceedings.neurips.cc/paper_files/paper/2025/file/ c7b5a35ea98b62512a869c19ea7b03cb-Paper-Conference.pdf. 1, 2 Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let’s verify step by step. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview. net/forum?id=v8L0pN6EOi. 4.1, 6 Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and LINGMING ZHANG. Is your code generated by chatgpt really correct? Rigorous evaluation of large language models for code generation. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (eds.), Advances in Neural Information Processing Systems, volume 36, pp. 21558–21572. Curran Associates, Inc., 2023. doi: 10.52202/075280-0943. URL https://proceedings.neurips.cc/paper_files/paper/2023/file/ 43e9d647ccd3e4b7b5baab53f0368686-Paper-Conference.pdf. 6 Anton Lozhkov, Hynek Kydlı́ček, Loubna Ben Allal, Guilherme Penedo, Edward Beeching, Quentin Gallouédec, Nathan Habib, Lewis Tunstall, and Leandro von Werra. OpenR1-Math-220k. https://huggingface.co/datasets/open-r1/OpenR1-Math-220k, 2025. C Mathematical Association of America. American invitational mathematics examination (AIME) 2024. https://maa.org/maa-invitational-competitions/, February 2024. 6 Mathematical Association of America. American invitational mathematics examination (AIME) 2025. https://maa.org/maa-invitational-competitions/, February 2025. 6 12

Preprint.

Sean Michael McLeish, Ang Li, John Kirchenbauer, Dayal Singh Kalra, Brian R. Bartoldson, Bhavya Kailkhura, Avi Schwarzschild, Jonas Geiping, Tom Goldstein, and Micah Goldblum. Teaching pretrained language models to think deeper with retrofitted recurrence. In Third Conference on Language Modeling, 2026. URL https://openreview.net/forum?id= PXVQTHYwgt. 1, 2, 2, 3.1, 4.1, 6 Nanbeige Lab, Chen Yang, Chengrui Huang, Fufeng Lan, Hanhui Chen, Hao Zhou, Huatong Song, Jiaqi Cao, Jiaying Zhu, Jinlin Niu, Kai Wang, Lisheng Huang, Qiliang Liang, Ran Le, Ruixiang Feng, Shuang Sun, Tao Gu, Tao Zhang, Tianyu Luo, Yang Song, Yun Xing, Yuntao Wen, Ziyao Xu, Zongchao Chen, and Zongqiang Li. Nanbeige4.2-3B: Unlocking agentic capabilities in a compact model, 2026. URL https://arxiv.org/abs/2607.22083. 1, 2 Taekhyun Park, Yongjae Lee, Dohee Kim, and Hyerim Bae. LoopUS: Recasting pretrained llms into looped latent refinement models, 2026. URL https://arxiv.org/abs/2605.11011. 2 Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc Le, Ed Chi, Denny Zhou, and Jason Wei. Challenging BIG-bench tasks and whether chain-of-thought can solve them. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (eds.), Findings of the Association for Computational Linguistics: ACL 2023, pp. 13003–13051, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.findings-acl.824. URL https://aclanthology.org/2023. findings-acl.824/. 6 Xiaomi MiMo Team. Mimo-v2.5-pro-fp4-dflash. https://huggingface.co/ XiaomiMiMo/MiMo-V2.5-Pro-FP4-DFlash, 2026. 6.3 Zihao Ye, Lequn Chen, Ruihang Lai, Wuwei Lin, Yineng Zhang, Stephanie Wang, Tianqi Chen, Baris Kasikci, Vinod Grover, Arvind Krishnamurthy, and Luis Ceze. FlashInfer: Efficient and customizable attention engine for LLM inference serving. In Eighth Conference on Machine Learning and Systems, 2025. URL https://openreview.net/forum?id= RXPofAsL8F. 6 Jun Zhang, Jue Wang, Huan Li, Lidan Shou, Ke Chen, Gang Chen, and Sharad Mehrotra. Draft & Verify: Lossless large language model acceleration via self-speculative decoding. In LunWei Ku, Andre Martins, and Vivek Srikumar (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 11263–11282, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/ 2024.acl-long.607. URL https://aclanthology.org/2024.acl-long.607/. 1, 2 Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng. SGLang: Efficient execution of structured language model programs. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang (eds.), Advances in Neural Information Processing Systems, volume 37, pp. 62557–62583. Curran Associates, Inc., 2024. doi: 10.52202/079017-2000. URL https://proceedings.neurips.cc/paper_files/paper/2024/file/ 724be4472168f31ba1c9ac630f15dec8-Paper-Conference.pdf. 1, 6, B Rui-Jie Zhu, Zixuan Wang, Kai Hua, Tianyu Zhang, Ziniu Li, Haoran Que, Boyi Wei, Zixin Wen, Fan Yin, He Xing, Lu Li, Jiajun Shi, Kaijing Ma, Shanda Li, Taylor Kergan, Andrew Smith, Xingwei Qu, Mude Hui, Bohong Wu, Qiyang Min, Hongzhi Huang, Xun Zhou, Wei Ye, Jiaheng Liu, Jian Yang, Yunfeng Shi, Chenghua Lin, Enduo Zhao, Tianle Cai, Ge Zhang, Wenhao Huang, Yoshua Bengio, and Jason Eshraghian. Scaling latent reasoning via looped language models, 2025. URL https://arxiv.org/abs/2510.25741. 1, 2, 2, 3.1, 6

13

Preprint.

A

E XPERIMENTAL S ETUP D ETAILS

Table 4 presents the complete mapping between evaluated Looped Transformer checkpoints and the model names used throughout this paper. Table 5 outlines the benchmark specifications, prompting protocols, and maximum generation token budgets across base and reasoning models. For reasoning models, evaluations apply chat templates with thinking mode enabled to support long chain-ofthought derivations. For code generation on HumanEval+ and MBPP+, we employ zero-shot custom prompts that instruct models to produce self-contained Python scripts within markdown code blocks. Table 4: Mapping between evaluated checkpoints and the model names used in this paper. Name in Paper

Hugging Face Checkpoint

Base Models Ouro-1.4B Ouro-2.6B Raven-Llama-3.2 Raven-OLMo-2-0425 Raven-TinyLlama-3T

ByteDance/Ouro-1.4B ByteDance/Ouro-2.6B smcleish/Recurrent-Llama-3.2-train-recurrence-32 smcleish/Recurrent-OLMo-2-0425-train-recurrence-32 smcleish/Recurrent-TinyLlama-3T-train-recurrence-32

Reasoning Models Ouro-1.4B-Thinking Ouro-2.6B-Thinking

ByteDance/Ouro-1.4B-Thinking ByteDance/Ouro-2.6B-Thinking

Table 5: Evaluation benchmark configurations for base and reasoning models.

B

Benchmark

Prompting Protocol

Base Max Gen Toks

Reasoning Max Gen Toks

GSM8K MATH-500 BBH HumanEval+ MBPP+ AIME 2024 AIME 2025

3-shot CoT 4-shot 3-shot CoT 0-shot (Custom Prompt) 0-shot (Custom Prompt) 0-shot 0-shot

512 2,048 1,024 1,024 1,024 – –

4,096 8,192 – – – 16,384 16,384

I MPLEMENTATION D ETAILS OF P IPELINED D ECODING

We implement L OOP S PEC on top of SGLang (Zheng et al., 2024). In this section, we cover key implementation details as below. KV Cache Management. The biggest challenge during L OOP S PEC decoding is correctly maintaining the KV states. Computation at different recurrent depths requires different KV states, even though they share model parameters. Moreover, pruning a rejected branch must preserve the prefix KV states still needed by other branches. We maintain the KV cache for each branch separately at each recurrent depth. A newly created branch reuses its parent’s cached prefix by copying the KV indices rather than the KV states and stores new KV separately. After verification, we keep the selected branch and its descendants and release only KV states no longer needed by the remaining branches. For Raven, we also maintain KV caches for P and C, with separate caches for C at each proposal and verification depth. CUDA Graph Execution. Each decoding step involves several GPU kernels, and launching them individually adds overhead to the computation. CUDA Graphs reduce this overhead by recording a sequence of operations and replaying it with updated inputs. However, drafting and pruning change the batch size, while each recorded graph requires fixed input shapes. Since the maximum number of branches can be determined in advance, we prepare graphs for the supported batch sizes separately for P, B, and C. At each step, we select the graph matching the number of branches processed by that component. Each branch’s hidden state is stored between steps and loaded into the selected 14

Preprint.

batch, so changing the batch size does not discard its progress. The graphs also allocate space for new KV entries and update the KV indices before model computation, avoiding separate launches for these cache-management operations. Depth Configuration. We require d2 and R to be multiples of d1 , with d1 < d2 < R. This aligns the proposal and verification stages with the execution schedule defined by d1 and synchronizes the pre-layers P and post-layers C across branches, allowing their computations to be batched.

C

DF LASH I MPLEMENTATION D ETAILS

In Section 6.3, we conducted an ablation study comparing L OOP S PEC and DFlash. The training and inference details are as follows. We trained 5-layer DFlash drafters with block size 16 on the Ouro thinking models. We used the SpecForge (Li et al., 2026) training framework, and trained on the OpenR1-Math-220k dataset (Lozhkov et al., 2025) with regenerated answers, following the recommended recipe. We trained for 10,000 steps with a batch size of 32 on both Ouro-1.4B-Thinking and Ouro-2.6B-Thinking. During inference, we use block size of 16 for DFlash, and identical hyperparameters and SGLang configurations for L OOP S PEC and DFlash for fair comparison.

D

E MPIRICAL R ESULTS ON G ATING S TRATEGY AND E FFECTIVE BATCH S IZE

Gating reduces the larger parallel computation introduced by second proposals. Each proposal spawns a child branch that continues drafting future tokens, so unconditionally issuing second proposals increases the number of active branches. We examine how often the gate opens and how gating affects the effective batch size on the two Raven models under sampling. Gate Opening Frequency. We measure how often the gate opens at d2 to trigger a second proposal across four benchmarks. As shown in Table 6, the opening frequency ranges from 18.75% to 27.10% for Raven-Llama-3.2 and from 20.11% to 40.40% for Raven-OLMo-2-0425. These results show that gating effectively reduces additional branch creation by selectively triggering second proposals. Table 6: Gate opening frequency under sampling, measured as the percentage of gate evaluations that trigger a second proposal. Model

GSM8K

MATH-500

HumanEval+

MBPP+

Raven-Llama-3.2 Raven-OLMo-2-0425

27.10% 24.19%

18.75% 40.40%

19.66% 20.11%

20.40% 20.49%

Effective Batch Size. We define effective batch size as the number of active branches processed by each recurrent forward pass and report its distribution over all recurrent forward passes during evaluation. Figure 5 compares decoding with and without gating on GSM8K. Gating reduces the mean effective batch size from 43.75 to 34.06 for Raven-Llama-3.2 and from 40.74 to 31.70 for Raven-OLMo-2-0425. The distributions also shift toward smaller batch sizes, with observed peaks decreasing from 239 to 176 and from 243 to 191, respectively. These results show that gating reduces both the mean and peak effective batch size during inference.

15

Preprint.

(a) Raven-Llama-3.2

Frequency (%)

40

With gate (mean: 34.06, peak: 176) Without gate (mean: 43.75, peak: 239)

30 20 10 0

0

10 20 30 40 50 60 70 80 90 100 110 120 130 140 150 160 170 180 190 200 210 220 230 240 250

Effective batch size

(b) Raven-OLMo-2-0425

Frequency (%)

40

With gate (mean: 31.70, peak: 191) Without gate (mean: 40.74, peak: 243)

30 20 10 0

0

10 20 30 40 50 60 70 80 90 100 110 120 130 140 150 160 170 180 190 200 210 220 230 240 250

Effective batch size

Figure 5: Effective batch size distributions with and without gating on GSM8K under sampling. Effective batch size is the number of active branches processed by each recurrent forward pass. Bars show the percentage of recurrent forward passes in each batch-size bin of width 10. Legends report the mean and peak batch sizes.

E

T OKEN C OMMITMENT S OURCES ACROSS P ROPOSAL D EPTHS

We examine how the two proposal depths contribute to token commitment in L OOP S PEC. We partition committed tokens into three categories: accepted first proposals, accepted second proposals after first-proposal rejection, and full-depth fallback when neither proposal is accepted. These categories identify the source of each committed token. Figure 6 shows the share of tokens committed from each source for Raven-Llama-3.2 with (d1 , d2 , R) = (2, 8, 32) and Ouro-1.4B with (d1 , d2 , R) = (1, 2, 4) in greedy decoding scenario for multiple benchmarks. First Proposals Cover Most Tokens. First proposals account for 91.59%–96.31% of committed tokens for Raven-Llama-3.2 and 86.57%–95.35% for Ouro-1.4B across the four benchmarks. Thus, the shallow proposal supplies most committed tokens in the evaluated settings, allowing the pipeline to retain computation started at the first proposal depth. Second Proposals Recover Early Rejections. Second proposals contribute a further 3.50%– 7.93% of tokens for Raven-Llama-3.2, leaving only 0.19%–0.48% to full-depth fallback. For Ouro-1.4B, the corresponding shares are 3.74%–10.48% and 0.90%–2.95%. Across both models and all four benchmarks, second proposals therefore recover more tokens than require full-depth fallback. These results support the complementary roles of the two depths: the first proposal enables early drafting, while the second preserves speculative progress for a substantial fraction of first-proposal rejections.

16

Share of committed tokens (%)

Preprint.

100 80 60 40 20 0

GSM8K

MATH-500 HumanEval+

MBPP+

GSM8K

(a) Raven-Llama-3.2 𝑑1 = 2, 𝑑2 = 8, 𝑅 = 32 First proposal (d1 )

MATH-500 HumanEval+

MBPP+

(b) Ouro-1.4B 𝑑1 = 1, 𝑑2 = 2, 𝑅 = 4 Second proposal (d2 )

Full-depth (R)

Figure 6: Share of committed tokens by different proposal sources.

F

E MPIRICAL P OWER -L AW D ECAY AND O PTIMAL P ROPOSAL D EPTHS FOR R AVEN -OLM O -2-0425

Total variation (TV)

In this section, we present the Raven-OLMo-2-0425 sampling (T = 1.0, top-p = 0.7) total variation (TV) distance to the target on GSM8K and MATH-500 in Figure 7. Similar to Figure 3(b), we obtain the power-law envelope ϕ(r) = βr−α , and it gives α = 1.08 and β = 0.207. With Theorem 2, we can compute the optimal proposal depths (d∗1 , d∗2 ) = (1.41, 9.17), which again aligns with the empirical optimal depths (d1 , d2 ) = (2, 8) in Section 6. φ(r) = 0.207r−1.08 0.1 0.02 GSM8K MATH-500

0.005 1

3

10

Recurrent depth r

30

Figure 7: Sampling (T = 1.0, top-p = 0.7) total variation (TV) distance between intermediate proposals at depth r and the target distribution at depth R = 32 for Raven-OLMo-2-0425 on GSM8K and MATH-500. TV is obtained on 10 depths and averaged per token position on the top-p distribution. The power-law envelope ϕ(r) is constructed as described in Section 4.1.

17

Preprint.

G

P ROOFS

We provide a formal definition of the total variation distance, which is used to evaluate the performance of speculative decoding. Definition 3 (Total variation distance). For distributions p and q over a finite sample space X , the total variation distance between p and q is defined as X X 1X TV (p, q) := |p(x) − q(x)| = [p − q]+ (x) = [q − p]+ (x). 2 x∈X

x∈X

x∈X

In the standard speculative decoding setting on vocabulary X , given the target distribution p and a proposal distribution q, the accepted probability is   X p(x) Pr[accept] = q(x) min 1, q(x) x∈X X = min (p(x), q(x)) x∈X

=

X

p(x) − [p(x) − q(x)]+ = 1 − TV (p, q) .

x∈X

In other words, TV (p, q) measures a probability of rejection. G.1

P ROOF OF T HEOREM 1

To prove Theorem 1, we need the following lemma in advance. Lemma 4. Let ν1 := pd2 ⊖pd1 , ν2 := pR ⊖pd1 and define ε1 := TV (pd1 , pR ), ε2 := TV (pd2 , pR ). Suppose ε1 > 0 and pd1 ̸= pd2 . Then TV (ν1 , ν2 ) ≤ εε12 . Proof of Lemma 4. We write u := [pR − pd1 ]+ , v := [pd2 − pd1 ]+ and δ := TV (pd1 , pd2 ). By the definition of the residual distribution in Equation (2), ν2 = εu1 , ν1 = vδ . Consider two cases depending on the relative sizes of δ and ε1 . Case 1: δ ≤ ε1 . ε1 TV (ν1 , ν2 ) = ε1

X

max(0, ν2 (x) − ν1 (x))

x

  ε1 max 0, u(x) − v(x) δ x X ≤ max(0, u(x) − v(x))

=

X

≤

X

x

max(0, pR (x) − pd2 (x)) = ε2

x

where the last inequality holds from the fact that [a]+ − [b]+ ≤ [a − b]+ . Similarly, Case 2: δ > ε1 . ε1 TV (ν1 , ν2 ) = ε1

X

max(0, ν1 (x) − ν2 (x))

x

  ε 1 max 0, v(x) − u(x) δ x X ≤ max(0, v(x) − u(x)) =

X

≤

X

x

max(0, pd2 (x) − pR (x)) = ε2

x

By combining both cases, TV (ν1 , ν2 ) ≤ εε12 . This completes the proof of Lemma 4. 18

Preprint.

We are now ready to prove Theorem 1. Proof of Theorem 1. Recall that the rejection probability from the first proposal is ε1 TV (pd1 , pR ) and that from the second one is TV (pR ⊖ pd1 , pd2 ⊖ pd1 ).

=

If ε1 = 0, the first proposal is always accepted, so the probability of both candidates are rejected is bounded by Pr[reject1 ∧ reject2 ] ≤ Pr[reject1 ] = ε1 = 0 ≤ ε2 . If pd1 = pd2 , then ε1 = ε2 , so the probability of both candidates are rejected is bounded by Pr[reject1 ∧ reject2 ] ≤ Pr[reject1 ] = ε1 . Without loss of generality, we assume that ε1 , δ > 0. When the gate is open, Lemma 4 gives ε2 Pr[reject1 ∧ reject2 ∧ gate open] ≤ ε1 · = ε2 . ε1 The gate in Equation (6) is closed when the primary candidate x e(1) ∼ pd1 satisfies pd2 (e x(1) ) ≥ (1) pd1 (e x ), and there is no second proposal. Under this condition, by Equation (3), the probability of drawing x and rejecting it is pd1 (x) − min (pd1 (x), pR (x)) = [pd1 − pR ]+ (x), and hence X Pr[reject1 ∧ gate closed] = [pd1 − pR ]+ (x) x : pd2 (x)≥pd1 (x)

X

≤

[pd2 − pR ]+ (x)

x : pd2 (x)≥pd1 (x)

≤

X

[pd2 − pR ]+ (x) = ε2 ,

x

where the first inequality holds from pd1 (x) ≤ pd2 (x) under the gate-closed condition. Adding the two bounds gives 2ε2 . This completes the proof of Theorem 1. G.2

P ROOF OF T HEOREM 2

Let C be a random variable of the number of recurrent steps to commit a single token. Then,  d1 if accepted in the first proposal, C = d2 else if accepted in the second proposal,  R otherwise. The expectation of C can be computed as E[C] = Pr[accept1 ]d1 + Pr[accept2 ]d2 + (1 − Pr[accept1 ] − Pr[accept2 ])R.

(10)

By the property of rejection sampling and Theorem 1, Pr[accept1 ] = 1 − TV (pd1 , pR ) = 1 − ε1 ≤ 1, Pr[accept2 ] ≤ Pr[reject1 ] = ε1 , 1 − Pr[accept1 ] − Pr[accept2 ] ≤ 2ε2 . For a prefix X, the assumption in Equation (7) yields EX [ε1 ] = EX [TV (pd1 , pR )] ≤ ϕ(d1 ) = βd−α 1 , EX [ε2 ] = EX [TV (pd2 , pR )] ≤ ϕ(d2 ) = βd−α 2 . Substituting these bounds into Equation (10) and taking the expectation over prefix X, we obtain −α EX [C] ≤ d1 + βd−α 1 d2 + 2βd2 R.

(11)

Our goal is to find optimal d1 and d2 that minimize the right-hand side in Equation (11). Denote the right-hand side by f (d1 , d2 ). The stationary conditions are obtained by taking derivative with respect to d1 and d2 , respectively, ∂f = 1 − αβd−α−1 d2 = 0, 1 ∂d1 ∂f −α−1 = βd−α = 0. 1 − 2αβRd2 ∂d2 19

Preprint.

Re-writing these conditions yields α+1

1

α

α+1

d∗1 = (αβ) α2 +α+1 (2αR) α2 +α+1 , d∗2 = (αβ) α2 +α+1 (2αR) α2 +α+1 . To verify above (d∗1 , d∗2 ) is the minimizer, we compute the Hessian of f as   α(α + 1)βd−α−2 d2 −αβd−α−1 1 1 H(d1 , d2 ) = −αβd−α−1 2α(α + 1)βRd−α−2 1 2 and show that the determinant of H(d∗1 , d∗2 ) is strictly positive. Using the fact that R(d∗2 )−α−1 = (d∗1 )−α /(2α), we have  det (H(d∗1 , d∗2 )) = αβ 2 (d∗1 )−2α−2 α2 + α + 1 > 0. Since this is the unique stationary point, (d∗1 , d∗2 ) is the global minimizer. Finally, the condition α2 +α+1 1 max{(αβ)1/α , (αβ)−α−1 , (αβ)1/α (2α) α2 } 2α ensures the validity of solution, i.e., 1 ≤ d∗1 < d∗2 < R. This completes the proof of Theorem 2.

R>

20

Preprint.

H

L OOP S PEC A LGORITHM P SEUDOCODE

Algorithm 1: L OOP S PEC with two proposals Input : prompt x1:n ; pre-layers P; recurrent-layers B; post-layers C; total recurrent depth R; proposal depths d1 < d2 < R Global: active branch set A ← ∅; output string Y ← ε (b) 1 A branch b has its prefix X , current hidden state Hb , recurrent depth rb , and proposal records Qb ; 2 Function Spawn(X): 3 c ← NewBranch(); 4 5 6 7

X (c) ← X; Hc ← P(X (c) ); rc ← 0; Qc ← ∅; A ← A ∪ {c}; return c;

8 Function Verify(b, p): 9 ρ ← p; 10 for j ← 1 to |Qb | do 11 12 13 14 15 16 17

(j)

(e xb , qj , cj ) ← Qb [j]; Uj ∼ U(0, 1); (j) (j) if Uj < min{1, ρ(e xb )/qj (e xb )} then (j) return (e xb , cj ); ρ ← ρ ⊖ qj ;

// accept // reject

x ∼ ρ; return (x, Spawn(X (b) ∥ x))

18 H ← P(x1:n ); 19 for r ← 1 to R do 20 H ← B(H); 21 x ∼ C(H); 22 Y ← Y ∥ x; 23 if generation has not stopped then 24 binit ← Spawn(x1:n ∥ x);

// commit first token after prefill

25 while generation has not stopped do 26 Acur ← A; 27 foreach b ∈ Acur in parallel do 28 Hb ← B(Hb ); rb ← rb + 1; 29 30 31 32 33 34 35 36 37 38 39 40

if rb = d1 ; then (1) q1 ← C(Hb ); x eb ∼ q1 ; (1) (b) c1 ← Spawn(X ∥ x eb ); (1) Qb [1] ← (e xb , q1 , c1 ); else if rb = d2 then (1) (e xb , q1 , c1 ) ← Qb [1]; qe2 ← C(Hb ); (1) (1) if qe2 (e xb ) < q1 (e xb ) ; then (2) q2 ← qe2 ⊖ q1 ; x eb ∼ q2 ; (2) c2 ← Spawn(X (b) ∥ x eb ); (2) Qb [2] ← (e xb , q2 , c2 );

// first proposal

// gating mechanism // residual second proposal

41 bmax ← arg maxb∈A rb ; 42 if rbmax = R then 43 p ← C(Hbmax ); 44 (x⋆ , c⋆ ) ← Verify(bmax , p); 45 Y ← Y ∥ x⋆ ; 46 A ← Reroot(A, c⋆ ); 47 output Y ;

// commit a new decoded token // keep c⋆ and its descendants

21

Record · ID 919396 · SHA-256 969e1932eed917dd
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.