Rare Event Estimation via Iterative Unalignment Hanming Yang Daksh Mittal
Jing Dong
Hongseok Namkoong
Decision, Risk, and Operations Division, Columbia Business School [email protected], {dmittal27, jing.dong, namkoong}@gsb.columbia.edu
arXiv:2609.24969v1 [cs.LG] 21 Sep 2026
Abstract As agents are deployed with increased autonomy, even extremely rare events along their stochastic output trajectories can occur and prove catastrophic. Safe deployment therefore does not depend on whether these events can occur, but on how often they might. We study the problem of estimating the probability of rare events that arise from stochastic variation in the agent’s own actions. Estimating this type of risk requires searching over the combinatorially vast space of trajectories. Naive Monte Carlo is computationally prohibitive in this regime, and constructing effective importance sampling (IS) proposals requires coordinated changes to a context-dependent chain of conditional distributions. We develop a new IS method that perturbs the original model’s weights to construct the proposal. The proposal is itself a differentiably parameterized language model, enabling gradient-based search over weight space. We formulate an objective that combines a differentiable surrogate for event amplification and an adaptive regularization scheme that dynamically balances amplification against estimator stability. We evaluate our approach on ∼120M and ∼2.6B models across three event families spanning 300+ rare events as rare as 10−9 , with reference probabilities computed with < 10% relative standard error. In our most verifiable settings, we observe that our IS estimator achieves over 800× compute-weighted efficiency gains over naive Monte Carlo for events with probabilities lower than 10−7 . Our implementation is available at https://github.com/namkoong-lab/iterative-unalignment.
1
Introduction
Agentic systems are being deployed at scale with increased autonomy; these agents continuously execute code [37, 38, 20, 45], manage file systems [44, 43], and interact with third-party APIs [25, 30]. As agents act over longer horizons without sustained human oversight, rare tail events can dominate safety concerns. A single rare trajectory that deletes a production database, exfiltrates sensitive data, executes an irreversible financial transaction, or takes a power-seeking action can cause catastrophic, irreversible harm, however unlikely it is on any single rollout [35, 26]. At deployment scale, potentially with more than a billion users continuously running agents simultaneously, even a one-in-a-billion failure probability can produce many failures across these runs. Estimating these rare failure probabilities is thus a prerequisite for safe deployment. Only with quantitative estimates can we decide whether the residual risk is tolerable, target pre-deployment remediation where it is most needed, and allocate resources for manual oversight. A substantial body of work has studied adversarial input search, in which user prompts (jailbreaks) are constructed to elicit unsafe behavior [47, 11, 40, 31]. In this work, we focus on a complementary source of risk: stochastic variation in the model’s own decoding trajectory. For a fixed prompt, which may be benign or adversarial, we estimate the probability of failure across the model’s possible continuations. Rare events of this kind arise from the autoregressive generation process itself and can accumulate into systemic agent failures. We therefore study rare-event estimation in the output space of language models (LMs), drawing on the classical literature on Monte Carlo and importance sampling [8, 27, 13, 28, 3, 34]. 1
Figure 1. The challenge is to amplify event trajectories broadly and by similar factors. Black and red numbers show token probabilities under the base and proposal models, respectively. Their products give the sequence probabilities shown on the right. Red shading marks harmful outputs. We seek a proposal that assigns higher probability to every event trajectory, with similar amplification factors so that importance weights remain stable. This requires coordinating probability changes across the full sequence and across the many ways the event can occur.
When failures are rare, even observing a single instance requires an impractically large number of Monte Carlo runs. Importance sampling (IS) can alleviate this burden by sampling from a proposal distribution that assigns greater probability mass to the failure region and reweighting the samples to preserve unbiasedness. Effective estimation involves finding where the original model places probability mass within the event region. This search may also reveal concrete ways the model can reach the specified failure. The statistical efficiency of this scheme is governed by the likelihood ratio between the proposal and original distributions. If the proposal assigns insufficient probability mass to rare-event trajectories that carry non-negligible probability mass under the original distribution, an issue we refer to as under-coverage, the corresponding importance weights can become extremely large, leading to high variance or even estimator collapse despite an increased rare-event frequency. This highlights a central challenge in the design of proposal distributions. How do we search over proposals with amenable likelihood ratio distributions? In classical applications of rare-event estimation such as finance and manufacturing, proposal distributions are often constructed using analytical insight or problem-specific structure [15, 18]. Canonical rare-event formulations often define the event by thresholding a scalar performance score g(y), that is E = {g(y) > γ}. Sequential methods such as Cross Entropy exploit this score directly, using intermediate thresholds to guide the proposal toward the event. Exponential tilting or low-dimensional parametric changes of measure, often guided by large deviations theory, can closely approximate the zero-variance distribution and yield asymptotically efficient estimators [29, 33, 13]. These constructions rely on strong modeling assumptions and exploitable low-dimensional structure, e.g., the proposal search space is typically a tractable parametric family. 2
Figure 2. Importance sampling with Iterative Unalignment for rare-event probability estimation. Naive Monte Carlo from the base model p rarely observes the rare target event (e.g., harmful generations triggered by the prompt), leading to severe under- or over-estimation of µ. Iterative Unalignment perturbs the base model’s weights to obtain the proposal qδ = pθ+δ , which yields a warped activation space in which the rare target is sampled with higher frequency. Reweighting the resulting samples by the importance ratio p(y)/qδ (y) yields an accurate estimator µ b.
For LMs, rare-event estimation has a distinctive sequential and high-dimensional structure. The proposal must be a stochastic process over a sequence of discrete tokens. At each step, a token is drawn from a conditional distribution over the full vocabulary V, with that distribution depending on the entire previously generated prefix. For a fixed output length L, the number of possible sequences is |V|L , and rare events typically have global, semantic dependencies on the entire trajectory. Effective proposal design must therefore modify a collection of prefix-dependent conditional distributions in a coordinated manner so that probability mass is shifted toward the event without severely under-covering relevant event trajectories. Moreover, the trajectory-level importance weight factorizes as the product of the corresponding token-level likelihood ratios. Even modest discrepancies at individual decoding steps can accumulate over long trajectories and produce highly variable or heavy-tailed importance weights. Figure 1 illustrates this structural challenge. These features make the low-dimensional, analytically specified changes of measure commonly used in classical rare-event estimation difficult to construct directly. Since the relevant proposal family consists of prefix-dependent autoregressive distributions over a combinatorially large space of token sequences, the starting point of this paper is to parameterize this proposal family through a high-dimensional language-model weight space. This parameterization itself appears to be novel. Concurrent methods either sample tilted trajectory ensembles by MCMC without learning a proposal model [14], or interpolate LM activations by assuming access to a set of event hits (trajectories satisfying the target event) [2]. While the MCMC approach of Dorman et al. provides estimates on moderate-frequency events, its reliance on local trajectory mutations makes it difficult to discover isolated modes in combinatorial sequence spaces, yielding zero event hits on rarer events with verified probabilities below ∼ 10−5 (see Section 6.3 and Appendix F). Answering whether gradient-based search over weight-space perturbations can serve as a viable foundation for rare-event proposal construction requires confronting two coupled difficulties. The first
3
is event amplification, which involves driving an autoregressive model toward an event so rare that it is almost unobservable under naive sampling. Classical methods often use a scalar score g both to define the rare event through a threshold and to guide proposal search [15]. Rather than impose this structure, we separate the event definition from the search signal. The rare event is specified by an arbitrary binary indicator Φ, while a differentiable surrogate Sδ supplies a dense signal that guides optimization of the proposal. The second is importance weight stability. The same optimization pressure that amplifies the rare event can also concentrate probability on a small set of easily discovered event trajectories, leading to highly variable estimates. Maintaining estimator stability while sufficiently amplifying the rare event, across widely varying rarities, sequence lengths, and vocabulary sizes, requires a carefully designed optimization scheme. Guided by extensive empirical analysis, we develop a new method, Iterative Unalignment (IU) (see Figure 2 for an overview). We validate IU in controlled and verifiable settings involving events with probabilities as small as 10−9 . We summarize our main contributions below. 1. LMs as expressive proposal distributions (Section 4). We formulate the IS proposal as a perturbed language model, i.e., qδ (y) = pθ+δ (y), where pθ is the base model (see Figure 2). This turns proposal construction from an explicit search over context-dependent autoregressive distributions into a differentiable optimization problem over model parameters, enabling gradientbased search over an expressive proposal family. We believe that using an LM as a proposal distribution may provide a tractable way to assess risks as agents work autonomously over longer horizons. 2. Regularized amplification with adaptive feedback (Sections 5–5.3). We develop IU, which separates rare-event amplification from proposal-shape control. A differentiable surrogate supplies a dense signal that pushes the proposal toward the event, while a dense per-step divergence regularizer provides a tractable measure of how aggressively the proposal is departing from the base model along sampled prefixes. Its strength is adjusted online using effective sample size, which serves as a diagnostic of observed importance-weight stability. Together, these signals form a feedback mechanism that allows the proposal to amplify the event while monitoring and correcting emerging weight instability. 3. A controlled deep-tail evaluation suite and a robust metric. Estimator accuracy in the deep tail can be difficult to verify rigorously because obtaining ground truth requires many rollouts, each generated by the model. Consequently, some prior work relies on sparse ground truth hits with high variance [42], model-based autoraters [2], or evaluates continuous scalar scores where independent tail ground truth is unavailable [14]. We construct three complementary event families that admit reliable reference probabilities and allow us to evaluate gains across event rarity levels, compositional requirements, and score-based criteria over natural language. Verification is only half the difficulty because finite samples may miss the rare, large errors that dominate metrics such as relative MSE. We therefore evaluate with quantile log error, which measures how many orders of magnitude separate estimates from the true probability in 95% of repeated runs. We detail both the events and the metric in Section 6. 4. Large empirical gains in diverse regimes. We evaluate IU on more than 300 events across three settings and two models: GPT-2 Small (∼120M parameters) and Gemma-2 (∼2.6B parameters). Evaluating the spread of the estimator across repeated estimates for hundreds of events requires substantial sampling, limiting the model sizes we can study. For our Token 4
Presence events with probabilities from 10−7 to 10−9 , IU achieves over 800× compute-weighted efficiency gains over naive Monte Carlo on average, including training and sampling costs. After training, IU only required 128 samples per estimate on GPT-2 to yield an average 95th-percentile error of 1.2 orders of magnitude for events below 10−7 (Section 7). Section 2 reviews related work. Section 3 formulates the estimation problem, and Sections 4–5 develop the language-model proposal family and IU. Sections 6–7 present the evaluation setup and results. Section 8 discusses implications and limitations. Our implementation is available at https://github.com/namkoong-lab/iterative-unalignment.
2
Related work
Our work concerns estimating rare-event probabilities under the autoregressive output distribution of a fixed agentic model. It relates to the broader literature on AI alignment and safety, including work on red teaming and failure elicitation, as well as to prior work on rare-event estimation in language models and neural networks. We discuss these connections below. Alignment. Our work sits within the broader literature on AI alignment, which seeks to reduce undesirable model behavior through methods such as reward models learned from human preferences [12, 5] and latent adversarial perturbations [10, 32]. Despite substantial progress, safety tuning can suppress rather than eliminate unsafe behaviors, leaving residual failures in the long tail of the output distribution [17]. Theoretical results similarly suggest that alignment processes may leave behind small, nonzero failure probabilities [41]. Our work focuses on quantifying these residual risks: rare-event estimation seeks to measure failure probabilities even when the corresponding behaviors have become exceedingly rare. Red teaming. A related line of work is red teaming, which instead seeks to expose failures by finding inputs that elicit them. Zou et al. [47] use gradients to optimize adversarial prompt suffixes, while Chao et al. [11] use an attacker LM to iteratively refine jailbreak prompts. While both red teaming and our approach aim to expose model failures, finding an input that elicits a failure does not by itself establish how frequently that failure occurs under the model’s output distribution. Rare-event estimation in LMs. Recent work has begun to study and quantify rare failures in language models, considering different ways such failures can arise in deployment. Some work studies how risk accumulates across repeated queries or varies across inputs, while other work focuses on rare behaviors that arise from stochastic generation for a fixed input. Our work falls in the latter setting, where the goal is to estimate the probability of a rare event under the model’s autoregressive output distribution. We next discuss representative approaches across these settings. Jones et al. [21] study how rare-behavior risk grows as a model receives more queries. They extrapolate elicitation probabilities to larger deployment scales, but do not construct importancesampling proposals to estimate rare-event probabilities. IU addresses the complementary problem of estimating an event’s probability under the model’s output distribution for a fixed input. A second line of work considers rare events arising from variation over inputs. Wu and Hilton [42] use IS to estimate the probability that the base model’s next argmax token equals a target token, constructing proposals over random inputs using gradient and MCMC search. Cao et al. [9] improve 5
computational efficiency in the same setting. These methods estimate probabilities over random inputs to a deterministic predictor, whereas we hold an arbitrary input fixed and learn a proposal over stochastic output trajectories. More closely related to our setting, recent work estimates rare events directly over the model’s output distribution. Angell et al. [2] construct importance-sampling proposals using linear activation interventions. Their activation direction requires examples of the target behavior, and their mixed sampling proportions require event hits for calibration. Obtaining such trajectories can itself be a rare-event problem. In the settings we evaluated, this construction did not yield reliable estimates. We include IU AS to compare activation updates with weight updates under the same objective. Angell et al. demonstrate 10–20× efficiency gains for harmful-output probabilities around 10−4 . We evaluate on verifiably rarer events, down to 10−9 , and obtain larger efficiency gains. Their event definition also relies on an LLM judge, making verification of rare-event probabilities depend on an additional model. Dorman et al. [14] also study rare events over the output space, using a different approach from IU. They use a scalar score on token sequences to both define the events of interest and guide MCMC sampling. Their experiments aim to map out the distribution of these scores under the base model. Since their reference values rely on naive Monte Carlo, the deep tails of these distributions remain unverified, even though that is where the rare events of interest lie. We discuss their method and compare it with IU on events with verified reference probabilities in Appendix F. Rare events in neural networks. Rare-event estimation has also been studied for neural networks outside the language-model setting. Webb et al. [39] use adaptive multilevel splitting to estimate probabilities of violating safety score thresholds. Their method uses intermediate score levels to make progressively rarer violations accessible in the computer vision setting. We provide an alternative approach in natural language settings.
3
Problem Formulation
We consider a language model pθ parameterized by θ, which defines a conditional distribution over token sequences. We use base model to mean the fixed reference distribution whose event probability we estimate. In contrast to conventions in the language modeling literature, the base model need not be a pretrained model prior to post-training—we are motivated by base models with agentic capabilities. Since its parameters remain fixed throughout, we write p := pθ from this point onward. Given an input x, the model generates an output sequence y = (y1 , . . . , yL ) ∈ V L autoregressively, i.e., yℓ ∼ p(· | x, y1:ℓ−1 ) for ℓ = 1, . . . , L. Let Φ : V L → {0, 1} be an indicator function specifying a property of interest, and define the corresponding rare event E := {Φ(y) = 1}. Our goal is to estimate the probability µ := p(E | x) = Ey∼p(·|x) [Φ(y)] .
(1)
For a clean exposition, we treat the input x as fixed throughout and suppress its dependence, focusing on the induced distribution over output sequences y. This formulation accommodates a broad range of trajectory-level events, from abrupt, localized occurrences to subtle semantic shifts that compound across the entire trajectory. While deployment risk may average over a distribution 6
of inputs x, conditioning on a fixed input isolates the problem of finding rare output trajectories. Appendix I.5 discusses the distinct roles of input context and autoregressive output length. Beyond threshold-based event definitions. Canonical rare-event formulations often begin with a scalar performance score g(y) and a threshold γ, defining Φ(y) = I{g(y) > γ}. The score may represent an option price, vehicle velocity, or queue length. In addition to defining the event, g ranks samples by their proximity to it, allowing multilevel or sequential methods to progress toward the rare event through a sequence of intermediate levels. In LM settings, however, the outcome of interest may be naturally specified only as a binary trajectory-level event, such as "whether an agent issued a combination of tool calls that deleted important data." One could still place this outcome in the canonical form by using an indicator as g. However, such a score alone is not sufficient for proposal search. To guide the proposal toward the event before it is observed, the score must also provide informative intermediate values on partial trajectories and on completed non-event trajectories. Another choice for g is the conditional probability that the event eventually occurs, g(y1:ℓ ) := p Φ(y) = 1 | y1:ℓ . (2) This score is perfectly aligned with the event of interest; it quantifies, at every prefix, the remaining probability of eventually reaching it. However, computing such a score requires the recursive sum X g(y1:ℓ ) = p(u | y1:ℓ )g(y1:ℓ u). (3) u∈V
with g(y1:L ) = Φ(y1:L ). That is, evaluating such a score requires summing over all possible continuations of the current prefix, of which there can be as many as |V|L−ℓ . For even a small language model, this continuation space is enormous. GPT-2 Small has |V| = 50,257, which yields about 1047 possible sequences even at L = 10. Moreover, before the first token is generated, this score is µ, so computing this ideal score already contains the original rare-event estimation problem. The other option is to learn an auxiliary scoring model from roll-out data. For such a model to define the event exactly, thresholding its output must correctly separate event and non-event trajectories across the relevant sequence space. For it to support proposal search, it must additionally assign informative intermediate values to prefixes and to trajectories on which the event does not occur. Rare-event data provide little direct supervision for either learning objective. When the event probability is µ, one expects only one event hit per 1/µ ordinary roll-outs. Hence training such a model is a difficult rare-event problem in its own right. Constructing proposals through a separate signal. We treat the event and the search signal as separate objects. The event indicator Φ is fixed by the risk we want to measure: did the trajectory exhibit the failure of interest? The search signal is introduced only after this event has been specified and instead answers an algorithmic question: in which direction should we move the proposal to observe that failure more often? We refer to this signal as a surrogate and denote it as Sδ . We do not require our surrogate to define the event, estimate its probability, or perfectly rank all trajectories to be a useful signal for proposal search. It may instead capture a correlated or enabling condition for the failure. Increasing it should tend to make Φ = 1 more likely, without needing to characterize the event region exactly. This separation leaves Φ free to represent an arbitrary binary property of a trajectory rather than forcing the risk of interest to conform to a naturally available scalar score. 7
Related methods use examples of the target behavior or event scores to guide proposal construction [2, 14]. Returning to our data-deletion example, whether important data were deleted is well defined, even when no natural score measures progress toward that outcome. Using a separate signal correlated with the event to guide search may allow the formulation to capture more realistic risks. This modest change weakens the requirement that one score both define the event and guide proposal search. Section 5.1 develops the properties of a useful search signal.
4
Language models as proposals
We now turn to proposal design. Given a proposal distribution qδ over trajectories, importance sampling (IS) estimates µ by drawing samples from qδ and reweighting as follows. N
µ b :=
1 X p(yi ) Φ(yi ), N qδ (yi )
iid
y i ∼ qδ .
i=1
The estimator is unbiased whenever qδ (y) > 0 for every trajectory satisfying p(y)Φ(y) > 0, but its efficiency depends critically on how the proposal qδ redistributes probability mass within the event. We use coverage to describe how the proposal’s relative mass across event trajectories compares with that of the base model. Ideally, conditioning on the event should preserve these relative probabilities, so that qδ (y | E) ≈ p(y | E), or equivalently, the amplification factor qδ (y)/p(y) is approximately constant across y ∈ E. We say that the proposal under-covers an event trajectory when it assigns that trajectory too little mass relative to its base-model probability, producing an unusually large importance weight p(y)/qδ (y). Such under-coverage can produce severe variance inflation. Increasing proposal mass on event trajectories is not harmful by itself, but concentrating that mass unevenly can leave other event trajectories under-covered, as illustrated in Figure 3. A useful proposal must therefore increase the event frequency while approximately preserving the original distribution conditional on the event to keep the in-event importance weights stable. In the classical IS literature, effective proposals are often derived by exploiting analytical structure in the underlying stochastic system [15, 8]. For LMs, this approach is considerably more difficult. The proposal must modify a sequence of context-dependent conditional distributions in a coordinated manner, while preserving sufficient probability mass across the many distinct trajectories that can realize the event. The exponentially large trajectory space and the multiplicative accumulation of token-level likelihood ratios make this coverage requirement especially challenging. Our goal is therefore to develop a tractable procedure for constructing good autoregressive proposals. Weight-space perturbations as a differentiable family of proposals Our central observation is that for autoregressive LMs, proposal distributions need not be specified explicitly over complete token sequences. Instead, they can be defined implicitly through the parameters of the generative model. For the base model p with weights θ, any perturbation δ defines a proposal qδ := pθ+δ . Because the same model parameters determine the conditional distribution at every decoding step, a single perturbation coherently changes the entire collection of prefix-dependent conditionals and thereby induces a new distribution over complete trajectories. This parameterization has an important computational advantage. The mapping δ 7→ qδ is differentiable, so proposal design becomes a gradient-based optimization problem in the LM’s parameter space rather than an explicit search over distributions on V L . By contrast, the CE optimization and twisted-SMC variants evaluated
8
Good proposal
100
Probability mass
10−2
Bad proposal c (1) ) ¿ Var(¹ c (2) ) Var(¹ q± q±
(1)
(2)
q±
q±
10−4 Tail coverage Tail undercoverage
p
10−6
p
10−8 10−10
⋯ Ec
E
Ec ⋯
⋯ Ec
E
Ec ⋯
1D projection of sequences
Figure 3. Proposal design and tail coverage. A one-dimensional projection of discrete sequence space, with probability mass on a logarithmic vertical axis; thin vertical lines indicate individual (1) sequences and the shaded band marks E. Left: the broad-tailed proposal qδ amplifies the event (2) while retaining tail coverage. Right: the concentrated proposal qδ leaves event tails under-covered; (2) dark-red hatching marks a severe case of under-coverage, where qδ < p, by shading the gap between the two masses. Here µ bq denotes the importance-sampling estimator using proposal q, with equal sample counts across proposals. The variance comparison indicates that the under-covering proposal has much larger estimator variance.
here use elite selection and particle resampling, respectively, without gradients through the LM’s weights (Appendices H.1 and H.2). The tilting can be implemented at three levels of the generation process. 1. Logit-level tilting. Add a trainable |V|-dimensional vector to the output logits. 2. Activation-level tilting. Add trainable offsets to the model’s hidden activations while keeping its weights fixed. 3. Weight-space tilting. Directly perturb the model weights using methods such as LoRA [19]. These choices differ in expressive power. Activation-level tilting can change hidden representations throughout the model, but additive offsets offer less flexibility than modifying the model’s transformations through full-parameter updates or LoRA. We focus on weight-space tilting and evaluate two questions separately. First, we compare weight-space tilting with activation- and logit-level tilting to assess the value of proposal expressivity. Second, within the activation and logit parameterizations, we compare gradient-based optimization with CE optimization to assess the value of the search algorithm. Interpretation of importance weights. Under this class of proposals, the trajectory-level weight factorizes into the following product of per-step likelihood ratios along the realized sequence. L
Y p(yℓ | y<ℓ ) p(y) w(y) = = . qδ (y) qδ (yℓ | y<ℓ ) ℓ=1
9
(4)
Thus, discrepancies between the base and proposal conditionals accumulate across decoding steps. Severe under-coverage at a single step, or moderate under-coverage across several steps, can produce a large trajectory-level weight. Proposal design must therefore coordinate probability shifts across the sequence. This adds constraints to the proposal search process that must be carefully managed. It is not enough for the proposal to simply trigger the rare event via arbitrary, out-of-distribution paths. If the proposal learns to satisfy the rare-event condition using sequences that are highly unnatural to the base model, it can under-cover the natural, high-probability trajectories that the base model would actually take to reach that same event. This failure to cover the base model’s preferred paths leads directly to exploding importance weights. Therefore, the proposal must be constrained to emit the rare event in a way that is natural under the base model, discovering and amplifying the exact paths to the rare event that the base model itself most prefers.
5
Iterative Unalignment
We now have a class of proposals qδ := pθ+δ that tilt the base model by changing its weights. Let rδ = qδ (E) > 0 be the proposal’s event probability, qδE = qδ (· | E) be its conditional distribution within the event, and q ∗ = p(· | E) be the base model conditioned on the event. The variance of the N -sample importance-sampling estimator decomposes as Var(b µ) =
1 µ2 Varqδ [w(y)Φ(y)] = N N
1 − rδ r | {zδ }
event scarcity
+
χ2 (q ∗ ∥qδE ) , r | {zδ }
(5)
weight dispersion
where χ2 (q ∗ ∥qδE ) = EqE [(q ∗ /qδE − 1)2 ]. Thus, a good proposal increases rδ while keeping χ2 (q ∗ ∥qδE ) δ small. The ideal proposal q ∗ (y) = p(y)Φ(y)/µ amplifies the base-model probability of every event trajectory by the same factor 1/µ, preserving their relative probabilities. It achieves zero variance because it places all mass on the event and matches p(· | E)’s relative mass exactly, making both contributing terms vanish. Note that while q ∗ is an informative construction, it is inaccessible as it depends on the unknown probability µ. Obtaining an optimization signal for the event and controlling distortion within the event region are again rare-event problems in their own right. The event indicator provides almost no optimization signal for increasing rδ when event hits are scarce. Measuring and controlling χ2 (q ∗ ∥qδE ) is also difficult because it compares conditional distributions within the rare-event region over all rare-event sequences. To amplify the event, we introduce a dense surrogate Sδ (y) based on differentiable components of the proposal model, such as logits evaluated along the prefixes of y. To restrain within-event distortion, we introduce a tractable regularizer Rδ (y) evaluated on proposal trajectories. At practical training budgets, sequences sampled from p typically contain no event hits and therefore provide no direct observations of within-event divergence. Sampling from the proposal makes this region more observable as amplification proceeds. We apply the regularizer to all proposal trajectories, including those outside the event, to provide feedback even before event occurrences become frequent. We propose a procedure called Iterative Unalignment (IU). At iteration t, we maximize the
10
empirical objective Jt (δ) =
1 B
B X i=1
Sδ (yi ) | {z }
event scarcity
−λt
R (y ) | δ{z i}
,
iid
y i ∼ q δt .
(6)
weight dispersion
where λt > 0 trades off the two terms. Learning an effective proposal requires discovering event trajectories while preserving their relative probabilities under the base model. These requirements are closely connected because each update changes the trajectories available to guide subsequent updates. Before event hits become frequent, amplification relies on the surrogate’s model-based signals. Optimizing these signals can favor a small subset of event trajectories and reduce the probability of others, including paths that have not yet been observed. Those paths then become harder to discover, and their contribution to importance-weight variability can remain hidden from the training batches. The regularizer must therefore restrain concentration even when direct evidence of within-event distortion is sparse. We develop the surrogate and regularizer in §5.1 and §5.2, respectively. We calibrate their relative strength using the importance weights observed during training, first across the full batch and then within the event once hits become frequent enough (§5.3). These observed weights cannot certify coverage of unobserved event paths. Since token generation is discrete, we treat the sampled trajectory y as fixed when computing an IU update. We differentiate Sδ (y) and Rδ (y) only through the proposal model’s conditional distributions qδ (· | y<ℓ ) evaluated along the realized prefixes. Thus, IU backpropagates through model evaluations on the sampled trajectory, but not through the discrete sampling operation that produced the trajectory. Our derivative-free CE optimization baselines provide an alternative that does not use these model gradients.
5.1
Differentiable surrogates
We first address the event-scarcity term in Eq. (5), which decreases as the proposal event probability qδ (E) increases. Directly maximizing qδ (E) = Eqδ [Φ(y)] provides little usable signal in the rare-event regime because almost every sampled trajectory has Φ(y) = 0, and the discrete sampled tokens do not admit an ordinary pathwise gradient. We therefore retain the binary indicator Φ as the sole definition of the event and introduce a separate differentiable surrogate Sδ to guide proposal optimization. Unlike traditional formulations in which a continuous score both defines and locates the event, this separation allows the event definition itself to remain arbitrary and binary. Since the effectiveness of a surrogate depends both on the event and the optimization trajectory, our goal in this section is to identify the practical properties that make a surrogate useful and demonstrate that such surrogates may realistically exist for many rare events. Three desiderata for a useful surrogate. A useful surrogate must possess three key properties to effectively guide the proposal distribution. 1. Event observability. The surrogate Sδ should provide informative values on trajectories for which the rare event does not occur. In the rare-event regime, Φ(y) = 0 for nearly all sampled trajectories and therefore provides little guidance for proposal optimization. In contrast, Sδ can be constructed from differentiable model quantities such as token probabilities. It can therefore supply a dense optimization signal before the event is observed. 11
Surrogate value S(y)
High recall surrogate
Low recall surrogate
1.0
Under-scored event region
0.5
0.0
⋯
Ec
E
Ec
⋯
⋯
Ec
E
Ec
⋯
1D projection of sequences
Figure 4. Conceptual illustration of recall. Surrogate optimization seeks to amplify the region where surrogate values are high, corresponding to an upper level set. Ideally, this region contains the entire event E, even if it also includes non-event trajectories. Left: the high-recall surrogate scores the whole event highly. Right: the low-recall surrogate leaves part of the event outside its preferred region (dark-red). The pale red band marks E. Thin gray bars denote individual sequence scores in a one-dimensional projection of discrete sequence space; black curves guide the eye.
2. Event alignment. Increasing the surrogate should tend to increase the probability of the true event. Informally, the gradient induced by Sδ , ∇δ Eqδ [Sδ (y)] should point in a direction similar to ∇δ Eqδ [Φ(y)]. The surrogate should provide a useful direction for steering the proposal toward the event. 3. Event recall. Optimization favors regions where the surrogate takes high values, corresponding to its upper level sets. It seeks to amplify proposal probability in these regions. Recall asks how much of the event E this preferred region contains, while precision asks how much of the preferred region lies in E. Ideally, the preferred region contains the entire event. To achieve this the surrogate should assign comparably high values throughout E (much like Φ) within the event, since otherwise optimization may only prefer a subset of the event. Whether training achieves this coverage also depends on the proposal family, regularizer, and optimizer. Figure 4 illustrates recall. Low precision causes the proposal to spend samples on trajectories that the exact indicator Φ later sets to zero. Low recall may cause the optimization process to assign too little proposal probability to a set of important event trajectories, making qδ (y) ≪ p(y) and the importance weight p(y)/qδ (y) extremely large. Thus, false positives mainly reduce event observability, whereas false negatives can produce the under-coverage and highly variable importance weights described in §3. For our purposes, only high recall is required, whereas high precision is optional. The region favored by optimization should retain the important event trajectories, even if it includes false positives. Low recall is more harmful because the surrogate is the sole mechanism that guides the proposal toward the event. The regularizer introduced in §5.2 can control importance weight dispersion, but it cannot recover important event trajectories that the surrogate systematically neglects. Potential surrogates for complex rare events. Suppose the event is that an autonomous agent successfully gains elevated permissions and issues a destructive delete command. Constructing a differentiable surrogate for the complete multi-stage event may be difficult. Issuing the delete 12
command is a necessary condition. A surrogate that scores trajectories containing this command highly can favor the full event, along with attempts that lack the required permissions. It therefore has imperfect precision. More generally, a differentiable surrogate for a necessary condition can be useful even when its preferred region extends beyond E. If the event is much more common in this region than under the base model, shifting proposal mass there can make the event more observable, provided its important trajectories retain coverage. Preliminary experiments suggest that IU can tolerate this broader surrogate coverage, though estimation error increases as the mismatch grows (Appendix I.1).
5.2
Controlling within-event distortion via regularization
The regularizer Rδ is intended to control χ2 (q ∗ ∥qδE ), the within-event distortion in Eq. (5), while allowing the surrogate to amplify the event. This requires assessing how the proposal’s conditional distribution within E differs from q ∗ . Recall that we evaluate Rδ using samples from qδ . Samples from the base model p rarely visit E, so evaluating the regularizer on them mainly constrains the proposal on typical sequences outside the event. As amplification raises rδ , proposal samples visit the event more often and provide more direct feedback on importance weights within it. Appendix I.2 compares joint forward KL on base-model rollouts, event-gated regularization, and on-policy forward KL. Written as an expectation under the proposal qδ , the chi-square divergence is: " 2 # r p(y) δ χ2 (q ∗ ∥qδE ) = 2 Ey∼qδ Φ(y) − 1. (7) µ qδ (y) Although controlling this divergence directly controls the within-event contribution to variance, estimating it from a finite batch can give a misleading picture of coverage. Even if the event weights in a finite batch show little dispersion, other event paths carrying substantial mass under q ∗ may be severely under-covered. The variance benefit depends on preserving the base model’s relative probability mass across the entire event region, while a finite batch reveals only a small subset. Agreement on the observed paths therefore does not establish that χ2 (q ∗ ∥qδE ) is small. This coverage problem is compounded by sampling from the proposal itself. Paths with substantial mass under qδ are sampled frequently, while event trajectories can remain unobserved despite their large importance weights p(y)/qδ (y). These large weights can make such trajectories major contributors to the squared-weight expectation in (7). Consequently, proposal batches may substantially underestimate the divergence by missing its largest contributors. The population identity remains exact, but estimating it for regularization can be least reliable where the proposal’s coverage is worst. Controlling concentration. To address this measurement difficulty, IU instead seeks to limit excessive concentration on the event paths that the surrogate is amplifying. At the level of the event-conditional distributions, this motivates the reverse divergence qδE (y) E ∗ KL(qδ ∥q ) = Ey∼qE log ∗ . (8) δ q (y) This direction penalizes disproportionate allocation of proposal mass to particular event paths. As those paths gain proposal mass, they also become more likely to appear in the samples used to 13
Frequency (%)
Probability mass
Overweighted
q± p Underweighted
-3
-1
1
Outcome, y
3
5
Population KL(q± (¢ j E)kq ¤ )
15
10
5
0
2
3
4
5
Frequency (%)
100
E
Population  2 (q ¤ kq± (¢ j E))
80 60 40 20
6
Batch estimates of KL(q± (¢ j E)kq ¤ )
0
1
10 100 10³ 10⁴ 10⁵
Batch estimates of  2 (q ¤ kq± (¢ j E))
Figure 5. Finite-batch measurement of event-conditional divergences in a discrete toy example. The proposal concentrates on one side of the event region, leaving the other side under-covered. Most proposal batches miss the under-covered region, so the chi-square estimates (right panel) severely underestimate the population value. In contrast, the reverse-KL estimates (center panel) cluster near their population value.
measure the penalty. Its ideal preference remains q ∗ . Over unrestricted conditional distributions, both KL(qδE ∥q ∗ ) and χ2 (q ∗ ∥qδE ) attain their unique minimum of zero at qδE = q ∗ . Thus, at any fixed event probability rδ , the two criteria favor the same allocation within the event. Figure 5 gives a complementary finite-batch view. Although both KL(qδE ∥q ∗ ) and χ2 (q ∗ ∥qδE ) are expectations under qδE , they behave very differently empirically. Reverse KL is dominated by regions that receive substantial probability under the proposal and is therefore well represented by typical proposal batches. In contrast, the chi-square divergence places its largest contributions on trajectories that are underrepresented by qδ ; these trajectories are precisely the least likely to appear in a finite batch. Consequently, empirical chi-square estimation becomes least reliable when proposal under-coverage is most severe. This observation does not imply that reverse-KL regularization guarantees coverage of the entire rare-event region. Event modes systematically missed by the surrogate cannot be recovered by the regularizer alone. The surrogate identifies the region to amplify, while the reverse-KL penalty limits how aggressively probability can concentrate within that region. Extending the regularizer’s control. Even under proposal samples, event hits can still be scarce early in training. We therefore apply the penalty to all proposal trajectories by replacing the c event-conditional divergence with KL(qδ ∥p). Writing qδE = qδ (· | E c ) and pE c = p(· | E c ), we have rδ 1 − rδ + (1 − rδ ) log µ 1−µ E ∗ + rδ KL(qδ ∥q )
KL(qδ ∥p) = rδ log
(9)
c
+ (1 − rδ ) KL(qδE ∥pE c ). The joint divergence penalizes distortion both within and outside the event. Extending regularization to non-event trajectories constrains the proposal even when the training batch contains no event hits. Estimating this penalty from trajectory log-likelihood ratios uses only the probabilities of tokens that happen to be sampled. At each visited prefix y<ℓ , however, both models provide complete 14
Figure 6. Reverse-KL regularization produces more concentrated in-event weights in this comparison. We show in-event importance weights under dense per-step reverse-KL (left) and chi-square (right) regularization. The chi-square weights span many more orders of magnitude. The solid black line marks the expected in-event weight µ/rδ , and the dashed blue line marks the observed mean.
conditional distributions over the vocabulary. We use this information by computing their reverse KL at every decoding step and summing over the trajectory, Rδ (y) :=
L X
KL(qδ (· | y<ℓ )∥p(· | y<ℓ ))
ℓ=1
=
L X X ℓ=1 u∈V
(10) qδ (u | y<ℓ ) qδ (u | y<ℓ ) log . p(u | y<ℓ )
By the autoregressive chain rule, Ey∼qδ [Rδ (y)] = KL(qδ ∥p). Thus, each visited prefix supplies feedback on all possible next tokens without changing the population divergence. The prefixes themselves are still sampled from qδ . We use Eq. (10) in the objective (6), holding the current proposal’s sampled prefixes fixed during each update. We apply it to every rollout without gating by Φ, so non-event completions also contribute their conditional distributions. Figure 6 compares dense per-step reverse-KL and chi-square regularization. For chi-square, we use the dense per-step sum, which gave the best empirical performance among the chi-square implementations we tested (Appendix I.3). In this comparison, reverse KL produces more concentrated in-event importance weights.
5.3
Adaptive regularization
In this section, we discuss how to set the strength of the regularizer in IU, i.e., λt in (6). Too little regularization lets the proposal concentrate on a few event paths, while too much suppresses amplification and leaves many evaluation batches without an event hit (Figure 7). The appropriate strength depends on the event’s rarity, the sequence length, and the proposal family, and changes as training proceeds. We therefore adjust λt in Eq. (6) using the observed importance weights. For a batch {yi }B i=1 ∼ qδt , let wi = p(yi )/qδt (yi ) and Φi = Φ(yi ). We define the empirical 15
Figure 7. Fixed regularization fails at both extremes. With too little regularization, the proposal collapses onto the surrogate mode, driving the importance weights and probability estimates toward severe underestimation and therefore inflated sample cost. With too much regularization, weak amplification reduces efficiency and causes the sample cost to rise as well. Adaptive regularization avoids both regimes. Triangles mark costs above the plotting limit; their heights are capped for readability. Here sample cost is the number of trajectories required to achieve the same estimation accuracy as adaptive regularization (Eq. (17)).
full-batch and in-event effective sample size ratios (ESSr) as 2 2 PB PB i=1 Φi wi i=1 wi , ESSrevent = PB ESSrbatch = PB . PB 2 B i=1 wi2 i=1 Φi i=1 Φi wi Each ratio decreases as the observed weights become more dispersed relative to their mean. The in-event ratio is defined only when the batch contains event hits. For a fixed proposal qδ with the same support as p, the population full-batch ratio equals exp[−D2 (p∥qδ )], where D2 is Rényi divergence of order two. This motivates limiting global departure from the base model during exploration, when event hits are scarce. Once events become sufficiently frequent, we use the in-event ratio, whose population value is 1/[1 + χ2 (q ∗ ∥qδE )] = exp[−D2 (q ∗ ∥qδE )]. The common scale factor µ/rδ in the event weights cancels from this ratio, so it isolates within-event distortion from common event amplification. P Let ĥt = B −1 B Φ be the empirical event frequency. At each iteration, we select the empirical i i=1 ratio and its target using ( (ESSrbatch , ρbatch ), ĥt < h⋆ , (ESSrt , ρt ) = (ESSrevent , ρevent ), ĥt ≥ h⋆ . We set h⋆ = 0.1 in our experiments and typically choose ρevent > ρbatch to target less relative weight dispersion within the event. We initialize λ0 ≥ λmin > 0 at a value suited to the scale of the regularizer and adapt it using λt+1 = max λmin , λt exp ηλ (ρt − ESSrt ) , (11) 16
15%
0.50
10%
Surrogate (Sδ)
0.25
Regularizer (δ)
5%
Event % (ĥt)
0.0 0
100
200
1.0
λt ESSr ρ
0.08
0% 300
0.6 0.06 0.4 0.04 0.02 0
Training step
0.8
ESSr, ρ
0.75
0.10
λt
Normalized Sδ, δ
20%
Event proportion (%)
25% 1.0
100
200
0.2 0.1 0.0 300
Training step
Figure 8. IU training dynamics over 300 optimization steps. Left panel. The empirical event proportion ĥt (in dashed teal) rises during training. The surrogate Sδ (blue) and regularizer Rδ (red) are each normalized by their terminal value for plotting. Right panel. Adaptive regularization brings the event-specific ESSrE (black) to its target ρE (dashed blue) by adjusting the multiplier λt (orange).
where ηλ > 0 is a step size. If the active empirical ESSr falls below its target, λt increases; if it exceeds the target, λt decreases toward its floor. Figure 8 illustrates this feedback during training. These population identities do not make empirical ESSr a certificate of population ESSr or a bound on chi-square divergence. Unobserved event modes can make χ2 (q ∗ ∥qδE ) arbitrarily large while empirical ESSr remains high, so reliable measurement faces the same finite-sample limitations discussed in §5.2. We use empirical ESSr as coarse feedback about observed weight dispersion to tune λt . The dense reverse-KL term in Eq. (10) supplies the optimization gradient across the vocabulary at every visited prefix, while ESSr adjusts how strongly it is applied. This lets observed weight dispersion guide regularization strength without requiring a reliable divergence estimate as the optimization objective. The indicator Φ selects weights for in-event feedback, while the penalty continues to apply to every rollout.
5.4
Putting it all together
Combining the surrogate from Section 5.1 with the regularizer from Section 5.2, IU updates the proposal parameters by maximizing the objective Jt in Eq. (6). The regularization strength λ is adjusted online according to the adaptive rule in §5.3. Because the training trajectories are sampled from the current proposal, the updates are concentrated on the regions of the trajectory space that the proposal actively explores. After training, we draw a fresh set of trajectories from the learned proposal qδ and estimate µ using exact importance weights. Algorithm 1 summarizes the end-to-end procedure, comprising gradient updates on δ, adaptive updates on λ, and a final importance-sampling stage under the learned proposal after training.
6
Experimental setup
6.1
Evaluation settings
Evaluating a rare-event estimator requires a reference ground truth value of µ to validate against. However, the very scarcity that motivates rare-event estimation makes obtaining reference values 17
Algorithm 1 Iterative Unalignment (IU) Inputs: p, Φ, Sδ , Rδ ; T, B, N ; ηδ , ηλ ; ρbatch , ρevent ; h⋆ ∈ (0, 1); λ0 ≥ λmin > 0. 1: δ1 ← 0, λ1 ← λ0 2: for t = 1, 2, . . . , T do iid
Sample y1 , . . . , yB ∼ qδt 4: wi ← p(yi )/qδt (yi ), Φi ← Φ(yi ) for i = 1, . . . , B P 5: Jt (δ) ← B1 B i=1 Sδ (yi ) − λt Rδ (yi ) 6: δt+1 ← δt + ηδ ∇δ Jt (δ) δ=δt P 7: ĥt ← B −1 B Φ i=1 ( i (ESSrbatch , ρbatch ), ĥt < h⋆ , 8: (ESSrt , ρt ) ← (ESSrevent , ρevent ), ĥt ≥ h⋆ . 9: λt+1 ← max λmin , λt exp ηλ (ρt − ESSrt ) 10: end for iid 11: Sample y1 , . . . , yN ∼ qδT +1 P p(yj ) 12: µ b ← N1 N j=1 qδ (yj ) Φ(yj ) 3:
▷ Fixed sampled trajectories ▷ Eq. (6) ▷ Update proposal ▷ Update event proportion
▷ Adapt regularization ▷ Fresh trajectories for estimation ▷ IS estimator
T +1
13: return (b µ, qδT +1 )
through naive Monte Carlo computationally prohibitive. Attaining a 10% relative standard error for an event at µ = 10−9 requires approximately 1011 raw-indicator rollouts via naive Monte Carlo (see Appendix G). Angell et al. [2] compare against Monte Carlo references with a minimum nonzero value of 10−4 , while Dorman et al. [14] focus on continuous scalar scores where independent deep-tail reference probabilities are unavailable. A good evaluation set should be both realistic, in the sense of faithfully representing deployment failures, and tractable enough to admit trustworthy reference probabilities. Existing works rely on LM autoraters whose reliability is itself a rare-event problem [2], or target continuous heuristic scores such as readability and output log probability [14]. Constructing events that are simultaneously realistic, rare, and verifiable is an entire research topic of its own. As our contribution is primarily algorithmic, we focus on events that preserve the difficulty of proposal design while allowing estimator accuracy to be verified. We use Token Presence for systematic evaluation and two complementary case studies to evaluate IU (Table 1). Token Presence allows ground truth to be computed efficiently in parallel across batches for the entire vocabulary, enabling systematic sweeps over deep event rarities (e.g., µ = 10−9 ) and longer generation lengths (up to 50 tokens). Compositional Token Presence extends this setting to multiple target tokens, either in a specified order (Ordered Sequence events) or in any order (Conjunctive events). Our Profanity Classifier setting connects the evaluation to the traditional formulation in which an event is defined by a continuous score crossing a threshold. All three settings are sequence-level and retain the core challenge of conducting importance sampling in our setting. Every importance weight carries a product of L step-wise likelihood ratios, so the challenges of the combinatorial |V|L output space and the need for qδ to remain disciplined across the full trajectory remain present.
18
Evaluation setting
Motivation
Token Presence
Highly verifiable reference values across models, lengths, and rarities, while preserving complex, multi-path trajectory dynamics.
Compositional Token Presence
Multiple target tokens in a specified order (Ordered Sequence events) or in any order (Conjunctive events).
Profanity Classifier
Traditional score-threshold rare-event definition. For LMs, this is a continuous score defined over a sequence of tokens. Table 1: Proposed evaluation suite.
6.1.1
Token Presence events
Let y = (y1 , . . . , yL ) ∈ V L be the output sequence. For a target token v ∈ V, the Token Presence event is ΦTOK (y; v) = I{∃ ℓ ∈ {1, . . . , L} such that yℓ = v} (12) so that ΦTOK (y; v) = 1 if and only if v appears at least once. Token Presence defines a highly multimodal event region. A target token can appear after many different prefixes and at many different positions. Conjunctive and Ordered Sequence events introduce further combinations of prefixes and positions through which the event can occur. A successful proposal must therefore amplify the event without collapsing onto a small subset of these trajectories. Finding a single adversarial path is insufficient for estimating the event’s total probability. For an event with base probability µ ≈ 10−9 , observing it reliably under a modest estimation budget (e.g., N = 512) requires amplification by over 106 , while retaining coverage across the event region. Weight perturbations must therefore amplify the event while preserving coverage of its trajectories. Token-level conditions also serve as natural proxies for real-world safety risks. For instance, prior work measures the risk of a model generating harmful instructions, such as for synthesizing chlorine gas, by tracking the probability assigned to critical keywords like bleach [21]. This direct connection to deployment safety further motivates our use of the Token Presence setting to systematically verify our algorithm. Token Presence also admits a low-variance, unbiased reference estimator constructed directly from the model’s next-token probabilities. Moreover, the next-token distributions recorded along a single rollout provide reference estimates for every token in the vocabulary simultaneously. For a rollout y ∼ p, we replace the sparse binary indicator with a sum of conditional first-hit probabilities: b TOK (y; v) = Φ
L X
p(v | y<ℓ ) I{v ∈ / y<ℓ }
(13)
ℓ=1
whose expectation is exactly µTOK (v) = Ep [ΦTOK (y; v)]; Appendix G.1 gives the derivation for this identity. This estimator remains precise even when the token is almost never observed. Appendix G.3 reports the resulting reference probabilities and their measured uncertainty, including the rarest event used in each evaluation setting. Characterizing estimator spread across many Token Presence events requires substantial sampling, which limits model size and generation length (Section 7).
19
To train the proposal, we use the target token’s step-averaged log-probability as the differentiable surrogate in the IU objective (6): SδTOK (y ; v)
L 1X = log qδ v | y<ℓ . L ℓ=1
6.1.2
Compositional Token Presence events
As a case study, we consider two multi-token extensions of Token Presence: Ordered Sequence events and Conjunctive events. We refer to both as compositional events. Let v = (v1 , . . . , vm ) be a tuple of target tokens. Ordered Sequence (SEQ) events. The Ordered Sequence event is ΦSEQ (y; v) = I ∃ 1 ≤ ℓ1 < · · · < ℓm ≤ L such that yℓj = vj for all j ,
(14)
which requires the targets to appear in order but not necessarily contiguously. Let cℓ be the number of targets already completed in order by the prefix y<ℓ . Our surrogate rewards the next required token, advancing to the following target when that token appears: SδSEQ (y ; v)
L 1X I{cℓ < m} log qδ vcℓ +1 | y<ℓ . = L ℓ=1
Conjunctive (AND) events. The Conjunctive event is ΦAND (y; v) = I{∀j ∈ {1, . . . , m}, ∃ℓ ∈ {1, . . . , L} such that yℓ = vj } ,
(15)
which requires every target token to appear in any order. The proposal must cover all target orderings, which makes it harder to learn. We use a surrogate that raises the probability of every target token throughout the trajectory: SδAND (y ; v)
m L 1 XX log qδ vj | y<ℓ . = L j=1 ℓ=1
Like Token Presence, both compositional events admit unbiased, low-variance reference estimators from the base model’s next-token probabilities along ordinary rollouts. Appendix G.2 defines these estimators and proves their unbiasedness. As it is harder to verify reference probabilities for compositional events in bulk, we present a set of 5 events per setting, and defer a systematic study for future work. 6.1.3
Profanity Classifier
As our second case study, we estimate the probability that an LM generates text whose classifier probability exceeds a profanity threshold κ ∈ (0, 1), so that ΦPROF (y; κ) = I{f (y) > κ}. This is an instance of the canonical threshold event of Section 3 where a scalar score defines E by crossing a cutoff. Traditional rare-event methods operate in this regime, as do the LM estimators of Angell et al. and Dorman et al., where the same score both specifies the event and supplies the search signal [2, 14]. Although we argue for separating the binary event Φ from the differentiable surrogate Sδ 20
used to train the proposal, we include this setting to show that IU remains effective when the event takes the traditional score-threshold form. The classifier f is a bag-of-words logistic regression that assigns every vocabulary token u ∈ V a learned weight βu . It converts the sum of token weights and a bias into a predicted probability: ! L X f (y) = σ βyℓ + b = σ β ⊤ count(y) + b , ℓ=1 |V|
where σ(z) = (1 + e−z )−1 , β ∈ R|V| is the per-token weight vector, b is a bias, and count(y) ∈ Z≥0 is the token-count (bag-of-words) vector of y. The event occurs when the predicted probability f (y) exceeds κ. Unlike token-based events, the Profanity Classifier does not admit an unbiased, low-variance estimator that recovers µ in the deep tail. We therefore rely on the raw binary indicator and choose κ so that brute-force verification remains feasible, targeting rarities around µ ≈ 10−5 , comparable to the verification regime of Angell et al. [2]. Differentiable surrogates for scores over discrete tokens. Although f defines the event, it is evaluated on discrete realized tokens and is not differentiable with respect to the LM parameters δ. We therefore use the differentiable surrogate below. We did not compare alternative surrogates in this setting. Let qδ (y<ℓ ) denote the vector of proposal next-token probabilities over V at prefix y<ℓ . The surrogate we choose is made differentiable by replacing each realized token weight βyℓ with its expectation under this distribution, β ⊤ qδ (y<ℓ ). Summing these expected contributions gives a soft version of the sequence-level classifier logit. We subtract the corresponding logit threshold log[κ/(1 − κ)] and apply a log-sigmoid to provide informative gradients below the threshold: ! L X κ PROF ⊤ Sδ (y ; κ) = log σ β qδ (y<ℓ ) + b − log . 1−κ ℓ=1
6.2
Evaluation metrics
We evaluate the spread of the estimator across repeated estimates. For an event with reference probability µ > 0, we evaluate each learned proposal using K independent estimates, {b µk,N }K k=1 , each based on N trajectories. We use an evaluation budget of 64,000 trajectories per estimator–event pair. Limitations of empirical relative mean squared error. A standard measure of estimator accuracy is relative mean squared error (rMSE), which we estimate from K independent runs as 1 PK 2 (b µ k=1 k,N /µ − 1) . This empirical average, however, can be unstable under practical evaluation K budgets. In our setting, importance weights are products of per-token likelihood ratios, which can yield highly skewed estimator distributions with rare, extreme overestimates even when most estimates are close to µ. Squared relative error penalizes multiplicative overestimation and underestimation asymmetrically. Underestimates each contribute at most 1 to the loss, since 0 ≤ µ bk,N /µ ≤ 1 for any underestimate. By contrast, large overestimates incur quadratically growing penalties. For example, estimates of 10µ and 100µ incur losses of 81 and 9801, respectively, far exceeding the maximum underestimation loss of 1. Consequently, population rMSE can be dominated by overestimates that 21
occur too rarely to be reliably represented in K evaluation runs. Empirical rMSE may substantially understate the population value when these extremes are absent, then increase sharply when a single extreme estimate appears. Although the empirical average is unbiased when population rMSE is finite, its variability can make comparisons between estimators unreliable. Prior work likewise shows that reliable variance-based comparisons of rare-event estimators can require prohibitively many simulation runs [22]. Figure 9 illustrates this failure mode; Figure 15 in the appendix gives an empirical example. In our experience, stabilizing empirical rMSE typically requires an impractically large K. Quantile log error. For rare-event probabilities spanning many orders of magnitude, a natural accuracy objective is to recover the correct order of magnitude [6]. A logarithmic scale expresses this directly. An error of one order of magnitude corresponds to a factor of ten at any event probability, giving accuracy a common interpretation across rarity levels. We therefore measure error as | log10 (b µN /µ)|, following the use of log accuracy ratios to measure errors in orders of magnitude [36, 24]. For example, estimates of 10µ and 0.1µ both incur a log error of 1. Thus, each additional factor of ten in either direction adds one unit of error. For equal additive deviations from µ, log error penalizes underestimation more heavily than overestimation. For example, estimates of 0.1µ and 1.9µ both differ from µ by 0.9µ, yet incur log errors of 1 and approximately 0.28, respectively. This property is useful for our risk-assessment objective, since severely understating a rare-event probability can create false reassurance. We retain zero estimates and assign them the maximum error of +∞, since they represent a complete failure to observe the event. We summarize log errors across repeated estimates using the quantile log error, the 95th percentile of absolute log error. For example, a value of 1 means that at least 95% of estimates fall within one order of magnitude of the true probability. We define µ bN E0.95 (N ; µ) := Q0.95 log10 , (16) µ where Q0.95 denotes the 95th percentile across repeated runs. We estimate E0.95 (N ; µ) using the empirical 95th percentile of the K observed absolute log errors. We use quantile log error to answer two questions. How does the required sampling budget change as required estimation error tightens? ⋆ For a chosen estimation error threshold E0.95 > 0, we compare the minimum sampling budget needed to meet it: ⋆ ⋆ N ⋆ (E0.95 ; µ) := min {N ∈ N>0 : E0.95 (N ; µ) ≤ E0.95 }. (17) ⋆ As an example, E0.95 = 1 requires at least 95% of repeated estimates to be within a factor of 10 of µ. As the error threshold tightens, some methods do not meet the target within our evaluation budget. We sketch the rest of the curve by first bootstrapping to a larger budget, then adding simulated naive-MC samples if the target remains unmet. We retain the initial IS estimates and count all IS and MC samples toward the required budget. Appendix C details the procedure and explains when these estimates are conservative. In plots that average across events, the error-threshold axis stops at 0.5 because tighter targets can require more naive-MC samples than our compute budget allows us to validate.
22
-2 -4
95% log-quantile interval Min–max
Mean Median
Unobserved tail
Reported estimation error
log10 μ ̂
-6
Repeated estimates
μÂ
μB̂
∼ 106
0.71
1.4k
E0.95 (ours)
∼ 102
1.1
1.5
rMSE ̂
∼ 102
0.63
0.62
Metrics
-8
rMSE ̂
-10 -12 -14
Misleadingly low
Repeated estimates ∼ 106 ∼ 102
μÂ
μB̂
Figure 9. Limited repetitions can miss the tails that dominate rMSE. This schematic compares two estimators at a fixed sampling budget per estimate. The constructed reference of 106 estimates represents a fuller evaluation of rMSE, including rare extremes that can dominate the loss in heavy-tailed regimes. With only ∼ 102 estimates, these extremes may go unobserved. A’s rMSE changes little between the two samples, whereas B’s drops sharply when its upper tail is missed, reversing their ranking. Quantile log error favors A in both samples, reflecting B’s persistent underestimation. The dashed line marks the true probability; hats denote metrics computed from the smaller sample. ⋆ (E ⋆ ; µ) be the budget naive MC requires to meet the same error threshold; the Let NMC 0.95 Estimator Efficiency Gain is ⋆ (E ⋆ ; µ) NMC 0.95 (18) ⋆ ; µ) . N ⋆ (E0.95
Values above one indicate fewer trajectories than naive MC at the same accuracy and reliability, after proposal training. To account for both proposal training and sampling costs, we define the total per-trajectory IS cost as costtrain ⋆ costtotal (19) IS (N ) := costIS + ⋆ ; µ) . ⋆ N (E0.95 We then report the Compute-Weighted Efficiency Gain as the ratio of total compute required by naive Monte Carlo to that of importance sampling: ⋆ (E ⋆ ; µ) cost NMC MC 0.95 . ⋆ ; µ) costtotal (N ⋆ ) N ⋆ (E0.95 IS
(20)
Here, costtrain is the upfront cost of training the proposal, costIS is the cost per IS sample, and costMC is the cost per naive-MC trajectory. Appendix E gives the measured costs across models and sequence lengths. Can a small fixed budget yield useful estimates? Another practically useful question for an estimator would be whether it can quickly pin down a rare event’s probability to be within certain orders of magnitude. We therefore also verify how much accuracy a small number of samples can 23
provide after proposal training, evaluating E0.95 using just N = 128 trajectories per estimate, and presenting that result as event probability becomes exceedingly rare.
6.3
Baselines for rare-event estimation
We compare Iterative Unalignment (IU) against two families of baselines, namely naive Monte Carlo and cross-entropy (CE) methods. Additionally, we evaluate variants of our own approach that operate on restricted proposal families to isolate the contribution of weight-space tilting. Naive Monte Carlo. The standard baseline draws N i.i.d. samples from the base model p and estimates µ via the sample mean of Φ(y). Iterative Unalignment (IU). We implement the IU proposal using LoRA [19] on both GPT-2 Small and Gemma-2. For Token Presence, we train each proposal with batch size B = 128 for T = 300 steps, using 38,400 trajectories. LoRA introduces trainable adapter weights that modify the model’s transformations. We found in preliminary experiments that full-parameter tilting yielded comparable performance, so we use LoRA to reduce training cost. Appendix D gives the optimization, regularization, and evaluation settings. IU variants (IU AS, IU Logit). To assess the effect of proposal expressivity, we evaluate two variants with fewer trainable parameters. IU AS adds a single |h|-dimensional vector to the residual stream at the output of every transformer block, after its attention and MLP computations. The vector is shared across layers and token positions. IU Logit adds a single |V|-dimensional vector to the output logits. Both variants keep the base-model weights fixed and optimize the objective in Eq. (6), treating sampled tokens as fixed. They use the same regularization and adaptive λ tuning as IU. Cross-entropy methods (CE AS, CE Logit). We evaluated several Cross-Entropy (CE) variants; Appendix H.1 provides detailed algorithm formulations and additional experimental results, and Appendix D gives the hyperparameters. CE AS and CE Logit use the same steering vectors as IU AS and IU Logit. These parameterizations have fewer parameters than LoRA, allowing us to search for tilts using CE. In preliminary experiments, providing CE with more flexibility through methods such as LoRA resulted in poorer performance. For our primary baselines (CE AS and CE Logit), we use the unweighted Cross-Entropy method for optimization [28, 7]. In each iteration, this approach evaluates sampled candidates using a score function and updates the sampling distribution toward high-scoring trajectories to drive generation toward the target event. Classical formulations, such as Importance-Weighted CE (Algorithm 2), either failed to sufficiently amplify the target event or collapsed onto a narrow subset of trajectories across experimental settings. Other output-space methods. We do not include Angell et al. [2] as a baseline as their method requires representative examples for constructing activation directions and event hits as calibration samples for proposal tuning. These requirements are restrictive in our setting, where obtaining such samples is itself a rare-event problem. We therefore include IU AS as a controlled activation-space baseline.
24
Dorman et al. [14] sample score-biased trajectory ensembles with MCMC. Their framework permits the target indicator to differ from the biasing scalar score, allowing both methods to be evaluated on the same Token Presence events. We use these events for their verified tail reference probabilities, which their original score distributions lack. We evaluate increasingly rare events while allocating Dorman a conservatively charged compute budget equal to 1.38× the full cost of an N = 128 IU estimate (including proposal training). While Dorman et al. estimate moderatefrequency events accurately, their sequence-space MCMC sampler records no event hits for events below ∼ 10−5 , precluding a finite E0.95 comparison in the deep tail. Appendix F gives the full comparison.
7
Results
Our evaluation covers 319 events across three settings and two models (GPT-2 Small and Gemma-2). Reference probabilities are computed with relative standard errors below 10% (Appendix G.3). For each estimator and event, we use an evaluation budget of 64,000 trajectories. For Token Presence alone, our training and evaluation budget amounts to approximately 25.0 million forward passes on batches of 128 sequences (Appendix E.1). Evaluating estimator spread across repeated estimates for many events limits the model sizes and generation lengths we can study. We summarize errors, required budgets, and efficiency gains across events using geometric means. This averages proportional changes symmetrically and gives aggregate gains no larger than the arithmetic mean. To compare sample requirements at a given accuracy, we plot estimator efficiency gains over naive MC as the error threshold tightens (higher is better). To compare accuracy at a small fixed sampling budget, we plot quantile log error at N = 128 trajectories across event probabilities (lower is better).
7.1
Token Presence
Token Presence events preserve the highly multimodal nature of rare events in our setting while enabling precise verification of estimator accuracy across all rarities, lengths, and models via low-variance reference values. Error threshold. On GPT-2 Small at L = 50, IU achieves geometric-mean estimator efficiency gains above 20,000× over naive Monte Carlo at every displayed error threshold (Figure 10, left). Compute-weighted efficiency gains exceed 800× at every displayed threshold for this rarity band at each length L = 10, 30, and 50 (Section 7.4). IU also retains more of its estimator efficiency gains as accuracy requirements tighten than IU AS and IU Logit. This suggests that IU benefits from the flexibility of weight-space perturbations. CE AS and CE Logit, while achieving substantial estimator efficiency gains at some accuracy thresholds on shorter sequences (Appendix Figure 24), provide no measurable estimator efficiency gains over naive Monte Carlo on these longer events. As accuracy requirements tighten, IU’s efficiency gains decline, suggesting a gradual degradation in control over in-event weights. The efficiency frontiers include naive-MC samples when the estimator itself did not meet the required error within our available compute (Appendix C).
25
Token Presence, GPT-2, L = 50
Estimation error at budget N = 128 (lower is better)
Efficiency gain across error thresholds (higher is better)
Estimation error, E0.95
Estimator efficiency gain
106 105 104 103 102 101
101
100 IU IU AS IU Logit
100 2.0
1.5
1.0
⋆ Estimation error threshold, E0.95
0.5
0
10−3
10−4
10−5
10−6
10−7
CE AS CE Logit
10−8
10−9
Event probability, μ
Figure 10. Token Presence estimation on GPT-2 Small at L = 50. Left: estimator efficiency gains as accuracy requirements tighten. Right: error with N = 128 samples across event probabilities. IU maintains larger gains and lower errors than the restricted updates and CE baselines. Curves show geometric means over ten events in 10−8 ≤ µ < 10−7 on the left and ten events per rarity band on ⋆ the right; shading shows ±0.5 standard deviations in log space. At E0.95 = 1, the geometric-mean required budgets are approximately 1,103 trajectories for IU and 9.24 × 107 for naive MC. Absolute budgets appear in Appendix Figure 18.
Rarity. As events become rarer, IU’s required sampling budgets are significantly smaller than those of naive Monte Carlo. Across all evaluated sequence lengths, IU’s estimator efficiency gains also consistently increase with event rarity (Figure 11). With just 128 trajectories per estimate after training, IU’s error remains stable across rarity bands. Its log error is only 1.17 orders of magnitude across the 60 GPT-2 Small events with 10−9 ≤ µ < 10−7 at L = 10, 30, and 50. Length. IU’s performance remains stable as sequence length increases. For the rarest evaluated ⋆ events (10−9 ≤ µ < 10−8 at E0.95 = 1), its geometric-mean estimator efficiency gains exceed 106 at every tested length (L = 10, 30, and 50) (Figure 11). This resilience may stem from the regularizer’s dense, per-step control over the proposal distribution, which limits importance-weight variance along longer trajectories. Model. We compare both models at L = 10. At this length, IU still achieves larger geometric-mean estimator efficiency gains than IU Logit at every displayed threshold on both models (Figure 12, left). Its gains remain substantial but are generally smaller on Gemma-2, where its geometric-mean errors at N = 128 are also larger across rarity bands (right). Gemma-2’s larger vocabulary (256,000 versus 50,257 tokens) may contribute to this decline. Improved surrogates or adaptive regularization may better control this larger space.
26
Estimator efficiency gain
Token Presence, GPT-2, IU
10−9–10−8 10−8–10−7
10−7–10−6 10−6–10−5
L = 10
10−5–10−4 10−4–10−3
L = 30
L = 50
107 105 103 101 2.0
1.5
1.0
0.5 2.0
1.5
1.0
0.5 2.0
⋆ Estimation error threshold, E0.95
1.5
1.0
0.5
Figure 11. IU estimator efficiency gains for GPT-2 Small Token Presence at L = 10, 30, and 50 (left to right). Gains tend to increase with event rarity across all three lengths. Each curve shows a geometric mean over ten events with neighboring rarity; shading shows ±0.5 standard deviations in log space. Higher gains are better, and accuracy requirements become stricter toward the right. ⋆ At E0.95 = 1 in 10−8 ≤ µ < 10−7 , geometric-mean IU versus naive-MC budgets are approximately 349 versus 8.73 × 107 , 255 versus 7.73 × 107 , and 1,103 versus 9.24 × 107 trajectories, respectively. Absolute budgets appear in Appendix Figure 19. Token Presence, GPT-2 and Gemma, L = 10
Estimation error at budget N = 128 (lower is better)
Efficiency gain across error thresholds (higher is better) 5
Estimation error, E0.95
Estimator efficiency gain
106 105 10
4
103 102 101
2 1
GPT-2 IU Gemma IU
100 2.0
1.5
1.0
⋆ Estimation error threshold, E0.95
0.5
0
10−3
10−4
10−5
10−6
GPT-2 IU Logit Gemma IU Logit
10−7
10−8
10−9
Event probability, μ
Figure 12. Token Presence estimation on GPT-2 Small and Gemma-2 at L = 10. Left: estimator efficiency gains as accuracy requirements tighten. Right: error with N = 128 samples across event probabilities. IU retains large estimator gains in both models. Plotted curves show geometric means over ten events per model in 10−8 ≤ µ < 10−7 on the left and ten events with neighboring rarity on ⋆ the right; shading shows ±0.5 standard deviations in log space. At E0.95 = 1, geometric-mean IU 7 versus naive-MC budgets are approximately 349 versus 8.73 × 10 trajectories on GPT-2 and 1,445 versus 1.35 × 108 on Gemma-2. Absolute budgets appear in Appendix Figure 20.
27
7.2
Case study: Compositional Token Presence
Compositional Token Presence extends the evaluation to events requiring multiple tokens. Reliably estimating events defined by token sets or orderings could support evaluation of more complex ⋆ failures. Tables 2 and 3 report results at E0.95 = 1, a factor-ten error tolerance met by 95% of estimates. IU successfully reduces sample requirements for all ten evaluated compositional events, achieving estimator efficiency gains up to 2,090× for Ordered Sequence events. Conjunctive events proved consistently more difficult, with maximum gains of 196×. Because Conjunctive events require all three target tokens to appear in any order anywhere in the trajectory, the proposal must maintain coverage over multiple valid orderings. This significantly complicates importance weight control compared to a fixed ordered sequence. Including all compute costs, IU still provides compute-weighted efficiency gains over naive Monte Carlo for the rarest events, though these gains are much smaller (reaching 35.3× and 11.3×, respectively). These more modest gains leave room for future work to improve surrogate design and adaptive regularization for multi-token conditions.
7.3
Case study: Profanity Classifier
Here a classifier assigns a profanity probability to a sequence of tokens, and the event occurs when f (y) > κ. Raising κ makes the event rarer. This follows the traditional rare-event formulation in which a score crossing a threshold defines the event. Across verified rarities in this setting, IU still achieved substantial gains over naive MC at both L = 50 and L = 10. At L = 10, IU has larger estimator efficiency gains across accuracy requirements and lower aggregate errors across rarity levels than IU Logit on both models (Figure 13). On Gemma-2, IU Logit falls below naive-MC estimator efficiency at stricter requirements, while IU retains a sampling advantage. On GPT-2 Small at L = 50, IU AS also loses its gains as requirements tighten, and both CE methods produce large errors and require more samples than naive Monte Carlo (Appendix Figure 28).
7.4
Computational cost
IU achieves geometric-mean compute-weighted efficiency gains above 800× over naive Monte Carlo at every displayed error threshold for GPT-2 Small Token Presence at L = 10, 30, and 50 (Figure 14). Each curve shows a geometric mean over ten events in 10−8 ≤ µ < 10−7 . We train every IU proposal for the same number of steps and measure training costs separately for each model and length (Appendix E). At lower accuracy requirements, training costs dominate the total compute budget because the fixed cost of training the proposal is much higher than drawing the few samples needed for estimation. When training dominates, differences in the number of required samples have little effect on IU’s overall computational cost, and the efficiency gains mostly reflect the changing cost of naive MC. For this reason, we defer the complete set of compute-weighted efficiency plots to Appendix K. However, as stricter accuracy requirements call for more samples, sampling costs increase and training accounts for a progressively smaller share of IU’s total cost. Consequently, the gap between estimator efficiency and compute-weighted efficiency gains narrows across the thresholds in Figure 14. On GPT-2 Small, IU’s compute-weighted efficiency gains increase with event rarity and remain
28
Target tokens the → was → and the → was → people a → the → king the → was → village the → was → connect
Rarity µ
MC samples ⋆ NMC
IU samples ⋆ NIU
Estimator efficiency gain
Compute-weighted efficiency gain
2.59 × 10−3 1.03 × 10−4 3.02 × 10−5 5.75 × 10−6 1.25 × 10−6
1,156 28,999 99,215 521,040 2,402,465
96 256 256 512 1,152
12× 113× 388× 1,020× 2,090×
0.0169× 0.423× 1.44× 7.55× 35.3×
Table 2. Ordered Sequence (SEQ) events on GPT-2 Small at L = 10, requiring factor-ten accuracy in 95% of estimates. Targets appear in order, but need not be contiguous. IU reduces sample requirements for every event; gains including training cost are largest for the rarest events. Target tokens {the, was, and} {the, was, people} {a, the, king} {the, was, village} {the, was, connect}
Rarity µ
MC samples ⋆ NMC
IU samples ⋆ NIU
Estimator efficiency gain
Compute-weighted efficiency gain
6.57 × 10−3 5.05 × 10−4 1.98 × 10−4 8.06 × 10−5 3.72 × 10−6
455 5,927 15,118 37,162 804,821
256 384 384 640 4,096
1.78× 15.4× 39.4× 58.1× 196×
0.00686× 0.0893× 0.228× 0.557× 11.3×
Table 3. Conjunctive (AND) events on GPT-2 Small at L = 10, requiring factor-ten accuracy in 95% of estimates. Every target appears, in any order. IU reduces sample requirements for every event; gains including training cost are largest for the rarest events. Profanity Classifier, GPT-2 and Gemma, L = 10
Estimation error at budget N = 128 (lower is better)
Efficiency gain across error thresholds (higher is better) 2
Estimation error, E0.95
Estimator efficiency gain
103
102
101
100
1
10−1 2.0
1.5
1.0
⋆ Estimation error threshold, E0.95
0.5
0
GPT-2 IU Gemma IU
100
10−1
10−2
GPT-2 IU Logit Gemma IU Logit
10−3
10−4
10−5
Event probability, μ
Figure 13. Profanity Classifier estimation on GPT-2 Small and Gemma-2 at L = 10. IU achieves larger estimator efficiency gains and lower errors than IU Logit in both models. Left: estimator gains as accuracy requirements tighten at classifier cutoff κ = 0.999. Right: geometric-mean error with N = 128 samples across event rarity. Each point combines 5 events with neighboring rarity; ⋆ shading shows ±0.5 standard deviations in log space. At E0.95 = 1, IU versus naive-MC budgets are 64 versus 34,858 trajectories on GPT-2 and 384 versus 47,311 on Gemma-2. Absolute budgets appear in Appendix Figure 21.
29
Token Presence, GPT-2, IU
Efficiency gain across error thresholds (higher is better) Estimator efficiency gain
Efficiency gain
L = 10
Compute-weighted efficiency gain L = 30
L = 50
105 103 101 2
1.5
1
0.5 2
1.5
1
0.5 2
⋆ Estimation error threshold, E0.95
1.5
1
0.5
Figure 14. IU estimator efficiency and compute-weighted efficiency gains on GPT-2 Small at L = 10, 30, and 50 (left to right). As accuracy requirements tighten, training accounts for less of the total cost and the gap narrows. Curves show geometric means over ten events per length in 10−8 ≤ µ < 10−7 ; ⋆ shading shows ±0.5 standard deviations in log space. At E0.95 = 1, the geometric-mean computeweighted efficiency gains over naive MC are approximately 1,750×, 1,660×, and 949×, respectively. Absolute sample and time budgets appear in Appendix Figure 22.
above 800× across the evaluated sequence lengths. These trends suggest that IU has the potential to remain useful as events become rarer and trajectories grow longer.
7.5
Qualitative examples
The following examples show high importance-weight trajectories from saved GPT-2 Small IU proposals for the four event types. IU model completions retain natural grammatical structure while satisfying the event conditions. This suggests that IU may help reveal ways a specified event can occur, in addition to estimating its probability. All samples include the prompt Once upon a time; L counts only generated tokens, and µ is the base-model event probability. Profanity is masked for display purposes. Token Presence. Contains Sorceress; L = 50, µ = 6.12 × 10−5 . Once upon a time, Ye Hao thought out how Sorceress Luo Yang should intervene and help him regain his immortality. Yukkun frowned and pouted. Even if it looked silly, laughing would be refreshing to her. It’s all simple, but Ye Hao couldn’t
Profanity Classifier. Classifier probability exceeds 0.999; L = 50, µ = 1.44 × 10−3 . Once upon a time he made such a f*****g joke out of his own blood... s**t was boiling. ’No one would f*****g tell you that I screwed him over,’ he said as he put his head together. ’Yeah, but nobody said that.
30
Conjunctive event. Contains the, was, and village in any order; L = 10, µ = 8.06 × 10−5 . Once upon a time, the village was fooling itself.
Ordered Sequence event. Contains the → was → village, with gaps allowed; L = 10, µ = 5.75 × 10−6 . Once upon a time the kingdom was abolished altogether and the village destroyed.
8
Discussion
We introduce Iterative Unalignment (IU), which searches over proposal distributions via weight-space perturbations to form an IS proposal for estimating rare-event probabilities under a fixed, potentially post-trained base model. While we demonstrate substantial gains in controlled environments, evaluating on complex real-world failures remains challenging due to the lack of verifiable groundtruth probabilities. Extending this evaluation requires benchmarks with more realistic failures and reference probabilities precise enough to assess estimator accuracy. Importance sampling and alignment. Efficient importance sampling seeks a proposal that makes the event more frequent while matching the base model’s conditional distribution p(· | E) within it. Learning such a proposal involves finding where the base model places probability mass in the rare region and preserving the relative probabilities of event trajectories. With event definitions that capture meaningful failures, this search has the potential to reveal the base model’s preferences among failure trajectories. These failure trajectories may help guide alignment efforts. Further developing IU. Effective rare-event estimation requires both frequent event samples and adequate coverage of the trajectories that contribute to the event probability. Greater flexibility can complicate this balance by allowing the proposal to concentrate on a narrow subset of trajectories. Yet restricting the proposal can also limit its ability to amplify the event while preserving the base model’s relative probabilities within it. Our results suggest that flexibility is useful for learning these changes across an autoregressive sequence. The sampling gains across all three event families demonstrate IU’s potential as a framework for rare-event estimation. The Gemma-2 and Compositional Token Presence results also suggest that controlling importance weights may be more difficult for larger models and more complex events. Better surrogate design and adaptive regularization schemes that control within-event distortion while allowing event frequency to increase are promising directions for improving this framework. Developing these components together may yield further gains in accuracy and efficiency. Aligning agentic models with safety constraints remains an urgent challenge. Although our current empirical findings reflect preliminary steps toward this goal, we hope the Iterative Unalignment framework offers a useful contribution.
Acknowledgements This work was supported by Coefficient Giving. 31
References [1] Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos, Matthieu Geist, and Olivier Bachem. On-policy distillation of language models: Learning from selfgenerated mistakes. In International Conference on Learning Representations, 2024. [2] Rico Angell, Raghav Singhal, Zachary Horvitz, Zhou Yu, Rajesh Ranganath, Kathleen McKeown, and He He. Estimating tail risks in language model output distributions. arXiv preprint arXiv:2604.22167, 2026. [3] Siu-Kui Au and James L Beck. Estimation of small failure probabilities in high dimensions by subset simulation. Probabilistic Engineering Mechanics, 16(4):263–277, 2001. [4] Tao Bai, Jinqi Luo, Jun Zhao, Bihan Wen, and Qian Wang. Recent advances in adversarial training for adversarial robustness. In Zhi-Hua Zhou, editor, Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI-21, pages 4312–4321. International Joint Conferences on Artificial Intelligence Organization, 8 2021. doi: 10.24963/ijcai.2021/591. URL https://doi.org/10.24963/ijcai.2021/591. Survey Track. [5] Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073, 2022. URL https://arxiv.org/abs/2212.08073. [6] Jose Blanchet and Henry Lam. Rare event simulation techniques. In Proceedings of the 2011 Winter Simulation Conference, 2011. URL https://www.columbia.edu/~khl2114/files/ BLWSC.pdf. [7] Zdravko I. Botev, Dirk P. Kroese, Reuven Y. Rubinstein, and Pierre L’Ecuyer. The cross-entropy method for optimization. In Venu Govindaraju and C. Radhakrishna Rao, editors, Handbook of Statistics: Machine Learning: Theory and Applications, volume 31, pages 35–59. Elsevier, 2013. doi: 10.1016/B978-0-444-53859-8.00003-5. [8] James Antonio Bucklew. Introduction to rare event simulation, volume 5. Springer, 2004. [9] Amanda Cao, Ivan Betancourt, Manish Rangan, and Yuqi Sun. Optimizing importance sampling methods for rare output estimation in language models. In AAAI Workshop on Trust and Control in Agentic AI, 2026. URL https://openreview.net/forum?id=Bp49EGSzto. [10] Stephen Casper, Lennart Schulze, Oam Patel, and Dylan Hadfield-Menell. Defending against unforeseen failure modes with latent adversarial training. Transactions on Machine Learning Research (TMLR), 2024. URL https://arxiv.org/abs/2403.05030. [11] Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J. Pappas, and Eric Wong. Jailbreaking black box large language models in twenty queries. 36, 2023. [12] Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences. In Advances in Neural Information Processing Systems (NeurIPS), volume 30, 2017. URL https://arxiv.org/abs/1706.03741.
32
[13] Pieter-Tjerk de Boer, Dirk P. Kroese, Shie Mannor, and Reuven Y. Rubinstein. A tutorial on the cross-entropy method. Annals of Operations Research, 134(1):19–67, 2005. doi: 10.1007/ s10479-005-5724-z. [14] Jake McAllister Dorman, Edward Gillman, Dominic C. Rose, Jamie F. Mair, and Juan P. Garrahan. Rare event analysis of large language models. arXiv preprint arXiv:2602.06791, 2026. ICML 2026 Oral. [15] Paul Glasserman. Monte Carlo Methods in Financial Engineering. Springer, New York, 2004. [16] Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adversarial examples. In International Conference on Learning Representations (ICLR), 2015. URL https://arxiv.org/abs/1412.6572. [17] Suvadeep Hajra, Palash Nandi, and Tanmoy Chakraborty. Exposing long-tail safety failures in large language models through efficient diverse response sampling. arXiv preprint arXiv:2603.14355, 2026. [18] Philip Heidelberger. Fast simulation of rare events in queuing and reliability models. ACM Transactions on Modeling and Computer Simulation (TOMACS), 5(1):43–85, 1995. [19] Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations (ICLR), 2022. URL https://openreview. net/forum?id=n1V4vR6Z9S. [20] Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R Narasimhan. SWE-bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations (ICLR), 2024. [21] Erik Jones, Meg Tong, Jesse Mu, Mohammed Mahfoud, Jan Leike, Roger Grosse, Jared Kaplan, William Fithian, Ethan Perez, and Mrinank Sharma. Forecasting rare language model behaviors. arXiv preprint arXiv:2502.16797, 2025. [22] Pierre L’Ecuyer, Jose H. Blanchet, Bruno Tuffin, and Peter W. Glynn. Asymptotic robustness of estimators in rare-event simulation. ACM Transactions on Modeling and Computer Simulation, 20(1):6:1–6:41, 2010. doi: 10.1145/1667072.1667078. [23] Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. Towards deep learning models resistant to adversarial attacks. In International Conference on Learning Representations (ICLR), 2018. URL https://arxiv.org/abs/1706.06083. [24] S. K. Morley, T. V. Brito, and D. T. Welling. Measures of model performance based on the log accuracy ratio. Space Weather, 16(1):69–88, 2018. doi: 10.1002/2017SW001669. [25] S. G. Patil, T. Zhang, X. Wang, and J. E. Gonzalez. Gorilla: Large language model connected with massive apis. arXiv preprint arXiv:2305.15334, 2023. [26] Alistair Reid, Simon O’Callaghan, Liam Carroll, and Tiberio Caetano. Risk analysis techniques for governed llm-based multi-agent systems. arXiv preprint arXiv:2508.05687, 2025. 33
[27] Gerardo Rubino, Bruno Tuffin, et al. Rare event simulation using Monte Carlo methods, volume 73. Wiley Online Library, 2009. [28] Reuven Y. Rubinstein and Dirk P. Kroese. The Cross-Entropy Method: A Unified Approach to Combinatorial Optimization, Monte-Carlo Simulation and Machine Learning. Springer Science & Business Media, New York, 2004. doi: 10.1007/978-1-4757-4321-0. [29] James S. Sadowsky and James A. Bucklew. On large deviations theory and asymptotically efficient Monte Carlo estimation. IEEE Transactions on Information Theory, 36(3):579–588, 1990. [30] T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom. Toolformer: Language models can teach themselves to use tools. arXiv preprint arXiv:2302.04761, 2023. [31] Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. "do anything now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models. arXiv preprint arXiv:2308.03825, 2023. [32] Abhay Sheshadri, Aidan Ewart, Phillip Huang Guo, Aengus Lynch, Cindy Wu, Vivek Hebbar, Henry Sleight, Asa Cooper Stickland, Ethan Perez, Dylan Hadfield-Menell, and Stephen Casper. Latent adversarial training improves robustness to persistent harmful behaviors in llms. Transactions on Machine Learning Research (TMLR), 2024. URL https://arxiv.org/abs/ 2407.15549. [33] David Siegmund. Importance sampling in the Monte Carlo study of sequential tests. The Annals of Statistics, 4(4):673–684, 1976. [34] Chris Snyder, Thomas Bengtsson, Peter Bickel, and Jeff Anderson. Obstacles to high-dimensional particle filtering. Monthly Weather Review, 136(12):4629–4640, December 2008. doi: 10.1175/ 2008MWR2529.1. [35] Hang Su, Jun Luo, Chang Liu, Xiao Yang, Yichi Zhang, Yinpeng Dong, and Jun Zhu. A survey on autonomy-induced security risks in large model-based agents. arXiv preprint arXiv:2506.23844, 2025. [36] Chris Tofallis. A better measure of relative prediction accuracy for model selection and model estimation. Journal of the Operational Research Society, 66(8):1352–1362, 2015. doi: 10.1057/jors.2014.103. [37] G. Wang, Y. Xie, Y. Jiang, A. Mandlekar, C. Xiao, Y. Zhu, L. Fan, and A. Anandkumar. Voyager: An open-ended embodied agent with large language models. arXiv preprint arXiv:2305.16291, 2023. [38] X. Wang, Y. Wang, Z. Liu, Y. Zhou, and G. Neubig. Executable code actions elicit better llm agents. arXiv preprint arXiv:2402.01030, 2024. [39] Stefan Webb, Tom Rainforth, Yee Whye Teh, and M. Pawan Kumar. A statistical approach to assessing neural network robustness. In International Conference on Learning Representations (ICLR), 2019. URL https://openreview.net/forum?id=S1xcx3C5FX. 34
[40] Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36, 2023. [41] Yotam Wolf, Noam Wies, Yoav Levine, and Amnon Shashua. Fundamental limitations of alignment in large language models. In International Conference on Machine Learning (ICML), 2024. [42] Gabriel Wu and Jacob Hilton. Estimating the probabilities of rare outputs in language models. In International Conference on Learning Representations, 2025. Spotlight. [43] Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh J Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, et al. Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments. Advances in Neural Information Processing Systems, 37:52040–52094, 2024. [44] John Yang, Akshat Gupta, Sagar Jain, Weishi Wang, and Karthik Narasimhan. Intercode: Standardizing and benchmarking interactive coding with execution feedback. In Advances in Neural Information Processing Systems (NeurIPS), 2023. [45] John Yang, Carlos E Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. Swe-agent: Agent-computer interfaces enable automated software engineering. Advances in Neural Information Processing Systems, 37:50528–50652, 2024. [46] Stephen Zhao, Rob Brekelmans, Alireza Makhzani, and Roger Baker Grosse. Probabilistic inference in language models via twisted sequential monte carlo. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research. PMLR, 2024. [47] Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrikson. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043, 2023.
35
A
Notation
Importance Sampling in LMs Symbol
Meaning
V, |V| u L ℓ y = (y1 , . . . , yL ) y<ℓ x θ pθ , p δ, δt qδ := pθ+δ Φ E µ µ b w(y) rδ , qδE q∗
Vocabulary and its size Generic vocabulary token Output sequence length Decoding position, ℓ ∈ {1, . . . , L} Output sequence (trajectory) in V L Prefix preceding position ℓ Input sequence; held fixed and suppressed from the notation Base-model parameters Base model; we write p := pθ since θ remains fixed Generic weight-space perturbation and its value at training iteration t Proposal model Binary event indicator, Φ : V L → {0, 1} Rare event {Φ(y) = 1} Target probability Ep [Φ(y)] Estimate of µ Trajectory importance weight p(y)/qδ (y) Proposal event probability qδ (E) and conditional distribution qδ (· | E) Zero-variance proposal p(y)Φ(y)/µ
Evaluation settings. The superscript names the event family. In expressions such as ΦTOK (y; v), y is the trajectory being evaluated, while the quantity after the semicolon specifies the event instance. The bare symbol ΦTOK denotes the corresponding family; the same convention applies to SδTOK . Event
Indicator
Token Presence Conjunctive (AND) Ordered Sequence (SEQ) Profanity Classifier
TOK
Φ (y; v) ΦAND (y; v) ΦSEQ (y; v) ΦPROF (y; κ)
Reference probability estimator b TOK
Φ (y; v) b AND (y; v) Φ b SEQ (y; v) Φ —
Reference probability
Differentiable surrogate
TOK
SδTOK (y; v) SδAND (y; v) SδSEQ (y; v) SδPROF (y; κ)
µ (v) µAND (v) µSEQ (v) µPROF (κ)
Symbol
Meaning
v v = (v1 , . . . , vm )
Target token for Token Presence Tuple of target tokens for Compositional Token Presence; m is the number of targets Profanity Classifier probability and its threshold; ΦPROF (y; κ) = I{f (y) > κ}
f, κ
Iterative Unalignment objective and training.
36
Symbol
Meaning
Sδ (y) Rδ (y) Jt λt , λ0 , λmin T, t B, i ηδ , η λ ESSrbatch , ESSrevent ρbatch , ρevent ESSrt , ρt ĥt , h⋆
Differentiable surrogate supplying the amplification signal Dense reverse-KL regularizer (10) Empirical IU objective at iteration t Regularization strength, its initial value, and its floor Total number of training iterations and the training iteration index Training batch size and batch sample index Learning rates for δ and for λ Full-batch and in-event effective sample size ratios Targets for the two effective sample size ratios Active effective sample size ratio and its target at iteration t Empirical event frequency and the phase-switch threshold
Evaluation protocol and metrics. Symbol
Meaning
N K, k
Estimation budget Number of independent repetitions and the repetition index, giving {b µk,N }K k=1 Quantile log error (16) Estimation error threshold: largest acceptable quantile log error Budget required to satisfy the error threshold (17) Exact Binomial naive-MC budget for the same threshold Per-trajectory wall-clock cost under MC and IS, and proposal training cost
E0.95 (N ; µ) ⋆ E0.95 ⋆ N ⋆ (E0.95 ; µ) ⋆ ⋆ NMC (E0.95 ; µ) costMC , costIS , costtrain
37
B
Additional experimental-setup diagnostic
Figure 15. The asymmetric penalty structure of rMSE leads to misleading comparisons in the rare-event regime. Left panel. The range of estimates µ bk versus true probability µ for tokens across six orders of magnitude in rarity. Each black dot represents an independent estimate, with error bars indicating variability across repetitions. The red dashed line denotes perfect estimation (b µ = µ). At µ ≈ 10−10 , the rare event (12) defined by the token TheNitrome (highlighted) shows the worst estimation quality, with estimates highly dispersed and systematically biased below the diagonal. Right panel. The improvement factor in rMSE relative to naive Monte Carlo. Paradoxically, we see the largest improvement in rMSE on this event. This occurs because estimators underestimate at this rarity, which does not get penalized much under rMSE’s asymmetry.
C
Efficiency gains across error thresholds
We bootstrap the stored estimates up to four times the evaluation budget, or 256,000 trajectories from 64,000 stored trajectories. If the target remains unmet, we add simulated naive-MC samples, since naive MC gives an unbiased estimate of the event probability. We simulate event counts from the Binomial distribution using the reference probability, retain the initial IS contributions, and count all IS and MC samples in the estimate and budget. This gives a conservative sketch when further IS sampling would outperform naive MC, since we assume no further IS advantage and retain the initial IS error and cost. Bootstrapping can also understate gains when the stored data miss rare event contributions. These estimates remain subject to bootstrap uncertainty and are not statistical lower bounds. The error-threshold axis stops at 0.5 because tighter targets can require more naive-MC samples than our compute budget allows us to validate.
D
Implementation and hyperparameters
Training and evaluation. For Token Presence, we use batches of 128 trajectories and a training budget of 300 steps per proposal. We then freeze the proposal and draw 500 evaluation batches,
38
totaling 64,000 trajectories. Appendix E reports the measured computational costs and accounts for the forward passes across events, methods, and sequence lengths. IU. Both GPT-2 Small and Gemma-2 use LoRA with rank r = 16 and scaling factor αLoRA = 32. The proposal learning rate is ηδ = 10−4 . We initialize the regularization strength at λ0 = 10, impose a floor of λmin = 0.1, and use dual learning rate ηλ = 10−2 . The controller targets population ESS ratio ρbatch = 0.005 while the event frequency is below h⋆ = 0.1. Once at least 10% of the batch satisfies the event, it targets event ESS ratio ρevent = 0.1. IU AS and IU Logit. The restricted variants use the same regularization settings as IU. Their proposal learning rates are 10−2 for IU AS and 1 for IU Logit. IU AS initializes each steering-vector coordinate to 0.01; IU Logit initializes its coordinates independently from a zero-mean Gaussian with standard deviation 0.01. CE AS and CE Logit. Both CE methods use a population of 128 steering vectors, elite fraction γelite = 0.3, and smoothing coefficient τCE = 0.5. The initial search standard deviation and its floor are both 0.1 for CE AS and 1 for CE Logit. The search mean is initialized to zero, and generation uses temperature 1. Evaluation fixes the steering vector at the learned search mean. After a warmup of 10 training steps, CE stops early if the event ESS ratio falls below 0.1 for two consecutive steps, each with at least two event hits. Appendix H.1 describes the CE formulations and their empirical behavior.
E
Computational costs Table 4. Wall-clock costs for Compute-Weighted Efficiency Gain, measured on an isolated NVIDIA A100-SXM4-80GB GPU. IU uses LoRA (r = 16, αLoRA = 32); training cost is for T = 300 steps at B = 128. Model L costMC costIS costtrain costIS /costMC (s / trajectory) (s) GPT-2 Small GPT-2 Small GPT-2 Small Gemma-2-2B-IT
E.1
10 30 50 10
0.0046 0.0138 0.0230 0.0501
0.0051 0.0153 0.0254 0.0569
213 640 1067 2718
1.10 1.10 1.10 1.14
Forward-pass accounting for Token Presence
We count each model forward pass on a batch of 128 sequences as one pass. GPT-2 has 60 events at each of L = 10, 30, 50, evaluated with five methods. Gemma-2 has 55 IU and 66 IU Logit runs, all at L = 10. This gives 60 · 3 · 5 + 55 + 66 = 1,021 runs, each with 300 training and 500 evaluation batches. Generating a batch of length L requires L forward passes. The generation budget is therefore Fgen = (300 + 500) 60 · 5 · (10 + 30 + 50) + (55 + 66) · 10 = 22,568,000.
39
Each batch also requires one proposal and one base-model scoring pass, giving 2 · (300 + 500) · 1,021 = 1,633,600 scoring passes. Reference estimation uses 1,024,000 sequences, or 8,000 batches, in each of the four model–length settings. Each batch requires L generation passes and one scoring pass, giving Fref = 8,000 (10 + 1) + (30 + 1) + (50 + 1) + (10 + 1) = 832,000. The total budget is 25,033,600 forward passes. Broken down, this is 22.57 million for generation, 1.63 million for scoring, and 0.83 million for reference estimation.
F
Comparison with score-guided MCMC
Dorman et al. [14] study rare-event estimation in language models by reconstructing the distribution of continuous scalar scores g(y) (such as output log probability and readability) under the base model p using score-biased MCMC and multistate Bennett acceptance ratio (MBAR) reweighting. While their primary focus is on continuous score distributions where exact tail reference probabilities are generally unavailable, their framework also supports estimating a target indicator Φ(y) using a correlated auxiliary scalar score to bias the MCMC transitions. This allows us to evaluate scoreguided MCMC on Token Presence events, where precise reference probabilities allow us to verify accuracy in the deep tail. We set the target to the binary event indicator Φ(y) and use the Token Presence surrogate as the biasing score, g(y) = SδTOK (y; v), evaluated with qδ = p. This score directly rewards prefixes that assign higher probability to the target token, providing an aligned search signal for the MCMC sampler. For each event, we run 10 chains across the positive bias schedule 0.1, 0.2, . . . , 1.0, with 1,219 transitions per chain and bias, 10% burn-in, and 6,090 direct samples. Including the 10 initial chain states, this gives 128,000 evaluated trajectories for one complete estimate. We retain rejected states and exclude a bias only when its Gelman–Rubin diagnostic exceeds 1.1, following the authors’ procedure. The schedule is fixed across events and receives no event-specific tuning. Dorman’s MCMC transition truncates each trajectory at a uniformly random output position and regenerates the remaining suffix. We conservatively charge each evaluated trajectory as only half a naive-MC inference. From Table 4, the charged cost is 128,000(0.5) costMC = 128,000(0.5)(0.0046) = 294.4 seconds. For comparison, the full cost of one N = 128 IU estimate (including proposal training) is costtrain + 128 costIS = 213 + 128(0.0051) = 213.65 seconds. Thus, in this experiment, Dorman’s charged budget is 294.4/213.65 = 1.38× the full cost of an N = 128 IU estimate. Table 5 reports event discovery and probability estimation. The event hit rate is the fraction of sampled trajectories satisfying the Token Presence event.
40
(a) Iterative Unalignment Event mere atten iminary sembly soDeliveryDate
Reference µ
Event hit rate
Final estimate
6.1594 × 10−4 1.1463 × 10−5 1.7840 × 10−6 1.1767 × 10−7 1.4414 × 10−8
0.620 0.167 0.311 0.345 0.264
6.1927 × 10−4 7.4793 × 10−6 2.4233 × 10−6 9.3520 × 10−8 1.2899 × 10−8
Reference µ
Event hit rate
Final estimate
−4
−3
6.0048 × 10−4 7.3987 × 10−6 0 0 0
(b) Dorman et al. Event mere atten iminary sembly soDeliveryDate
6.1594 × 10 1.1463 × 10−5 1.7840 × 10−6 1.1767 × 10−7 1.4414 × 10−8
1.04 × 10 6.67 × 10−5 0 0 0
Table 5. IU and score-guided MCMC on GPT-2 Small Token Presence events at L = 10. Under the compute accounting above, IU makes every event common under its proposal, while Dorman et al. yield zero event hits for events with reference probabilities below ∼ 10−5 .
These results illustrate the structural distinction between trajectory-space MCMC and global proposal optimization. Score-guided MCMC relies on local suffix mutations, which can struggle to discover isolated target modes in high-dimensional discrete sequence spaces when the event is exceedingly rare (µ < 10−5 ). In such regimes, MCMC chains may fail to record event hits despite satisfying standard internal diagnostics (such as Gelman–Rubin R̂ ≤ 1.1). In contrast, IU optimizes model parameters globally, steering autoregressive decoding dynamics across all prefix positions to reliably sample deep-tail events and provide accurate importance-sampling estimates.
G
Ground truth estimation
We establish the unbiasedness of the token-based reference estimators and report the relative standard errors of the reference probabilities used in our evaluation.
G.1
b TOK Unbiasedness of Φ
Let µTOK (v) = Ep [ΦTOK (y; v)]. Decomposing over the first hitting time of v gives the following mutually exclusive events. ΦTOK (y; v) =
L X
I{yℓ = v} I{v ∈ / y<ℓ }
(21)
ℓ=1
Conditioning on y<ℓ and using E[I{yℓ = v} | y<ℓ ] = p(v | y<ℓ ) gives " L # X b TOK (y; v)]. Ep [ΦTOK (y; v)] = Ep p(v | y<ℓ ) I{v ∈ / y<ℓ } = Ep [Φ ℓ=1
b TOK (y; v) are therefore unbiased estimators of µTOK (v). Both ΦTOK (y; v) and Φ 41
(22)
G.2
Compositional Token Presence first-completion estimators
The same first-completion argument gives unbiased estimators from next-token probabilities for the Compositional Token Presence events in §6.1. For the Conjunctive event, let v = (v1 , . . . , vm ) be a tuple of distinct target tokens and let Mℓ = {v1 , . . . , vm } \ {y1 , . . . , yℓ−1 } be the targets still missing immediately before position ℓ. The event is first completed at ℓ exactly when one target remains and that target is emitted. Its reference estimator is therefore b AND (y; v) = Φ
L X
I{|Mℓ | = 1}
X
p(u | y<ℓ ).
(23)
u∈Mℓ
ℓ=1
Conditioning on y<ℓ replaces the indicator that the final missing token is sampled with its next-token probability. Because the possible first-completion positions are mutually exclusive, b AND (y; v)] = Ep [ΦAND (y; v)]. Ep [Φ For the Ordered Sequence event, define the prefix state cℓ = max{d ∈ {0, . . . , m} : (v1 , . . . , vd ) is a subsequence of y<ℓ } . The event is first completed at position ℓ exactly when cℓ = m − 1 and yℓ = vm . Thus, b SEQ (y; v) = Φ
L X
I{cℓ = m − 1}p(vm | y<ℓ ).
(24)
ℓ=1
Applying the tower property at every position and summing the mutually exclusive first-completion events gives b SEQ (y; v)] = Ep [ΦSEQ (y; v)]. Ep [Φ b over ordinary rollouts from the In both cases, the reference probability is the sample mean of Φ frozen base model. The estimator integrates out the final token draw that completes the event, yielding a much denser reference signal than the raw indicator.
G.3
Accuracy of the reference probabilities
For Ngt independent base-model rollouts, let Zi be the unbiased first-completion estimate for tokenbased events or the binary indicator for the Profanity Classifier. We compute the reference probability and its estimated relative standard error as Ngt
µ bgt =
1 X Zi , Ngt
d = p sZ , RSE Ngt µ bgt
i=1
where sZpis the sample standard deviation of the Zi . For naive MC, the population relative standard error is (1 − µ)/(Ngt µ), so 10% precision at µ = 10−9 requires approximately 1011 rollouts. We use 1,024,000 reference rollouts per model and length for Token Presence and 512,000 for compositional and profanity events. These budgets are fixed across events; all estimated relative standard errors are below 10%. Table 6 reports the rarest event in each listed setting.
42
Table 6. Reference probability and relative standard error for the rarest reported event in each listed setting. All relative standard errors are below 10%. Setting GPT-2 Small token, L = 50 Gemma-2B token, L = 10 Ordered Sequence event, L = 10 Conjunctive event, L = 10 GPT-2 Small profanity, L = 50 Gemma-2B profanity, L = 10
Rarest reported event ThumbnailImage meio the → was → connect {the, was, connect} classifier threshold 0.999 classifier threshold 0.95
H
Related methods
H.1
Cross-entropy baseline selection
Reference Reference Relative samples probability SE 1,024,000 3.97 × 10−9 0.74% 1,024,000 1.08 × 10−9 1.32% 512,000 1.25 × 10−6 7.16% 512,000 3.72 × 10−6 9.17% 512,000 1.44 × 10−3 3.68% 512,000 7.50 × 10−4 5.10%
We evaluated two CE formulations: importance-weighted CE for rare-event estimation and unweighted CE for optimization. Both rank sequences using Gδ (y) = Sδ (y) for Token Presence and Compositional Token Presence, and Gδ (y) = log f (y)/(1−f (y)) for the Profanity Classifier, the logit of its predicted probability. We write Topa for the indices of the top a fraction of scores. Importance-weighted CE. Classical rare-event CE maximizes elite-sequence likelihood with weights proportional to p(y)/qδt (y) [13, 28]. We approximately maximize this objective by gradient ascent, holding the sampled sequences, elite membership, and weights fixed (Algorithm 2). These updates did not amplify rare events enough for reliable estimation; neither activation nor logit steering observed the event in Table 7. Algorithm 2 Importance-weighted CE with gradient updates Inputs: p, qδ , Φ, Gδ ; T, B, η; ρ, h⋆ . 1: δ1 ← 0 2: for t = 1, 2, . . . , T do 3: 4:
iid
Sample y1 , . . . , yB ∼ qδt P ĥt ← B −1 B i=1 Φ(yi )
▷ Fixed sampled trajectories ▷ Update event proportion
if ĥt ≥ h⋆ then 6: return qδt 7: end if 8: It ← Topρ {Gδt (yi )}B i=1 P 9: wi ← p(yi )/qδt (yi ), w̄i ← wi / j∈It wj for i ∈ It P 10: Ft (δ) ← i∈It w̄i log qδ (yi ) 11: δt+1 ← δt + η∇δ Ft (δt ) 12: end for 13: return qδT +1 5:
43
▷ Event-proportion target reached ▷ Select elite trajectories ▷ Fixed elites and weights ▷ Gradient ascent ▷ Proposal for IS estimation
Parameterization Activation Logit
Iterations
Training trajectories
Event hits
44 96
45,056 98,304
0 0
Table 7. Importance-weighted CE with gradient updates for ActionCode on GPT-2 Small at L = 30 (µ = 2.812 × 10−9 ). Counts cover completed training rounds.
Unweighted CE optimization. In our unweighted likelihood trials, gradient updates concentrated the proposal too strongly on the observed elite sequences, resulting in severe underestimation. The standard CE optimization search performed better empirically: it samples steering vectors, ranks them by the scores of their generated sequences, and moves the search mean toward the unweighted average of the elite vectors [28, 7] (Algorithm 3). We retain this update for activation steering (CE AS) and logit steering (CE Logit). At L = 10, it achieves substantial estimator and compute-weighted efficiency gains over naive Monte Carlo at some accuracy thresholds (Figure 24). Algorithm 3 Unweighted CE optimization with activation/logit tilting 2 ; τ Inputs: p, qδ , Gδ ; T, B; γelite , σCE CE ∈ (0, 1]. 1: m1 ← 0 2: for t = 1, 2, . . . , T do iid
2 I) Sample δ (1) , . . . , δ (B) ∼ N (mt , σCE 4: Sample yi ∼ qδ(i) for i = 1, . . . , B 5: It ← Topγelite {Gδ(i) (yi )}B i=1 P 6: mt,elite ← |It |−1 i∈It δ (i) 7: mt+1 ← (1 − τCE )mt + τCE mt,elite 8: end for 9: return mT +1
3:
H.2
▷ Steering vectors ▷ One sequence per vector ▷ Select elite perturbations ▷ Elite mean ▷ Update search mean ▷ Learned perturbation
Twisted sequential Monte Carlo
Twisted sequential Monte Carlo (TSMC) uses intermediate scores to guide particles toward promising continuations. These scores can be supplied by learned twist functions [46]. In our implementation, the surrogate guides particle resampling through look-ahead evaluations, without updating the LM’s weights. IU instead uses surrogate gradients to train a proposal that changes the token conditionals throughout generation. We evaluated a TSMC variant using our surrogate Sδ (y) as the resampling score with a two-step look-ahead, averaged across 20 look-ahead particles. Across 10 token events with probabilities between 10−6 and 10−9 , this variant yielded an average efficiency ratio of 0.87× relative to naive Monte Carlo. It did not consistently improve upon naive Monte Carlo, so we did not retain it in the main baseline comparisons. In these runs, repeated resampling concentrated particles onto a small number of high-scoring prefixes, reducing diversity among downstream trajectories. This behavior highlights the need to preserve diverse event paths while directing sampling toward the event, which motivates the combination of amplification and regularization in IU. 44
I
Additional considerations in methodology
I.1
Surrogate coverage
We examine surrogate mismatch using a fixed GPT-2 Small proposal trained to amplify the token ␣connect. We evaluate this proposal on five events. The central event requires this token alone. Adding alternative tokens through an OR condition expands the event beyond the condition targeted by the surrogate. Adding required tokens through an AND condition instead makes ␣connect a necessary but insufficient condition for the event. This latter case provides a simple instance of the necessary-condition surrogate discussed in Section 5.1. Figure 16 shows the lowest error when the event matches the token targeted by the surrogate. Error increases as either alternative or additional required tokens are introduced. These preliminary results suggest that a surrogate can remain useful when it targets a broader condition than the event, with a loss of estimation accuracy as the mismatch grows. A systematic study of this tradeoff for more complex events remains open. 3.5
undercoverage 3.1
overcoverage
3.0
E0.95
2.5 2.0
2.0
1.9
1.5
1.4 0.9
1.0 0.5 0.0
" connect" OR "uss" OR "aed"
" connect" OR "uss"
" connect"
" connect" AND " the"
Tokens in Φ
" connect" AND " the" AND " was"
Figure 16. Estimation error under surrogate mismatch. A fixed proposal trained to amplify ␣connect is evaluated on the five token-presence events shown. OR conditions (left) include event trajectories without this token; AND conditions (right) require additional tokens that the surrogate does not target. The central event matches the surrogate’s target. Lower error is better. Undercoverage and overcoverage refer to the surrogate’s target condition relative to the event.
I.2
Alternative regularizers
We also considered three alternatives to the regularizer in Eq. (10). Joint forward KL. For the ActionCode event, regularization based on base-model rollouts failed to control concentration within the event (Figure 17, left). This alternative uses the joint forward
45
divergence KL(p∥qδ ) = Ey∼p
" L # X p(y) log = Ey∼p KL(p(· | y<ℓ )∥qδ (· | y<ℓ )) . qδ (y)
(25)
ℓ=1
At practical training budgets, rollouts from p almost exclusively contain non-event trajectories. A batch of B such rollouts misses the event with probability (1 − µ)B , so the penalty receives little direct feedback about distortion within E. This causes the proposal to perform poorly. Restricting regularization to the event. Gating the penalty by the event indicator, using Φ(y)Rδ (y) on proposal rollouts, also degraded IU’s performance. When a batch contains no event hits, this penalty is zero on every trajectory, so the surrogate update receives no regularization in early steps. Applying Rδ to every rollout constrains these early updates before event hits become frequent (Figure 17, center). On-policy forward KL. Finally, we kept proposal rollouts and reversed the divergence between the next-token distributions, " L # X Don (δ) = Ey∼qδ KL(p(· | y<ℓ )∥qδ (· | y<ℓ )) . (26) ℓ=1
The outer expectation follows qδ , while the inner KL weights tokens under p. This is the forward-KL objective used in on-policy distillation, where a student learns from teacher feedback on its own generated sequences [1]. It differs from joint forward KL in Eq. (25), whose prefixes follow p. Both on-policy directions provide feedback across the full vocabulary at each visited prefix. Neither directly compares conditionals at unvisited prefixes. The two variants gave similar pooled probability estimates in this comparison (0.98µ for reverse KL and 0.94µ for forward KL), despite different event frequencies and weight distributions (Figure 17, right). The preferred direction remains problem specific, consistent with the findings of Agarwal et al. [1].
Figure 17. Alternative regularizers, compared through in-event importance weights p/qδ at L = 10. Red denotes our choice in each panel. Left: joint reverse KL on proposal rollouts and joint forward KL on base-model rollouts for ActionCode. Center: regularization on every proposal rollout and regularization gated by Φ. Right: reverse and forward next-token KL, both evaluated on proposal rollouts. The center and right panels use the Profanity Classifier. The green line marks µ, the weight under the zero-variance proposal; colored dashed lines mark each proposal’s empirical mean in-event weight.
46
I.3
Per-step chi-square and importance-weight variance
A related alternative to the reverse-KL penalty is an unweighted sum of per-step chi-square divergences: " L # X 2 Dχ (q) := EY∼q χ (pℓ ∥qℓ ) . (27) ℓ=1
P 2 2 The exact joint chi-square, χ2 (p∥q) = Eq [ L ℓ=1 Wℓ−1 χ (pℓ ∥qℓ )], scales these local terms by the 2 importance weights of the prefixes, Wℓ−1 . Since these weights multiply out over the sampled tokens, they can vary sharply between trajectories and introduce severe instability into the penalty itself. By dropping the prefix weights, the unweighted sum avoids this compounding variance while retaining dense optimization feedback at every visited prefix. This unweighted surrogate remains sound for controlling the proposal divergence. With a finite vocabulary V and fixed generation length L, the unweighted penalty Dχ (q) is zero if and only if q = p. More generally, for proposals qn satisfying the same support assumption, driving the unweighted penalty to zero guarantees convergence of the true joint divergence: Dχ (qn ) −→ 0
I.4
=⇒
χ2 (p∥qn ) −→ 0.
(28)
Further fine-tuning the proposal
While IU employs a surrogate to provide a continuous gradient for the base proposal distribution, it is natural to then ask whether the proposal can be further improved after the initial IU optimization. Once a high-coverage proposal qδ is established, it triggers the rare event frequently enough that the exact indicator Φ(y) provides a usable signal. At this point, one could discard the surrogate and fine-tune the proposal further using other techniques. Exploring these subsequent refinement steps is an interesting direction, but it lies outside the scope of this work. We view such fine-tuning as naturally building upon the initial high-coverage proposal that IU provides.
I.5
Input context and autoregressive output length
Input context and generated output play different roles in IU. Conditional on a fixed input x, the trajectory importance weight is a product of L token-level likelihood ratios, one for each generated token. Increasing the output length therefore adds stochastic decoding steps over which discrepancies between the proposal and base model can accumulate. A longer fixed input supplies conditioning information but does not itself add factors to this product. The output length thus determines the size of the sequence space over which IU constructs its proposal.
47
J
Absolute budgets for the main-text efficiency plots
Figures 18–22 show the absolute budgets underlying each main-text efficiency figure, using the same events, methods, error thresholds, and geometric-mean summaries. Lower budgets are better. Dashed or dotted curves show naive MC; in model comparisons, their colors match the corresponding model and method. Shading shows ±0.5 standard deviations in log space across events, as in the main text. Required sampling budgets include any MC continuation described in Appendix C.
Required samples, N ⋆
Token Presence, GPT-2, L = 50
108 107 106 105 104 103 102 2.0
1.5
1.0
⋆ Estimation error threshold, E0.95
IU IU AS IU Logit
0.5
CE AS CE Logit Naive MC
Figure 18. Absolute sampling budgets for the Token Presence method comparison in Figure 10 (left), on GPT-2 Small at L = 50. Each curve summarizes the same ten events in 10−8 ≤ µ < 10−7 . ⋆ At E0.95 = 1, geometric-mean IU and naive-MC budgets are approximately 1,103 and 9.24 × 107 trajectories, respectively.
48
Token Presence, GPT-2, IU
10−9–10−8 10−8–10−7
Required samples, N ⋆
L = 10
10−7–10−6 10−6–10−5
10−5–10−4 10−4–10−3
L = 30
IU Naive MC L = 50
108 106 104 102 2.0
1.5
1.0
0.5 2.0
1.5
1.0
0.5 2.0
⋆ Estimation error threshold, E0.95
1.5
1.0
0.5
Figure 19. Absolute sampling budgets for the sequence-length comparison in Figure 11, with L = 10, 30, and 50 from left to right. Each color denotes a rarity band with ten events; solid curves show ⋆ IU and dashed curves show naive MC. At E0.95 = 1 in 10−8 ≤ µ < 10−7 , geometric-mean IU versus naive-MC budgets are approximately 349 versus 8.73 × 107 , 255 versus 7.73 × 107 , and 1,103 versus 9.24 × 107 trajectories, respectively.
49
Required samples, N ⋆
Token Presence, GPT-2 and Gemma, L = 10
108 107 106 105 104 103 102 2.0
1.5
1.0
⋆ Estimation error threshold, E0.95
GPT-2 IU Gemma IU GPT-2 IU Logit
0.5
Gemma IU Logit Naive MC
Figure 20. Absolute sampling budgets for the Token Presence model comparison in Figure 12 (left), ⋆ at L = 10. Each curve summarizes the same ten events per model in 10−8 ≤ µ < 10−7 . At E0.95 = 1, 7 geometric-mean IU versus naive-MC budgets are approximately 349 versus 8.73 × 10 trajectories on GPT-2 and 1,445 versus 1.35 × 108 on Gemma-2.
50
Profanity Classifier, GPT-2 and Gemma, L = 10
Required samples, N ⋆
106 105 104 103 102 101
2.0
1.5
1.0
⋆ Estimation error threshold, E0.95
GPT-2 IU Gemma IU GPT-2 IU Logit
0.5
Gemma IU Logit Naive MC
Figure 21. Absolute sampling budgets for the Profanity Classifier model comparison in Figure 13 ⋆ (left), at L = 10 and κ = 0.999. At E0.95 = 1, IU versus naive-MC budgets are 64 versus 34,858 trajectories on GPT-2 and 384 versus 47,311 on Gemma-2.
51
Token Presence, GPT-2, IU
IU
Required samples, N ⋆
L = 10
L = 30
L = 50
107 105 103
2.0
1.5
1.0
0.5 2.0
L = 10
Required GPU time (s)
Naive MC
1.5
1.0
0.5 2.0
L = 30
1.5
1.0
0.5
1.0
0.5
L = 50
106 105 104 103 2.0
1.5
1.0
0.5 2.0
1.5
1.0
0.5 2.0
⋆ Estimation error threshold, E0.95
1.5
Figure 22. Absolute budgets underlying the compute-weighted efficiency comparison in Figure 14, on GPT-2 Small at L = 10, 30, and 50 from left to right. Top: required trajectories. Bottom: estimated GPU time including IU training. Each curve summarizes the same ten events per length in ⋆ 10−8 ≤ µ < 10−7 . At E0.95 = 1, geometric-mean IU versus naive-MC times are approximately 230 versus 4.02 × 105 , 645 versus 1.07 × 106 , and 2,235 versus 2.12 × 106 seconds, respectively.
52
K
Full-detail result plots
Token-rarity comparisons use ten randomly selected tokens per rarity decade. In the main result figures, markers show geometric means across events and shading spans ±0.5 standard deviations around the mean in log space. ±0.5 was chosen for visual separation. The main text uses summaries for readability. Each figure here gives four full-detail views for the corresponding setting. Upper left: empirical estimate quantiles for IU only N = 128, including both models in model comparisons. The outer markers show the 2.5% and 97.5% quantiles, inner bars show the interquartile range, plus markers show medians, and diamonds show per-event arithmetic means. Upper right: quantile log error E0.95 against event probability at N = 128, with individual event points and per-decade geometric-mean lines and spread bands. Lower left: Estimator Efficiency Gain across error thresholds. Lower right: Compute-Weighted Efficiency Gain across the same thresholds, including proposal training and sampling costs from Appendix E. The error and frontier panels include all methods available in the setting. Efficiency frontier panels retain every selected event’s curve and markers beneath the geometricmean lines. As in the main figures, token frontiers use the ten matched events per model and method in 10−8 ≤ µ < 10−7 , and profanity frontiers use the classifier cutoff κ = 0.999. Error bands span ±0.75 standard deviations in log space. Dashed horizontal lines in the frontier panels indicate parity with naive Monte Carlo.
53
Token Presence, GPT-2, L = 50
Token Presence, GPT-2, L = 50
102
Estimation error, E0.95
Estimated probability, μ̂
10
Estimation error at budget N = 128 (lower is better)
Empirical estimate quantiles (IU)
−2
10−4
10−6
10−8
10−10
101
100
0 10−2 10−3 10−4 10−5 10−6 10−7 10−8 10−9
10−3
105
3
101
2.0
1.5
1.0
⋆ Estimation error threshold, E0.95
IU IU AS IU Logit
10−7
10−8
10−9
CE AS CE Logit
Token Presence, GPT-2, L = 50
Efficiency gain across error thresholds (higher is better)
Compute-weighted efficiency gain
Estimator efficiency gain
Token Presence, GPT-2, L = 50
10
10−6
IU IU AS IU Logit
Mean estimate
IU (N=128)
107
10−5
Event probability, μ
Event probability, μ truth (μ ̂ = μ)
10−4
0.5
CE AS CE Logit
Compute-weighted efficiency gain across error thresholds (higher is better) 103 102 101 100 2.0
1.5
1.0
⋆ Estimation error threshold, E0.95
IU IU AS IU Logit
CE AS CE Logit
Figure 23. Full-detail Token Presence rarity results on GPT-2 Small at L = 50: IU empirical quantiles (N = 128), all-method error at N = 128, and all-method estimator and compute-weighted efficiency frontiers.
54
0.5
Token Presence, GPT-2, L = 10
Token Presence, GPT-2, L = 10
Estimation error at budget N = 128 (lower is better)
10−4
Estimation error, E0.95
Estimated probability, μ̂
Empirical estimate quantiles (IU)
10−6
10
−8
10−10
101
100
0
10−2
10−4
10−6
10−8
10−10
10−3
105
3
101
2.0
1.5
1.0
⋆ Estimation error threshold, E0.95
IU IU AS IU Logit
10−7
10−8
10−9
CE AS CE Logit
Token Presence, GPT-2, L = 10
Efficiency gain across error thresholds (higher is better)
Compute-weighted efficiency gain
Estimator efficiency gain
Token Presence, GPT-2, L = 10
10
10−6
IU IU AS IU Logit
Mean estimate
IU (N=128)
107
10−5
Event probability, μ
Event probability, μ truth (μ ̂ = μ)
10−4
0.5
CE AS CE Logit
Compute-weighted efficiency gain across error thresholds (higher is better) 103 102 101 100 2.0
1.5
1.0
⋆ Estimation error threshold, E0.95
IU IU AS IU Logit
CE AS CE Logit
Figure 24. Full-detail Token Presence results on GPT-2 Small at L = 10: IU empirical quantiles (N = 128), all-method error at N = 128, and all-method estimator and compute-weighted efficiency frontiers.
55
0.5
Token Presence, GPT-2, L = 30
Token Presence, GPT-2, L = 30
Estimation error at budget N = 128 (lower is better)
102
Estimation error, E0.95
Estimated probability, μ̂
Empirical estimate quantiles (IU)
10−4
10−6
10−8
10−10
101
100
0 10−3
10−4
10−5
10−6
10−7
10−8
10−9
10−3
Token Presence, GPT-2, L = 30
Compute-weighted efficiency gain
Estimator efficiency gain
107 105 103
101
1.5
1.0
IU IU AS IU Logit
10−7
10−8
10−9
CE AS CE Logit
Token Presence, GPT-2, L = 30
Efficiency gain across error thresholds (higher is better)
⋆ Estimation error threshold, E0.95
10−6
IU IU AS IU Logit
Mean estimate
IU (N=128)
2.0
10−5
Event probability, μ
Event probability, μ truth (μ ̂ = μ)
10−4
0.5
CE AS CE Logit
Compute-weighted efficiency gain across error thresholds (higher is better) 103 102 101 100 2.0
1.5
1.0
⋆ Estimation error threshold, E0.95
IU IU AS IU Logit
CE AS CE Logit
Figure 25. Full-detail Token Presence results on GPT-2 Small at L = 30: IU empirical quantiles (N = 128), all-method error at N = 128, and all-method estimator and compute-weighted efficiency frontiers.
56
0.5
Token Presence, GPT-2, L = 50
Token Presence, GPT-2, L = 50
102
Estimation error, E0.95
Estimated probability, μ̂
10
Estimation error at budget N = 128 (lower is better)
Empirical estimate quantiles (IU)
−2
10−4
10
−6
10−8
10−10
101
100
0 10−2 10−3 10−4 10−5 10−6 10−7 10−8 10−9
10−3
105
3
101
2.0
1.5
1.0
⋆ Estimation error threshold, E0.95
IU IU AS IU Logit
10−7
10−8
10−9
CE AS CE Logit
Token Presence, GPT-2, L = 50
Efficiency gain across error thresholds (higher is better)
Compute-weighted efficiency gain
Estimator efficiency gain
Token Presence, GPT-2, L = 50
10
10−6
IU IU AS IU Logit
Mean estimate
IU (N=128)
107
10−5
Event probability, μ
Event probability, μ truth (μ ̂ = μ)
10−4
0.5
CE AS CE Logit
Compute-weighted efficiency gain across error thresholds (higher is better) 103 102 101 100 2.0
1.5
1.0
⋆ Estimation error threshold, E0.95
IU IU AS IU Logit
CE AS CE Logit
Figure 26. Full-detail Token Presence results on GPT-2 Small at L = 50: IU empirical quantiles (N = 128), all-method error at N = 128, and all-method estimator and compute-weighted efficiency frontiers.
57
0.5
Token Presence, L = 10
Token Presence, L = 10
Estimation error at budget N = 128 (lower is better) 101
10−4
Estimation error, E0.95
Estimated probability, μ̂
Empirical estimate quantiles (IU)
10−6 10−8 10
−10
10−12
100
10−14 10−2
10−4
10−6
10−8
0
10−10
Event probability, μ truth (μ ̂ = μ) GPT-2 IU (N=128)
10−3
Gemma IU (N=128) Mean estimate
GPT-2 IU Gemma IU
Compute-weighted efficiency gain
Estimator efficiency gain
105 104 3
102 101 100 2.0
1.5
1.0
⋆ Estimation error threshold, E0.95
GPT-2 IU Gemma IU
10−6
10−7
10−8
10−9
GPT-2 IU Logit Gemma IU Logit
Token Presence, L = 10
Efficiency gain across error thresholds (higher is better)
106
10
10−5
Event probability, μ
Token Presence, L = 10
107
10−4
0.5
GPT-2 IU Logit Gemma IU Logit
Compute-weighted efficiency gain across error thresholds (higher is better) 103 102 101 100 2.0
1.5
1.0
⋆ Estimation error threshold, E0.95
GPT-2 IU Gemma IU
0.5
GPT-2 IU Logit Gemma IU Logit
Figure 27. Full-detail Token Presence model comparison on GPT-2 Small and Gemma-2 at L = 10. Empirical quantiles show IU for both models at N = 128; error at N = 128 and the two efficiency frontiers show both IU and IU Logit for each model.
58
Profanity Classifier, GPT-2, L = 50
Profanity Classifier, GPT-2, L = 50
Empirical estimate quantiles (IU)
Estimation error at budget N = 128 (lower is better) Estimation error, E0.95
Estimated probability, μ̂
10−1 10−2 10−2 10−2 10−3 10−4
101
100
0 10−2
10−1
10−3
IU IU AS IU Logit
Mean estimate
IU (N=128) Profanity Classifier, GPT-2, L = 50
Compute-weighted efficiency gain
Estimator efficiency gain
102 101 100 10−1 10−2 1.5
1.0
⋆ Estimation error threshold, E0.95
IU IU AS IU Logit
CE AS CE Logit
Profanity Classifier, GPT-2, L = 50
Efficiency gain across error thresholds (higher is better)
2.0
10−3
Event probability, μ
Event probability, μ truth (μ ̂ = μ)
10−2
10−1
0.5
CE AS CE Logit
Compute-weighted efficiency gain across error thresholds (higher is better) 100
10−1
10−2
2.0
1.5
1.0
⋆ Estimation error threshold, E0.95
IU IU AS IU Logit
0.5
CE AS CE Logit
Figure 28. Full-detail Profanity Classifier results on GPT-2 Small at L = 50: IU empirical quantiles, all-method error at N = 128, and all-method estimator and compute-weighted efficiency frontiers at classifier cutoff 0.999.
59
Profanity Classifier, L = 10
Profanity Classifier, L = 10
Estimation error at budget N = 128 (lower is better)
Empirical estimate quantiles (IU) 10−1
Estimation error, E0.95
Estimated probability, μ̂
100
10−2 10−3 10−4 10−5 10−6
100
10−7 100
10−1
10−2
10−3
10−4
0
10−5
Event probability, μ truth (μ ̂ = μ) GPT-2 IU (N=128)
100
Gemma IU (N=64) Mean estimate
GPT-2 IU Gemma IU
Profanity Classifier, L = 10
Compute-weighted efficiency gain
Estimator efficiency gain
102 101 100 10−1 1.5
1.0
GPT-2 IU Gemma IU
10−3
10−4
10−5
GPT-2 IU Logit Gemma IU Logit
Profanity Classifier, L = 10
103
⋆ Estimation error threshold, E0.95
10−2
Event probability, μ
Efficiency gain across error thresholds (higher is better)
2.0
10−1
0.5
GPT-2 IU Logit Gemma IU Logit
Compute-weighted efficiency gain across error thresholds (higher is better) 100
10−1
2.0
1.5
1.0
⋆ Estimation error threshold, E0.95
GPT-2 IU Gemma IU
GPT-2 IU Logit Gemma IU Logit
Figure 29. Full-detail Profanity Classifier model comparison on GPT-2 Small and Gemma-2 at L = 10. Empirical quantiles show IU for both models at budget 128. Efficiency frontiers use classifier cutoff 0.999.
60
0.5
L
All experimental tokens and events
This section lists all 321 distinct model–length–event combinations in the experimental inventory: 246 Token Presence events, 65 Profanity Classifier events, 5 Conjunctive events, and 5 Ordered Sequence events. The listed µ values are the reference probabilities used by the experiments. Lengths L count generated tokens after the prompt. GPT-2 Small uses Once upon a time; Gemma-2 uses Tell me a story with its chat template. Table 8. All 60 Token Presence events for GPT-2 Small at L = 10 and their reference probabilities. Read the left list first, then the right; events are ordered from rarer to more common. Token ActionCode ␣\u30b5\u30fc\u30c6\u30a3 reportprint ␣RandomRedditor ␣externalToEVA ␣externalTo oreAndOnline ␣antidepress ␣teasp ÍÍ \u30b4\u30f3 soDeliveryDate soType natureconservancy \u30fc\u30c6 ␣pestic ␣unfocusedRange \u9f8d\u559a\u58eb ItemImage Thumbnail sembly ␣SolidGoldMagikarp Delivery zbek Database ␣BELOW 797 767 ␣978 ␣Ukrain
Ground truth µ
Token
1.151 × 10−9
iminary 388 ricia à employment html ␣SELECT article Si ␣Aside atten ␣indicative ␣Harbaugh ␣Ci ␣Acad ␣definitions ␣ingredient ␣loads ␣boost ␣Tomb aed uss ␣connect ␣fortunate ria ␣Arch ␣Let ␣promise ␣desired ␣mere
1.41 × 10−9 1.41 × 10−9 1.413 × 10−9 1.413 × 10−9 1.445 × 10−9 1.463 × 10−9 1.607 × 10−9 3.942 × 10−9 5.2 × 10−9 1.4403 × 10−8 1.4414 × 10−8 1.4863 × 10−8 2.0748 × 10−8 3.8521 × 10−8 4.0529 × 10−8 4.43 × 10−8 7.4453 × 10−8 7.6178 × 10−8 8.9868 × 10−8 1.17665 × 10−7 3.50913 × 10−7 4.43993 × 10−7 5.28826 × 10−7 6.0715 × 10−7 6.77883 × 10−7 6.97073 × 10−7 8.14367 × 10−7 9.14544 × 10−7 9.74397 × 10−7
61
Ground truth µ 1.783951 × 10−6 1.814818 × 10−6 1.85303 × 10−6 2.314821 × 10−6 2.863334 × 10−6 3.451055 × 10−6 4.080043 × 10−6 4.326181 × 10−6 4.56803 × 10−6 9.890819 × 10−6 1.1462555 × 10−5 1.2699662 × 10−5 1.4530803 × 10−5 2.11075 × 10−5 2.5733549 × 10−5 2.808293 × 10−5 3.4664296 × 10−5 4.3123397 × 10−5 5.2364278 × 10−5 7.9511388 × 10−5 1.25505903 × 10−4 1.44590405 × 10−4 1.58662006 × 10−4 2.23511597 × 10−4 2.35639032 × 10−4 2.56191561 × 10−4 2.63755705 × 10−4 4.82588483 × 10−4 5.50281373 × 10−4 6.15937519 × 10−4
Table 9. All 60 Token Presence events for GPT-2 Small at L = 30 and their reference probabilities. Read the left list first, then the right; events are ordered from rarer to more common. Token ActionCode Orderable cloneembedreportprint rawdownload ␣RandomRedditor DeliveryDate ÃÂÃÂÃÂÃÂ GoldMagikarp ␣volunte ␣srf ÃÂÃÂ ␣unintention ␣teasp inventoryQuantity ␣referen ␣looph \u30bc\u30a6\u30b9 ␣opio ÛÛ aution ␣mosqu \u30d8\u30e9 ␣SetTextColor glomer escription ␣contrace ␣resil ␣sshd AppData ␣\u30b5
Ground truth µ
Token
2.812 × 10−9
748 568 652 021 arijuana ␣Wem \u30aa Layout ␣363 Border ␣Requires ␣Petroleum ␣cowork aten loaded ␣marginally ␣unfounded ␣vib ␣massage fing ␣glam ␣Armed stock ␣Deity ␣voluntary ␣Baby ␣pathetic inating ␣couch ␣Energy
3.176 × 10−9 3.343 × 10−9 3.419 × 10−9 3.422 × 10−9 4.485 × 10−9 4.929 × 10−9 7.148 × 10−9 8.268 × 10−9 9.695 × 10−9 1.1427 × 10−8 1.1591 × 10−8 1.6259 × 10−8 2.697 × 10−8 4.4671 × 10−8 5.4553 × 10−8 6.8845 × 10−8 8.1701 × 10−8 9.6763 × 10−8 9.8635 × 10−8 1.06689 × 10−7 1.34227 × 10−7 3.78981 × 10−7 4.20284 × 10−7 4.38592 × 10−7 4.65958 × 10−7 5.52023 × 10−7 7.90799 × 10−7 8.59069 × 10−7 9.75121 × 10−7
62
Ground truth µ 2.385615 × 10−6 2.52272 × 10−6 3.088749 × 10−6 3.58389 × 10−6 4.099484 × 10−6 4.666225 × 10−6 4.693658 × 10−6 5.44136 × 10−6 7.860511 × 10−6 8.441393 × 10−6 1.2932401 × 10−5 1.6701184 × 10−5 3.0257203 × 10−5 3.8312104 × 10−5 4.0012317 × 10−5 4.2360556 × 10−5 5.1715971 × 10−5 6.1999679 × 10−5 7.2618226 × 10−5 9.8132645 × 10−5 1.13032715 × 10−4 1.2007045 × 10−4 1.22980506 × 10−4 1.3891251 × 10−4 1.67620776 × 10−4 1.87197977 × 10−4 2.29775411 × 10−4 2.3309172 × 10−4 3.03255045 × 10−4 3.25971749 × 10−4
Table 10. All 60 Token Presence events for GPT-2 Small at L = 50 and their reference probabilities. Read the left list first, then the right; events are ordered from rarer to more common. Token ThumbnailImage Nitrome Orderable reportprint ␣RandomRedditor ␣TheNitrome ÃÂÃÂÃÂÃÂÃÂÃÂÃÂÃÂÃÂÃÂÃÂÃÂÃÂÃÂÃÂÃÂ ÃÂÃÂÃÂÃÂÃÂÃÂÃÂÃÂ DeliveryDate ␣pione ␣subur ␣carbohyd ␣srf ␣earthqu ikuman ortunately \u899a\u9192 conservancy ModLoader ␣\u30b5\u30fc\u30c6\u30a3\u30ef\u30f3 BuyableInstoreAndOnline ␣proport ategory \u30fc\u30c6\u30a3 ␣Pwr ␣SetTextColor ␣gobl acebook escription ␣SolidGoldMagikarp
Ground truth µ
Token
3.969 × 10−9
LEASE intendent emort pkg ␣discrep paralleled enhagen udeau anchez flush ␣metabolites ␣Croat HHHH inho ILE ␣normalized ␣Sorceress ␣Arlington ␣forfeit ␣Plymouth ␣Lies nex ␣Antarctica jected ␣queue ␣Ale ␣artifacts ␣Dun ␣offices ␣Sal
4.643 × 10−9 4.971 × 10−9 5.398 × 10−9 5.416 × 10−9 5.453 × 10−9 6.567 × 10−9 7.011 × 10−9 7.108 × 10−9 9.779 × 10−9 1.1113 × 10−8 1.1206 × 10−8 1.6348 × 10−8 2.8878 × 10−8 3.0623 × 10−8 3.0704 × 10−8 5.6321 × 10−8 6.3908 × 10−8 7.9891 × 10−8 8.1083 × 10−8 1.02812 × 10−7 2.43806 × 10−7 2.90383 × 10−7 3.73774 × 10−7 6.12012 × 10−7 6.24561 × 10−7 6.65429 × 10−7 7.85046 × 10−7 8.97874 × 10−7 9.54816 × 10−7
63
Ground truth µ 2.497735 × 10−6 2.531015 × 10−6 2.591963 × 10−6 3.469834 × 10−6 4.214965 × 10−6 5.011348 × 10−6 5.647727 × 10−6 5.69348 × 10−6 6.136545 × 10−6 9.453041 × 10−6 1.5720263 × 10−5 2.0371159 × 10−5 2.58156 × 10−5 4.3729255 × 10−5 5.5443441 × 10−5 5.7964509 × 10−5 6.1175284 × 10−5 7.5363656 × 10−5 9.0493806 × 10−5 9.4129842 × 10−5 1.10739747 × 10−4 1.1683334 × 10−4 1.55974645 × 10−4 1.9980935 × 10−4 2.3909146 × 10−4 4.79594106 × 10−4 5.78806968 × 10−4 6.11178344 × 10−4 7.24431942 × 10−4 9.52690141 × 10−4
Table 11. All 66 Token Presence events for Gemma-2 at L = 10 and their reference probabilities. Read the left list first, then the right; events are ordered from rarer to more common. Token
Ground truth µ
Token
␣meio ␣\u043f\u0440\u044a ␣Cita \u6b63\u5728 cei opho McC Hb endid ␣Dial ␣decoration Animals ␣utilized \u6743 Zoo Vish vano ␣sweats ␣articulated helia ␣Sylvain ␣Skinny avor Cassie ␣Kjell ␣hastily ␣whe ␣unsteady ␣disliked ␣pep Madame ␣Ace Maxwell
1.084 × 10−9 1.334 × 10−9 1.406 × 10−9 1.45 × 10−9 1.601 × 10−9 2.166 × 10−9 2.345 × 10−9 2.65 × 10−9 4.994 × 10−9 9.451 × 10−9 1.0744 × 10−8 1.0839 × 10−8 1.0885 × 10−8 1.2994 × 10−8 1.793 × 10−8 1.9079 × 10−8 1.9914 × 10−8 2.0707 × 10−8 2.2923 × 10−8 5.1906 × 10−8 6.5284 × 10−8 7.6772 × 10−8 1.36062 × 10−7 1.64704 × 10−7 1.65677 × 10−7 1.88567 × 10−7 2.76225 × 10−7 3.01208 × 10−7 3.98665 × 10−7 4.54393 × 10−7 5.48416 × 10−7 6.51134 × 10−7 6.83873 × 10−7
␣shedding attered ␣coating ␣mostly ␣mor ␣disapproval ␣nurturing ␣tank ␣groans ␣bison then ␣mural ␣Aur mills ␣eccentric ␣sharper ␣cutter ␣ramp ␣Casper ␣scenic ␣longer ␣railing ␣Robin ␣shifting ␣Avon ␣unsettling ␣resolute ␣librarian ␣steady ana ␣brooding ␣whimsical ␣distant
Ground truth µ 9.74816 × 10−7 1.015341 × 10−6 1.229574 × 10−6 1.364986 × 10−6 1.376871 × 10−6 1.803206 × 10−6 2.100988 × 10−6 2.570783 × 10−6 5.677432 × 10−6 6.391679 × 10−6 6.78363 × 10−6 7.759179 × 10−6 9.377181 × 10−6 1.2066338 × 10−5 1.218224 × 10−5 1.23574 × 10−5 1.5257824 × 10−5 1.9968653 × 10−5 2.0307949 × 10−5 2.1126934 × 10−5 2.2296295 × 10−5 4.2092146 × 10−5 5.653832 × 10−5 7.2919989 × 10−5 7.9591409 × 10−5 1.10149827 × 10−4 1.18333068 × 10−4 1.21119221 × 10−4 1.23186823 × 10−4 1.38252377 × 10−4 1.72549393 × 10−4 2.31835613 × 10−4 5.34711289 × 10−4
Table 12. All Profanity Classifier events and reference probabilities. The event is f (y) > κ. A dash means that this model–length–threshold combination is absent from the experimental inventory. Cutoff κ
GPT-2, L = 10
GPT-2, L = 50
Gemma, L = 10
0.1 0.15 0.2 0.25 0.3 0.35 0.4 0.45 0.5 0.55 0.6 0.65 0.7 0.75 0.8 0.85 0.9 0.95 0.99 0.995 0.999
2.8353515625 × 10−1
6.888671875 × 10−2
2.11609375 × 10−1 1.6591015625 × 10−1 1.33453125 × 10−1 1.085703125 × 10−1 8.918359375 × 10−2 7.340625 × 10−2 6.056640625 × 10−2 4.980859375 × 10−2 4.091015625 × 10−2 3.321484375 × 10−2 2.64765625 × 10−2 2.061328125 × 10−2 1.58671875 × 10−2 1.1671875 × 10−2 7.86328125 × 10−3 4.6953125 × 10−3 2.14453125 × 10−3 4.921875 × 10−4 2.8125 × 10−4 8.59375 × 10−5
5.83203125 × 10−2 5.11953125 × 10−2 4.6140625 × 10−2 4.18125 × 10−2 3.82890625 × 10−2 3.507421875 × 10−2 3.22734375 × 10−2 2.970703125 × 10−2 2.7375 × 10−2 2.5140625 × 10−2 2.32109375 × 10−2 2.10859375 × 10−2 1.877734375 × 10−2 1.6671875 × 10−2 1.43359375 × 10−2 1.187109375 × 10−2 8.4296875 × 10−3 4.0546875 × 10−3 2.9375 × 10−3 1.44140625 × 10−3
7.9390625 × 10−2 5.4046875 × 10−2 3.921875 × 10−2 3.0921875 × 10−2 2.615625 × 10−2 2.28125 × 10−2 1.965625 × 10−2 1.66875 × 10−2 1.4265625 × 10−2 1.21875 × 10−2 1.0421875 × 10−2 8.4375 × 10−3 7.109375 × 10−3 5.828125 × 10−3 4.578125 × 10−3 3.28125 × 10−3 1.9375 × 10−3 7.5 × 10−4 3.73527547 × 10−4 2.18535188 × 10−4 6.3319311 × 10−5
64
Table 13. All Compositional Token Presence events and reference probabilities for GPT-2 Small at L = 10. Conjunctive events require every listed token in any order; Ordered Sequence events require the listed order but allow intervening tokens. Leading spaces are part of each target token. Event type
Target tokens
Conjunctive (AND) Conjunctive (AND) Conjunctive (AND) Conjunctive (AND) Conjunctive (AND) Ordered Sequence (SEQ) Ordered Sequence (SEQ) Ordered Sequence (SEQ) Ordered Sequence (SEQ) Ordered Sequence (SEQ)
{␣connect, ␣the, ␣was} {␣the, ␣was, ␣village} {␣a, ␣the, ␣king} {␣the, ␣was, ␣people} {␣the, ␣was, ␣and} ␣the → ␣was → ␣connect ␣the → ␣was → ␣village ␣a → ␣the → ␣king ␣the → ␣was → ␣people ␣the → ␣was → ␣and
65
Ground truth µ 3.7222278697068845 × 10−6 8.061106367685 × 10−5 1.981384612154 × 10−4 5.053878165782 × 10−4 6.57273504138 × 10−3 1.246940501071 × 10−6 5.749512319003 × 10−6 3.01940971556 × 10−5 1.033016124585 × 10−4 2.589375125567 × 10−3