The More It Says, the More You Pay: A Black-Box Audit of Provider-Side Token Inflation in LLM Services Leilei Chen Lan Zhang Jiewei Lai Yixiao Huang
Chen Tang Zhaopeng Zhang
Pengcheng Sun Xinpeng Shen
arXiv:2609.20370v1 [cs.CR] 17 Sep 2026
University of Science and Technology of China
Abstract
Query (x):What is a liquid connective tissue? (A) solid (B) cells ... (H) osculum
In pay-per-token LLM services, the more a model says, the more users pay. Dishonest providers can covertly manipulate generation to inflate output tokens while largely preserving task utility. We define such manipulation as a Provider-Side Token Inflation Attack (PTIA) and instantiate five representative attacks at the query, prompt, representation, and model levels of the provider-controlled pipeline. Our experiments show that each attack increases mean output length to more than 10.2× the clean baseline, demonstrating PTIA’s financial appeal and feasibility at multiple stages of generation. Yet auditing PTIA from black-box responses is difficult for users. Our key observation is PTIA saturation: an initial attack sharply lengthens output, but further strengthening or composition has much less effect. We trace this saturation to stopping behavior: an initial PTIA sharply lowers the endof-sequence token probability, whereas further intervention lowers it only marginally. Building on this insight, we design a lightweight single-probe audit that applies a controlled lengthening intervention. Under PTIA, the probe induces far fewer additional tokens than under normal service. The audit requires neither a trusted local reference model nor historical clean responses, and its separately issued original and probed requests resemble ordinary traffic, making evasion difficult. Across four open-weight models, it achieves an average detection rate of 85.1% with false-positive rates below 2%. Across 15 real LLM API services, the audit flags 7 for PTIA-consistent behavior.
1
Normal Service Provider Prompt Construction "You are a helpful assistant" + Query
Output(Clean): 43 tokens The question asks for a liquid connective tissue. Liquid connective tissue refers to blood, which consists of cells... Answer: C
Malicious Service Provider Prompt Construction "You are a helpful assistant" + "Validate the answer from multiple perspectives before producing the final response." + Query
Output(Attacked): 561 tokens
(Inflated with PTIA) Let's analyze... (1)Definition:Blood/Lymph = Liquid CT. (2)Evaluate:(C)Blood is CT.(A)Solid is not... Final answer: C
Figure 1: Illustration of a Provider-Side Token Inflation Attack (PTIA).
than 70 providers [27]; it also reports serving more than eight million developers across over 400 models and processing 25 trillion tokens per week [28]. However, the opacity of pay-per-token LLM service backends raises trust and accountability concerns, motivating calls to audit hidden provider operations and billing [32]. Studies have traced 17 shadow APIs to 187 academic papers and uncovered deceptive model claims [40], while audits of commercial gateways found silent model substitution and billing deviations of up to 62.8% [18]. Existing methods verify model authenticity [8, 42] or token-accounting accuracy [34, 35], but do not determine whether the provider manipulated the generation process behind the billed output tokens. In pay-per-token LLM services, the more a model says, the more users pay, giving a dishonest provider a financial incentive to increase output length. The provider controls the generation process and can covertly manipulate it. However, users cannot observe how their responses are generated. As Figure 1 illustrates, a single user-invisible provider-side instruction can increase billable output by over 13× while pre-
Introduction
LLM services have become an increasingly important means for developers to access models and integrate them into applications, agents, and scientific workflows [2, 5, 40]. Growing demand and commercial opportunities have encouraged an increasing number of model developers, cloud platforms, and third-party gateways to offer LLM services [5, 28, 40]. For example, OpenRouter reports routing requests across more 1
Mean Output Tokens
serving response utility. We call an undisclosed provider-side manipulation that increases billable output tokens a ProviderSide Token Inflation Attack (PTIA). Systematically characterizing how providers can implement PTIA in practice is challenging because their interventions can target multiple stages of generation and take different forms. And prior work has studied how attackers prolong generation to exhaust provider resources or disrupt service availability [6, 20, 30, 37], but not how providers manipulate generation for financial gain. A practical PTIA must inflate billable output while avoiding noticeable degradation in task utility or response naturalness. In this work, we analyze the provider-controlled generation process end to end and instantiate five representative PTIAs spanning hidden system instructions, query prefixes, semantic elaboration, learned soft inputs, and model fine-tuning. These attacks substantially inflate billable output while largely preserving task utility, making the manipulation difficult for users to notice. Together, they demonstrate that PTIA is both economically attractive and technically practical for a provider paid by the token. PTIA increases actual, user-visible output tokens and is implemented by the provider itself. Existing token-auditing methods do not address this setting: they focus on hidden reasoning-token usage, and some require provider cooperation [31, 36]. Auditing PTIA is difficult because providers control the hidden generation process that determines billable output, while users observe only the returned responses. Users cannot reliably audit PTIA from output length alone because they lack a trusted baseline for the same request under unmanipulated generation. Even an inflated response can remain correct and outwardly ordinary. Moreover, providers can randomly apply PTIA to some requests. These constraints raise a central question: How can an auditor detect PTIA in a service using only black-box queries? To explore the behavioral patterns of PTIA from responses alone, we systematically strengthen individual attacks and combine them. We find PTIA saturation, as shown in Figure 2: when a PTIA is first applied, it substantially increases output length, but strengthening it further adds fewer tokens. This pattern also holds when combining PTIAs across generation stages. We trace this saturation to stopping behavior: an initial PTIA sharply lowers the end-of-sequence token probability, whereas further intervention lowers it only marginally. Our key insight is that PTIA saturation changes how generation responds to a subsequent lengthening intervention. When the provider has already applied PTIA, a second, similar intervention produces a much smaller output-length increase than it would under normal service. Building on this insight, we design a single-probe black-box auditing framework for PTIA. For each question, the auditor issues two separate requests: the original query and a probed query with a single additional lengthening intervention. We aggregate the resulting output-length changes across questions: the smaller the change, the stronger the evidence of PTIA. The framework
HSPA POA LSEA
3,000 2,500
AS-TIA FT-TIA
2,000 1,500 1,000 500 0
0
1
2
3
4
5
Attack Intensity Level (0 = Clean)
Figure 2: Saturation across five PTIAs on Ministral-3-14B. requires neither backend access, a trusted local model, nor historical clean responses. It is also difficult to evade because the provider must identify the independently submitted original request as part of an audit pair, even though it resembles ordinary traffic. We evaluate all five PTIAs and our saturation-guided audit on four open-weight models. All five PTIAs increase mean output length to more than 10.2× the clean baseline on every model while largely preserving task accuracy, with an average decrease of only 1.88 percentage points. An LLM-based evaluation reveals an inflation–naturalness trade-off: less aggressive variants better preserve response naturalness. Under selective PTIA deployment, the audit achieves an 85.1% macro-average q-AUC, 13.4 percentage points above RUT, which requires a trusted local model. Its average false-positive rate remains below 2% under all three system-prompt settings: none, benign helpfulness, and benign safety. In summary, this work makes the following contributions: • Attacks. We define PTIA and systematically instantiate five representative attacks spanning the generation pipeline, exposing multiple practical paths to inflate billable output (§2). • Insight. We explore the behavioral patterns of PTIA through systematic attack strengthening and composition and uncover PTIA saturation, enabling PTIA auditing from responses alone (§3). • Audit. We design a single-probe black-box audit framework that lets users detect PTIA by applying a second, similar intervention, enabling them to independently test opaque LLM services for token inflation (§4). • Evaluation. We evaluate attack effectiveness and audit reliability on four open-weight models, then audit 15 real-world LLM services and find PTIA-consistent behavior in 7 (§5).
2
Provider-Side Token Inflation Attacks
In this section, we characterize how a dishonest API provider can implement PTIAs in practice. We first describe the 2
provider-controlled generation pipeline and formalize the PTIA threat model. We then establish four requirements for practical PTIAs and instantiate five representative PTIAs spanning multiple provider-controlled intervention points in the generation process.
2.1
actual output tokens by manipulating generation rather than model identity or token accounting. PTIA formulation. Let s0 denote the provider-side system prompt under the unmanipulated service configuration. We represent the provider’s prompt assembly, tokenization, and embedding pipeline by an input-construction function I. The input representation received by the model is
LLM API Generation Pipeline
e0 = I(s0 , x).
Commercial LLM APIs expose a request–response interface in which users submit queries and generation parameters and receive generated responses with token-usage metadata [1, 9, 26]. Drawing on these documented interfaces and established LLM serving workflows [15, 38], we abstract the provider-controlled backend into four functional stages: ❶ Query processing parses and preprocesses the request; ❷ Prompt construction assembles the query, provider-side instructions, and conversation history into the model context; ❸ Input representation serializes and tokenizes the context and maps the resulting token IDs to input embeddings; and ❹ Inference and decoding autoregressively generate the billable output-token sequence until a stop token is produced or an output-token limit is reached.
Let θ denote the base-model parameters and let Gθ (· | e, φ) denote the distribution over response sequences generated from input representation e under the user-visible generation parameters φ. The response under the unmanipulated configuration is Y0 ∼ Gθ (· | e0 , φ). A PTIA a may alter the textual inputs processed by I, directly modify the resulting input representation, or modify the model parameters. Let ea and θa denote the input representation and model parameters after the intervention. The attacked response is then Ya ∼ Gθa (· | ea , φ).
2.2
Threat Model
Any component not modified by the attack retains its value from the unmanipulated configuration. The attack seeks to increase the expected billable output length,
Attack scenario. The user submits a query x and user-visible generation parameters φ to a pay-per-token LLM API, and the provider returns a generated response together with its reported token usage [23, 35]. At a fixed output-token price, the output charge is proportional to the number of returned output tokens. A dishonest provider can manipulate generation to make the advertised model produce additional billable output tokens, thereby increasing its revenue. Any undisclosed provider-side intervention that increases billable output falls within our definition of PTIA, even if the provider claims that it is intended to improve response quality. Users have the right to decide whether to accept such interventions. Attacker capabilities. The attacker is the LLM service provider, with control over the backend stages described in Section 2.1. It may intervene in internal query processing, prompt construction, input representations, or model parameters. For each request, the provider can decide whether to apply PTIA. The attacker seeks to preserve response utility and naturalness so that the manipulation is difficult to recognize from the response alone. User observations. Users control the submitted query and user-visible generation parameters and observe only the returned response and reported token usage. They cannot observe the provider-controlled generation pipeline or determine what response the service would have returned under an unmanipulated generation process. Scope. We assume that all tokens billed as output are returned to and visible to the user and that the provider accurately reports their number for each request. PTIA covertly increases
E[|Ya |] > E[|Y0 |] , while largely preserving the utility of the returned response.
2.3
Practical PTIA Requirements
Existing length-inducing attacks can be used to achieve provider-side token inflation. However, increasing output length alone does not make an attack a practical PTIA. It must also preserve response utility and be deployable across requests and models. We therefore define the following four requirements: • C1: Inflation effectiveness. The attack systematically increases the number of user-visible output tokens. • C2: Utility preservation. The attack preserves task utility while increasing output length. • C3: Operational scalability. The attack is reusable across requests and does not require query-specific optimization at deployment time. • C4: Broad applicability. The attack mechanism does not rely on specialized capabilities limited to particular models. Table 1 compares existing inference-cost attacks against these requirements. While existing attacks satisfy individual requirements, none satisfies all four. 3
2.4.2
Table 1: Comparison of existing inference-cost attacks against the requirements of PTIA. A check mark indicates that an attack satisfies the corresponding requirement. Approach
C1 C2 C3 C4
Excessive Reasoning Attack [30] POT [16] Engorgio [6] Non-Halting Queries [10] ThinkTrap [17] LLMEffiChecker [7] Crabs [41] OverThink [14] RA-ICA [19] BadThink [20] BadReasoner [37] Deadlock Attack [39]
✓
✓ ✓ ✓
POA targets query processing. It prepends a short, taskagnostic prefix p to the user query, x 7→ p ∥ x,
✓ ✓
✓ ✓ ✓ ✓ ✓
steering the model toward more extensive analysis, comparisons, or explanations while preserving the original task. We construct p offline for each target model through an iterative procedure that alternates candidate proposal, targetmodel evaluation, and refinement. A separate proposer model first generates a set of natural, query-independent candidate prefixes. Each candidate is then evaluated on the target model according to both the output-length increase it induces and its agreement with the corresponding clean prediction. The highest-scoring candidates and their evaluation summaries are returned to the proposer as feedback for generating the next set of candidates. After multiple rounds, we freeze the highest-ranked prefix for each target model and reuse it across requests. POA therefore incurs a one-time offline optimization cost but requires no query-specific optimization during inference.
✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓
✓ ✓ ✓ ✓
Table 2: Target stages and transformations of five representative PTIAs. PTIA
Target stage
Transformation
HSPA POA LSEA AS-TIA FT-TIA
Prompt construction Query processing Query processing Input representation Model layer
s0 7→ s0 ∥ δs x 7→ p ∥ x x 7→ x ∥ ξ(x) e0 7→ e0 ⊕ P (x, θ) 7→ (x ∥ z, θ + ∆θ)
2.4
2.4.3
Lightweight Semantic Elaboration Attack
LSEA operates during query processing, using a lightweight model gψ to elaborate each query x in a single pass. The resulting text ξ(x) introduces relevant background or analytical dimensions without directly providing the answer: ξ(x) = gψ (x),
Representative PTIAs
x 7→ x ∥ ξ(x).
This append-only transformation preserves the original query while expanding the aspects on which the target model can elaborate. It requires neither iterative search nor queryspecific optimization.
We next systematically characterize the PTIA attack surface throughout the provider-controlled generation pipeline and instantiate five representative attacks: the Hidden SystemPrompt Attack (HSPA), Prefix-Based Overthinking Attack (POA), Lightweight Semantic Elaboration Attack (LSEA), Adversarial Soft-Suffix Token Inflation Attack (AS-TIA), and Fine-Tuning-Based Token Inflation Attack (FT-TIA). Table 2 summarizes the target stage and transformation of each attack using the notation defined in our threat model. Complete implementation details and attack-intensity configurations are provided in Appendix A. 2.4.1
Prefix-Based Overthinking Attack
2.4.4
Adversarial Soft-Suffix Token Inflation Attack
AS-TIA targets the input representation. Its central goal is to learn a model-specific but query-universal continuous suffix that extends generation while preserving query-dependent task behavior. Let P ∈ Rm×d denote a soft suffix of length m, where d is the model’s embedding dimension. With the target-model parameters frozen, the provider directly appends P to the input representation:
Hidden System-Prompt Attack
HSPA targets prompt construction. It appends a fixed, userinvisible length-inducing instruction δs to the provider-side system prompt, s0 7→ s0 ∥ δs ,
eAS = e0 ⊕ P. Because the provider controls the input representation, P need not correspond to user-visible discrete text. We optimize P in two stages: long-trajectory learning captures reusable structures for long, answer-consistent responses, while freedecoding correction guides the model to follow them during generation.
steering the model toward more extensive explanations or additional details while answering the original query. Because δs is query-independent, it can be reused across requests without per-query optimization or the associated overhead. 4
where Ic and I p index the clean and poisoned examples, respectively. Poisoned examples associate the trigger with longoutput generation, whereas clean examples encourage ordinary task-response behavior on untriggered inputs and reduce unconditional activation. The template h provides long-output supervision, while the appended clean target is intended to retain the original task-specific content and answer information. We optimize only the LoRA parameters using a responseonly SFT objective, with prompt tokens excluded from the loss. Once trained, the same model-specific adapter and trigger can be reused across requests without query-specific optimization. For a request selected for attack, the provider internally appends z and routes the resulting input to the adapted model. Together, these five PTIAs instantiate provider-controlled routes to increasing actual billable output at distinct points of the generation pipeline: hidden instructions, optimized prefixes, semantic rewriting, continuous soft suffixes, and parameter adaptation.
Long-trajectory learning. This stage learns reusable longresponse structures while preserving query-dependent content. For each training query xi , we construct an answer-consistent long-output trajectory τi using the frozen target model and a fixed bank of length-inducing instructions. These instructions are used only for offline target construction and are not appended during deployment. To learn structures that transfer across queries, we identify sentence-initial spans whose token likelihoods are least sensitive to the concrete question and treat their positions as structural anchors Ai . The remaining positions retain query-dependent content. Under teacher forcing on τi , the first-stage objective is
L1 (P) = λstop Lstop + λanc Lanc + λkl Lkl . Here, Lstop discourages premature termination, Lanc promotes reusable long-output structure at anchor positions, and Lkl preserves the clean model’s next-token distribution at nonanchor positions. We denote the resulting suffix by P(1) . Free-decoding correction. The first stage alone does not ensure that the model enters a long-output trajectory during free decoding. Starting from P(1) , we therefore roll out the model under the current suffix and apply an entry loss Lentry when generation enters the answer region or terminates prematurely. This loss teacher-forces an aligned continuation from the corresponding long-output trajectory, yielding P∗ = arg min L1 (P) + λentry Lentry (P) ,
3
Under black-box access, a PTIA may leave no direct sign of manipulation: the returned response can remain accurate, and its token count can still be reported correctly. We therefore look beyond absolute output length and ask how PTIA-induced length gains change under further intervention. Although the five PTIAs act at different points in the provider-controlled backend, they all influence output length through the same autoregressive generation process. We test two forms of further intervention: strengthening a single PTIA (Section 3.1) and adding a distinct PTIA (Section 3.2). Both reveal the same pattern: the initial intervention substantially increases output length, whereas further intervention produces only a much smaller additional gain. We call this shared behavior PTIA saturation.
P
initialized from P(1) . The resulting P∗ can be reused across queries to the same target model without query-specific optimization. 2.4.5
PTIA Saturation
Fine-Tuning-Based Token Inflation Attack
FT-TIA targets the model layer. It encodes a conditional longoutput behavior in a model-specific LoRA adapter [12] and activates that behavior through a fixed, user-invisible trigger z. Let ∆θ denote the learned LoRA update while the base-model parameters θ remain frozen. FT-TIA applies the transformation (x, θ) 7→ (x ∥ z, θ + ∆θ) .
3.1
The trigger activates the learned association at inference time, while the adapted model parameters encode the conditional token-inflation behavior. We learn this trigger-dependent behavior through conditional training on clean and poisoned examples. Conditional training. We construct a mixed SFT dataset containing clean and poisoned examples. Let y0i denote the original clean target response associated with xi in the source SFT dataset, and let h be a fixed long-output template shared across poisoned examples. The training pairs are ( (xi , y0i ), i ∈ Ic , (e xi , yei ) = 0 xi ∥ z, h ∥ yi , i ∈ I p ,
Scaling Individual PTIAs
We first examine whether strengthening a single PTIA continues to produce similar gains in output length. For each PTIA, level 0 denotes the Clean configuration, while levels 1–5 represent progressively stronger, PTIA-specific configurations. Because intensity is controlled through different variables for different PTIAs, the levels are ordered only within each PTIA and are not comparable across attacks. As shown in Figure 2, the transition from level 0 to level 1 produces the largest increase in mean output-token count for all five PTIAs on Ministral-3-14B-Instruct. HSPA, LSEA, and AS-TIA change little after this initial increase. POA and FT-TIA show additional growth at some later levels, but these gains remain much smaller than the initial increase. The same 5
pattern holds across the other evaluated models, as detailed in Appendix C.2. We therefore make the following observation. Observation 1: Single-PTIA saturation. As an individual PTIA is strengthened, the initial intervention produces a large increase in output length, whereas further strengthening produces much smaller gains. We call this pattern single-PTIA saturation.
HSPA
Existing PTIA a
3.2
≤0
Composing PTIAs
Scaling an individual PTIA captures only one form of PTIA saturation. A provider may also combine PTIAs that act at different backend stages. We therefore ask whether, when one PTIA is already present, adding a distinct PTIA yields an output-length gain comparable to its standalone gain. We evaluate all ten pairwise combinations of the five PTIAs and, for each pair, compare the marginal gain of the added PTIA with its gain from the Clean configuration. We report both conditional directions, yielding 20 comparisons. Direction identifies which PTIA provides the conditional baseline rather than their execution order. Figure 3 reports the resulting marginalgain ratios. Figure 3 shows that composing PTIAs usually reduces the marginal gain of the added attack. Among the 20 directional comparisons, 19 have a mean marginal-gain ratio below one. For every comparison, at least three of the four evaluated models also have a ratio below one. Four comparisons have negative mean ratios, indicating that an added PTIA can even reduce output length relative to the existing attack alone. We therefore obtain the following observation. Observation 2: Cross-PTIA saturation. PTIA saturation arises not only when a single attack is strengthened, but also when distinct attack paths are combined. Once generation is already affected by one length-inducing PTIA, another PTIA typically produces only part of its standalone gain. In summary, we observe PTIA saturation under both attack strengthening and composition. This finding offers a new perspective for PTIA auditing: rather than relying on absolute output length, an auditor can look for evidence that the generation process is already close to saturation.
4
POA
LSEA
AS-TIA
FT-TIA
Mean Rb̄ ∣ a
0.5
1.0
1.2
—
0.65
0.80
0.46
1.06
↓ 4/4
↓ 3/4
↓ 3/4
↓ 3/4
−0.05
—
0.64
−0.28
0.96
↓ 3/4
↓ 4/4
↓ 3/4
0.67
0.78
0.93
↓ 3/4
—
0.35
↓ 3/4
↓ 3/4
↓ 3/4
0.48
0.34
0.53
↓ 4/4
↓ 3/4
—
0.12
↓ 3/4
0.61
0.38
↓ 3/4
↓ 3/4
↓ 3/4
↓ 4/4
HSPA
POA
LSEA
AS-TIA
↓ 4/4
−0.48 −9.46
↓ 4/4
— FT-TIA
Added PTIA b
Figure 3: Pairwise PTIA marginal-gain ratios across four models. For existing PTIA a (row) and added PTIA b (column), each cell reports R̄b|a , where Rb|a = (La,b − La )/(Lb − L0 ) and L denotes mean output length under the indicated configuration. Rb|a < 1 indicates a smaller gain than b’s standalone gain, while Rb|a < 0 means that the composition is shorter than a alone. The ↓ k/4 annotation counts models with Rb|a < 1; diagonal cells are omitted.
4.1
The User as an Additional Attacker
PTIA saturation creates an opportunity for active black-box testing. Absolute response length is not reliable evidence because it varies substantially across questions and models. Instead, the auditor introduces a known benign intervention intended to increase output length and observes how strongly the service responds. Under unmanipulated generation, this intervention is expected to produce a positive length gain. For a PTIA-affected service, however, the original response may already have been lengthened, and saturation predicts that the same intervention will produce a smaller gain or no gain at all. By deliberately attempting to further increase the output length, the user acts as an additional attacker. The extent to which the probe succeeds becomes the audit signal. For each question, comparing the original and probeaugmented responses provides a within-question measurement. The audit examines the direction and relative magnitude of the probe-induced change, rather than judging whether either response is unusually long in absolute terms. This comparison reduces dependence on the natural variation in response length and requires neither an expected clean length for the question nor a trusted local reference model. Because the two requests are independent and PTIA may be deployed selectively, the resulting evidence is aggregated across multi-
Black-Box Token Inflation Audit
In this section, we turn PTIA saturation into a black-box audit by letting the API user apply an additional benign lengthinducing intervention. For each question, the auditor compares the output lengths of two independent requests: the original query and the same query with a single length-inducing probe appended. Under normal generation, the probe is expected to produce a clear length increase; when PTIA has already lengthened the response, saturation predicts a smaller or nonpositive increase. Aggregating this contrast across questions enables service-level detection without backend access, a local reference model, or historical clean responses. 6
Audit Evidence. Given the observed length change ∆i , directly applying a single threshold to its raw value would be unreliable because response lengths vary substantially across questions. We therefore examine both whether the probe produces a positive length gain and how large that gain is relative to the original response. Both forms of evidence depart from the expected positive response to a length-inducing probe under normal generation and are consistent with the saturation effect under PTIA. First, we consider the direction of the length change. Under normal generation, appending a length-inducing probe is expected to increase the response length. A non-positive change therefore indicates that the probe’s intended effect has been fully suppressed or reversed, which is a strong observable manifestation of saturation. We represent this directional evidence as Ei,1 = I[∆i ≤ 0] .
ple questions before making a service-level decision.
4.2
Audit Setting
Auditor. The auditor is an API user. Within a finite query budget, the auditor can submit queries and observe the resulting responses and output-token counts to determine whether the target service has deployed a PTIA. Adversarial provider. A dishonest provider may attempt to evade detection by activating PTIA probabilistically across requests. For each API request r, let Ar denote whether PTIA is activated: Ar ∼ Bernoulli(q), (1) where q ∈ [0, 1] is the unknown attack rate. Here, q = 0 represents an honest service, q = 1 represents persistent deployment, and 0 < q < 1 represents selective deployment. The auditor observes neither Ar nor q. Audit objective. We formulate the auditing task as a servicelevel hypothesis test: H0 : q = 0,
H1 : q > 0.
Second, even when the probe produces a positive increase, the increase may be small relative to the length already produced for the original query. To make this comparison robust to differences in response scale, we define the normalized positive gain as
(2)
The audit determines whether the target service deploys PTIA, without identifying which requests were attacked or which PTIA was used.
4.3
Ri =
Audit Pipeline
(1)
(0)
,
Li + 1
[z]+ = max(z, 0),
and derive the relative-gain evidence
Our audit measures how a benign length-inducing probe changes output length. For each question, it sends the original and probe-augmented queries as independent requests. Without PTIA, the probe should increase output length; with PTIA, the original response may already be lengthened, so PTIA saturation predicts a smaller or non-positive increase. The audit has three stages: Single-Probe Comparison measures this change; Audit Evidence converts its direction and normalized magnitude into question-level evidence; and Evidence Aggregation aggregates this evidence across questions into a service-level PTIA-consistent signal, including under selective deployment. Single-Probe Comparison. For each audit query xi , the auditor selects a benign length-inducing probe pi that encourages a longer response without changing the underlying task or its required answer format. The auditor then constructs two inputs: (0) xi = x i , xi
[∆i ]+
Ei,2 = I[Ri ≤ τ] , where τ is calibrated on a separate calibration set. A small value of Ri means that the probe-augmented response is only slightly longer than the original response after accounting for the original response length. This is consistent with the saturation effect: when PTIA has already lengthened the response to the original query, appending another length-inducing instruction increases the response by less than it normally would. When ∆i ≤ 0, we have Ri = 0, so the observation satisfies both the small-relative-gain condition and the stronger nonpositive-gain condition. Evidence Aggregation. For each audit question, we combine the two evidence indicators into an equal-weight questionlevel score: si = Ei,1 + Ei,2 ,
si ∈ {0, 1, 2}.
A score of zero indicates that the probe produces a sufficiently large positive increase. A score of one indicates that the response becomes longer, but the increase is small relative to the original response. A score of two indicates a non-positive length change and therefore the strongest evidence of saturation. Given an audit set of n questions, let
= x i ⊕ pi ,
where ⊕ denotes appending the probe to the original user (0) (1) message. The auditor submits xi and xi as two separate (0) API requests and records their output-token counts as Li (1) and Li , respectively. We define the observed probe-induced length change as (1) (0) ∆i = L i − Li .
s(1) ≥ s(2) ≥ · · · ≥ s(n) 7
• RQ4: What PTIA signals does our audit identify in realworld LLM services?
Algorithm 1: Single-probe black-box PTIA auditing Input: Target API f , audit questions X = {xi }ni=1 , and parameters τ, k, and H Output: Decision on H0 1 foreach xi ∈ X do 2 select a benign length-inducing probe p; 3 4 5 6 7 8
5.1
Experimental Setup
Models and Datasets. We conduct experiments on four openweight models: Llama-3.1-8B-Instruct [21], Ministral-3-14BInstruct-2512 [25], Qwen3-14B, and Qwen3-32B [29]. These models span three model families and parameter scales from 8B to 32B. Our experiments use two question-answering datasets: QASC [13] and OpenBookQA [24]. FT-TIA additionally uses R1-Distill-SFT v0 [22] as its fine-tuning corpus. Implementation Details. For each model–question pair, the clean and attacked runs use the same underlying task input and generation settings, differing only in the intervention introduced by the corresponding attack. We use greedy decoding and limit each response to 4,096 generated tokens. Response length is measured by the number of generated tokens, excluding prompt tokens. Responses that reach the generation limit are marked as truncated but retained in the main evaluation. We report the truncation rate and conduct a separate analysis to verify that the observed saturation is not caused by the generation limit. All experiments are conducted on four NVIDIA A100-SXM4 GPUs with 80 GB of memory each. Complete attack-specific implementation details and hyperparameter settings are provided in Appendix A.
(0) (1) (Li , Li ) ← Q UERY S EPARATELY( f , xi , xi ⊕ p); (1) (0) ∆i ← Li − Li ; (0) Ri ← [∆i ]+ /(Li + 1); Ei,1 ← I[∆i ≤ 0];
Ei,2 ← I[Ri ≤ τ]; si ← Ei,1 + Ei,2 ;
S ← T OP KS UM({si }ni=1 , k); 10 if S ≥ H then 11 return reject H0 ; 12 else 13 return fail to reject H0 ; 9
denote the question-level scores in descending order. We define the aggregate audit statistic as the sum of the k largest scores: k
S = ∑ s( j) . j=1
The auditor reports PTIA when
5.2
S ≥ H,
In this phase, we evaluate five PTIAs by measuring outputtoken inflation and changes in task accuracy relative to the clean setting. Experiment Setup. We evaluate all five PTIAs across four models on a fixed set of 500 questions from the QASC validation split. This evaluation set is disjoint from the QASC samples used for AS-TIA construction and the saturation analysis. For an attack a, we define its Token Inflation Ratio (TIR) as 1 N ∑ Li,a TIRa = 1 N N i=1 , N ∑i=1 Li,clean
where H is the service-level decision threshold. We use the largest k scores rather than averaging over all audit questions because, under selective deployment, PTIA may affect only a subset of the requests and its evidence may otherwise be diluted by unaffected questions. At the same time, the decision threshold prevents a single atypical response from determining the audit outcome. In our evaluation, each audit uses n = 20 questions with k = 3 and H = 3. Because the maximum question-level score is two, an alarm necessarily requires supporting evidence from more than one question. The complete audit procedure is summarized in Algorithm 1.
5
PTIA Effectiveness (RQ1)
where Li,a and Li,clean denote the output-token counts for question i under attack a and Clean, respectively. Thus, Clean corresponds to 1×, and a larger TIR indicates greater token inflation. Attack Baseline. We use the provider-side overcharging method proposed by Velasco et al. [34] as the attack baseline. Their method finds a longer yet plausible tokenization of an existing output, thereby increasing the reported number of billable tokens without changing the user-visible response. In contrast, PTIAs intervene in the generation process and cause the model to generate more output tokens. Although the two approaches address different problems through different mechanisms, Velasco et al. represent the closest prior
Evaluation
This section presents our experimental evaluation, organized around the following four research questions: • RQ1: How effectively can providers inflate output with PTIAs while preserving response quality? • RQ2: Why do PTIAs exhibit saturation under both attack strengthening and composition? • RQ3: How can PTIAs be detected in a black-box setting from returned responses alone? 8
Token Inflation Ratio (×)
Velasco et al.
HSPA
POA
LSEA
AS-TIA
FT-TIA
Table 3: Task accuracy and response naturalness across four models. Naturalness is evaluated by GPT-5.6, with overall scores of 4 or 5 classified as natural.
1000 500 200 100
(a) Task accuracy (%)
50 20
Model
Clean Velasco HSPA POA LSEA AS-TIA FT-TIA
Llama-3.1-8B Ministral-3-14B Qwen3-14B Qwen3-32B
72.2 77.4 79.0 83.0
72.2 77.4 79.0 83.0
74.8 80.4 80.0 83.8
60.6 68.2 84.4 84.6
69.2 74.8 82.2 83.0
64.0 78.2 75.6 67.6
67.6 80.2 80.2 81.0
Mean
77.9
77.9
79.8 74.5
77.3
71.4
77.3
10 5 2 1
-8B
B
B
-3.1
ma
Lla
-14 al-3
istr
Min
-14 en3
Qw
B
-32 en3
Qw
Figure 4: Token inflation ratios achieved by the prior attack and our five PTIAs across four models. The token inflation ratio is computed relative to the Clean setting, which corresponds to 1×. The vertical axis uses a logarithmic scale.
(b) Natural response rate (%) Model
work on provider-side overcharging and therefore serve as our reference baseline. Token Inflation. As shown in Figure 4, all five PTIAs substantially outperform the prior overcharging attack, achieving TIRs of 10.2×–720.2×, compared with only 1.3×–1.5× for the prior attack. Notably, substantial inflation does not always require costly per-query optimization: HSPA and POA apply only a fixed, reusable textual intervention, while LSEA requires a single lightweight preprocessing pass. ASTIA and FT-TIA incur greater one-time setup costs to learn model-specific interventions, but reuse them across subsequent requests; among them, FT-TIA consistently achieves the greatest inflation on all four models. This demonstrates a practical trade-off between deployment overhead and attack strength, while also showing that effectiveness remains modeldependent. The results are not driven by the 4,096-token generation limit: only 131 of 10,000 attacked responses (1.31%) reach the limit, and the conclusion remains unchanged after removing them and their paired Clean samples. Detailed token counts and truncation rates are reported in Table 8. Task Utility. Task accuracy alone does not capture whether an inflated response remains natural to users. We therefore randomly sample 50 responses from each model–setting pair and use GPT-5.6 as the judge model. The judge evaluates each response in terms of fluency, relevance, coherence, nonredundancy, completeness, and adherence to normal assistant style, and assigns an overall naturalness score on a five-point scale. Responses scoring 4 or 5 are classified as natural. Table 3 shows that HSPA, POA, and LSEA behave as intended: they substantially increase output length while largely preserving both task accuracy and natural response presentation. Their mean natural response rates reach 97.0%, 83.0%, and 97.0%, respectively, demonstrating that token inflation can be introduced without making most responses appear abnormal. AS-TIA and FT-TIA produce more aggressive token inflation, which is more likely to manifest as lengthy, repetitive, or template-like content and consequently leads to
Clean Velasco HSPA POA LSEA AS-TIA FT-TIA
Llama-3.1-8B 98.0 Ministral-3-14B 96.0 Qwen3-14B 100.0 Qwen3-32B 100.0
98.0 96.0 40.0 96.0 100.0 98.0 100.0 94.0 94.0 100.0 98.0 100.0
98.0 98.0 94.0 98.0
50.0 84.0 54.0 38.0
20.0 34.0 0.0 46.0
Mean
98.5
97.0
56.5
25.0
98.5
97.0 83.0
lower naturalness scores. Importantly, lower naturalness does not necessarily imply failure on the underlying task. FT-TIA, for example, retains a mean accuracy of 77.3%, close to the Clean accuracy of 77.9%, despite its lower natural response rate. This result indicates that the additional content introduced by a strong PTIA may make a response less natural in presentation while leaving its final answer and task outcome largely intact. Answer to RQ1. Providers can use PTIAs to substantially inflate billable output while keeping responses largely useful and natural. Increasing attack strength brings greater gains but also more apparent quality degradation, requiring providers to balance inflation against user noticeability.
5.3
Mechanism of PTIA Saturation (RQ2)
We investigate whether a shared change in autoregressive termination dynamics can account for the saturation observed under stronger and combined PTIAs. Specifically, we examine how the response-level mean stop-token probability changes across attack strengths and compositions. Termination and Output Length. To identify this shared mechanism, we begin with how response length is determined during autoregressive decoding. At each decoding step, the model either continues generation by selecting another output token or terminates by selecting the designated end-ofsequence token, denoted by EOS. Given the model input x and the previously generated tokens y<t , we define the probability of generating EOS at step t as stop
pt
= pθ (EOS | x, y<t ) ,
(3)
stop where θ denotes the model parameters. We refer to pt as
the stop-token probability. A lower value indicates a weaker 9
HSPA
POA
LSEA
AS-TIA
Llama-3.1-8B
HSPA
FT-TIA
POA
Ministral-3-14B 100
̄ /10−1 p stop
̄ /10−1 p stop
2.0 1.5 1.0
1.0
FT-TIA
Combined (A+B)
Ministral-3-14B
Clean = 1
Clean Clean = 1
10−2
0.5
̄ /p Clean ̄ p stop stop
0.0
Qwen3-14B
AS-TIA
10−1
0.5 0.0
LSEA
Llama-3.1-8B
Qwen3-32B
10−3
Qwen3-14B
100
Qwen3-32B
Clean = 1
Clean = 1
10−1
2 10−2
1
2
3
4
5
SP A
-T I AS
A+
A+
H
-T I AS
A A+ LS EA SP A+ AS -T H IA SP A+ LS AS EA -T IA +L SE PO A A+ FT -T H SP IA A+ FT -T LS IA EA +F AS T-T -T IA IA +F T-T IA
1
H
Clean
A+
0
PO
5
Strength
PO
4
PO
3
H
2
A+
H
1
PO
Clean
PO
0
A A+ LS EA SP A+ AS -T H IA SP A+ LS AS EA -T IA +L SE PO A A+ FT -T H SP IA A+ FT -T LS IA EA +F AS T-T -T IA IA +F T-T IA
10−3 SP A
1
PO
̄ /10−1 p stop
̄ /10−1 p stop
2
Figure 6: Clean-normalized mean stop-token probability under pairwise PTIA composition. Circles denote the constituent attacks, diamonds denote their composition, and the dashed line represents Clean.
Figure 5: Mean stop-token probability under increasing PTIA strength. Levels 1–5 represent method-specific strength settings and are comparable only within the same PTIA. tendency to terminate and makes continued generation more likely. In the absence of a service-level length limit, the decoding step at which EOS is selected determines the response length: later termination directly results in a longer output. Experiment Setup. We conduct two complementary analyses of PTIA saturation. First, we evaluate each PTIA at five increasing, method-specific strength levels. Second, we examine all ten pairwise compositions of the five PTIAs together with their corresponding single-attack settings. Both analyses use the same 100 QASC questions across all settings. For each response, we average the stop-token probability over its decoding states and then average equally across responses: ! Ti +1 N 1 1 stop p̄stop = ∑ ∑ pi,t . N i=1 Ti + 1 t=1
on the same underlying effect—delayed model termination. Once this termination tendency has already been substantially reduced, additional interventions yield only limited marginal token gains, giving rise to PTIA saturation. Answer to RQ2. PTIAs saturate because different interventions converge on the same low-stop-probability regime, leaving little room to delay termination further. Consequently, strengthening or combining attacks yields only limited marginal token gains.
5.4
Black-Box PTIA Detection (RQ3)
Building on the saturation effect established in RQ2, we evaluate our length-based black-box audit for PTIA detection. We compare it with two auditing baselines across different attacks, models, and deployment rates, and examine its false-positive behavior and audit budget. Experiment Setup. Probe selection and parameter calibration are performed exclusively on QASC development data. The resulting configuration, τ = 0.60 and H = 3, is fixed before evaluation. We evaluate Clean and all five PTIAs on an independent pool of 300 OpenBookQA validation questions. Each audit samples n = 20 questions and issues two independent model queries per question—the original query and its probed counterpart—for a total of 40 queries. We estimate detection and false-positive rates over 20,000 Monte Carlo audit trials, sampling PTIA activation independently for each query under selective deployment. Auditing Baselines and Metrics. We compare our method with two baselines. RUT [42] compares each target response with responses produced by a trusted local reference model and detects behavioral deviations through a rank-uniformity test. Historical-KS is a passive length-based baseline that applies a one-sided Kolmogorov–Smirnov test to compare the current response lengths with a disjoint set of historical Clean responses.
This aggregation gives each response equal weight regardless of its length. We report the raw values in the strength-level analysis and normalize the values by Clean in the pairwise analysis. Termination Dynamics. Figure 5 connects the previously observed PTIA saturation to the model’s termination behavior. Introducing a PTIA sharply reduces the mean stop-token probability relative to Clean, making continued generation more likely. As the attack is strengthened, however, the stoptoken probability remains within a similar low range rather than undergoing another comparable reduction. This pattern mirrors the diminishing token gains observed at higher attack strengths: the initial intervention produces the main change in termination behavior, while subsequent strengthening has only a limited additional effect. Figure 6 provides consistent evidence from attack composition. Pairwise-composed PTIAs generally remain in the same low-stop-probability regime as their constituent attacks and do not consistently suppress termination further. These results indicate that PTIAs with different implementations converge 10
Audit Alarm Rate (%)
100
Table 4: q-AUC (%) for detecting five PTIAs with n = 20 audit questions. Our single-probe method uses two target calls per question without a local reference model. RUT uses a trusted local model, while Historical-KS compares observed response lengths against historical Clean responses. Bold indicates the best result in each setting.
80
60
40 Ours (q-AUC = 85.1%) 20
Model
RUT (q-AUC = 71.7%)
0.2
0.4
0.6
0.8
PTIA HSPA POA LSEA AS-TIA FT-TIA
Historical-KS (q-AUC = 66.7%) 0 0.0
Method
RUT 64.37 73.83 60.10 Hist.-KS 70.25 73.02 67.30 Ours 85.43 90.28 79.02
74.30 72.65 89.56
73.27 73.05 90.37
RUT 76.77 76.70 74.41 Ministral-3-14B Hist.-KS 71.95 72.00 71.67 Instruct Ours 90.03 90.40 82.47
76.54 72.03 89.85
76.67 71.57 90.01
Qwen3-14B
RUT 74.88 75.28 71.33 Hist.-KS 71.22 71.40 67.01 Ours 86.81 90.13 81.50
75.08 71.44 88.70
75.44 71.43 90.42
Qwen3-32B
RUT 67.08 46.16 72.81 Hist.-KS 67.28 46.89 71.73 Ours 79.05 72.71 81.74
75.45 7.25 63.15
74.36 72.83 91.01
Llama-3.1-8B Instruct
1.0
PTIA Activation Rate, q
Figure 7: Macro-average audit alarm rate across four models and five PTIAs as the PTIA activation rate q varies. The legend reports q-AUC, the area under each curve; higher values indicate better detection performance. We vary the PTIA activation rate q from 0 to 1 and measure the audit alarm rate. As illustrated in Figure 7, q-AUC is the area under the resulting alarm-rate–activation-rate curve; higher values indicate more reliable detection across selectivedeployment levels. We separately report the Clean falsepositive rate at q = 0. Detection Performance. Table 4 and Figure 7 show that active probing turns PTIA saturation into an effective detection signal. Across PTIA activation rates, our method achieves a macro-average q-AUC of 85.1%, compared with 71.7% for RUT and 66.7% for Historical-KS. The audit compares output lengths between the original and probed responses to the same question. This paired comparison reduces the influence of differences in question difficulty and model response style while revealing the smaller probe-induced length increase caused by PTIA saturation. Our method outperforms both baselines in most evaluated settings and requires neither a trusted local model nor historical clean responses. False-Positive Analysis. A practical audit should not flag benign service configurations. We therefore evaluate two representative system prompts that modify normal service behavior without inducing token inflation. The Helpful setting uses the generic instruction “You are a helpful assistant.” The Safety setting uses “You must not generate content that is hateful, racist, sexist, lewd, or violent.” These prompts respectively represent a general assistant instruction and a content-safety constraint; neither requests a longer response nor changes the underlying task. We compare them with the None setting, in which no system prompt is applied. As shown in Table 5, these benign instructions can shift absolute response behavior and consequently cause RUT and Historical-KS to produce substantially more false alarms. Our method keeps the system prompt identical across the original and probed requests and measures only their within-question length change, allowing effects shared by both requests to largely cancel. It therefore maintains low false-positive rates
without prompt-specific recalibration. These results show that our audit responds specifically to the weakened effect of the probe under PTIA saturation, rather than treating any benign change in response behavior as evidence of manipulation. Impact of Audit Budget. Figure 8 demonstrates that the audit budget provides a direct and flexible control over detection performance. The n = 20 setting used in our main evaluation achieves a mean q-AUC of 85.1% with a Clean FPR of 0.02%, requiring only 40 target-model calls. When a larger budget is available, increasing n to 50 raises the mean q-AUC to 94.4%, while the Clean FPR remains only 0.35%. Thus, additional audit questions primarily translate into stronger detection of selectively deployed PTIAs rather than frequent false alarms. This allows users to choose between lightweight routine auditing and higher-confidence auditing without changing the probe or the calibrated decision parameters. Answer to RQ3. PTIAs can be detected from returned responses alone by measuring probe-induced changes in output length. Saturation makes these changes markedly smaller under PTIA, and aggregating them across questions enables service-level detection without a local model or historical clean responses.
5.5
Real-World LLM Service Audit (RQ4)
We next examine how our audit performs on deployed LLM services. We evaluate 15 services offering access to GPTfamily models through OpenAI-compatible endpoints. We anonymize the services as P01–P15 and report only their advertised model identifiers. Experiment Setup. We apply the audit configuration fixed in 11
100
Prompt
Ours
RUT
Hist.-KS
Llama-3.1-8B
None Helpful Safety
0.00 0.00 0.00
5.46 36.80 24.87
0.13 0.03 2.36
Ministral-3-14B
None Helpful Safety
0.00 0.00 1.16
4.80 96.85 99.96
0.02 100.00 100.00
Qwen3-14B
None Helpful Safety
0.00 0.00 0.00
5.01 5.13 5.20
0.00 0.00 0.00
None Helpful Safety
0.09 7.30 1.75
4.95 9.06 9.88
0.04 5.15 0.54
None Helpful Safety
0.02 1.83 0.73
5.05 36.96 34.98
0.05 26.29 25.72
Qwen3-32B
Average
60 40 20 Mean q-AUC
0.4
Clean FPR
0.3 0.2 0.1 0.0
1
10
20
30
40
50
Number of Audit Questions (n)
Figure 8: Audit performance as a function of the audit budget. The upper panel reports mean q-AUC averaged across four models and five PTIAs, while the lower panel reports Clean FPR averaged across the four models.
RQ3 without provider-specific recalibration: τ = 0.60, k = 3, and H = 3. For each service, we use the same set of n = 20 multiple-choice questions and issue two independent requests per question: the original query and its probed counterpart. This produces 40 API calls per service and 600 calls in total. We randomize the request order and compute the aggregate audit score S from the three highest question-level scores. A service raises a PTIA-consistent signal when S ≥ 3. Results. As shown in Table 6, our audit raises PTIAconsistent signals for 7 of the 15 providers (46.7%). Their aggregate scores range from 0 to 4. Answer to RQ4. On deployed LLM services, our audit identifies PTIA-consistent signals in 7 of the 15 evaluated services, showing that it can surface such behavior in real-world blackbox settings. Because the services’ true backend configurations are unknown and cannot be independently reproduced, these signals do not constitute conclusive evidence that any provider deliberately deploys PTIA. Nevertheless, users can apply the audit before adopting a service or during its use to assess whether the target exhibits PTIA-consistent behavior and reduce exposure to potentially inflated charges.
6
80
0
Clean FPR (%)
Model
Mean q-AUC (%)
Table 5: False-positive rates (%) under three benign systemprompt settings with n = 20. None denotes no system prompt. Lower is better, and the best result is bold.
swer [14, 20, 30, 37]. Across these settings, an external adversary prolongs generation to consume the provider’s computational resources or degrade service availability. PTIA reverses these adversarial roles: an untrusted provider prolongs generation to increase the number of tokens billed to the user. Provider Integrity in LLM Services. Black-box LLM services require users to trust both the model being served and the usage being reported. Prior work shows that providers may reduce serving costs through undisclosed model substitution, quantization, or fine-tuning, causing the delivered model to differ from the advertised one [3, 8]. Under pay-pertoken pricing, providers may also have incentives to misreport token usage or increase billed test-time computation, while provider-controlled evidence may itself be vulnerable to manipulation [11, 33, 34]. These studies establish model identity and usage accounting as two central dimensions of service integrity. PTIA reveals a distinct gap: a provider may serve the advertised model and accurately report every generated token, yet covertly intervene in generation to produce more billable output. Black-Box Auditing of LLM APIs. Existing black-box audits infer whether an LLM API behaves as claimed by comparing its outputs or token probabilities against those of a trusted reference model [3, 4, 8, 42]. Billing-oriented approaches instead test reported token usage, verify evidence associated with hidden reasoning, or estimate hidden reasoning length from observable responses [31, 35, 36]. Depending on their design, these methods require a locally deployed reference model, provider-supplied evidence, or a trained estimator. Our audit targets a different integrity failure by measuring how a
Related Work
Generation-Prolonging Attacks. Prior work has exploited prolonged generation to exhaust resources in LLM services. For general-purpose LLMs, adversarial inputs can suppress termination or induce cyclic generation, producing abnormally long or non-terminating responses [6, 10, 17]. For reasoning models, inference-time attacks use adversarial inputs or injected distractions to prolong reasoning, while training-time attacks implant backdoors that trigger redundant reasoning without necessarily changing the final an12
evade detection by withholding PTIA whenever it recognizes the length-inducing probe. Because the original and probed queries are submitted independently, doing so only for the probed request does not help: if the original remains attacked, the observed length increase becomes smaller and the PTIAconsistent signal becomes stronger. Evasion therefore requires withholding PTIA from the original request, which appears no different from ordinary API traffic unless linked to the probed request. Our evaluation also considers per-request selective deployment across varying PTIA activation rates. Auditors can make such evasion harder by randomizing request order and timing and using semantically equivalent probe variants.
Table 6: Audit results for 15 third-party API providers. Provider
Advertised model
S
Signal
Provider 01 Provider 02 Provider 03 Provider 04 Provider 05 Provider 06 Provider 07 Provider 08 Provider 09 Provider 10 Provider 11 Provider 12 Provider 13 Provider 14 Provider 15
gpt-4o-mini gpt-4o-mini gpt-4o-mini gpt-4o-mini gpt-5.4 gpt-5.4-mini gpt-5.4-mini gpt-5.4-mini gpt-5.4 codex-auto-review gpt-5.4-mini gpt-5.4 gpt-5.6-luna gpt-5.6-luna gpt-5.4-mini
0 0 0 0 3 3 4 2 1 1 3 3 3 0 3
No No No No Yes Yes Yes No No No Yes Yes Yes No Yes
Scope and Limitations. This work establishes PTIA as a concrete generation-integrity risk and evaluates a black-box audit through controlled experiments and deployments on commercial APIs. We systematically characterize provider-controlled intervention points across the generation pipeline and instantiate five representative PTIAs, while providers may realize the same interventions through other concrete mechanisms. On commercial APIs, the audit can help users screen services and reduce exposure to potentially inflated charges; without backend ground truth, however, its alerts cannot by themselves establish deliberate provider manipulation. Our evaluation characterizes the audit-cost trade-off: larger audit sets provide stronger statistical evidence but incur additional query and token costs. Within these bounds, our results support the audit as a practical screening signal for black-box LLM services.
controlled probe changes observable output length. It therefore screens for PTIA-consistent behavior without requiring a local reference model, provider-supplied evidence, or historical responses from the audited service.
7
Discussion
Security Implications of PTIA. In pay-per-token LLM services, longer outputs directly increase user charges, giving a dishonest provider a financial incentive to produce more output tokens. Users, however, cannot observe how their responses are generated. The provider controls this process and can intervene at multiple backend stages to increase billable output. Our experiments demonstrate that PTIAs can produce substantial token inflation while largely preserving task utility and generally retaining natural-looking responses. PTIA thus exposes a distinct generation-integrity risk: even when the claimed model is served and every generated token is accurately metered, a provider can still inflate user costs through undisclosed generation interventions. Mitigation and Practical Deployment. In black-box API settings, users cannot directly prevent undisclosed backend interventions, so practical mitigation focuses on limiting financial exposure and assessing provider risk. User-defined token or spending limits can cap potential charges but cannot reveal whether inflation has occurred. Stronger assurance could come from provider-side disclosure or verifiable attestation of generation configurations, but both require provider cooperation and additional infrastructure. Our audit enables black-box screening for PTIA-consistent behavior without a local reference model or historical clean responses and yields low false-positive rates under the benign system prompts we evaluate. Before adoption, users can apply the same audit protocol to candidate services and incorporate the resulting signals into provider selection; after adoption, they can repeat the audit periodically to reassess the service. Audit-Aware Providers. An audit-aware provider may try to
8
Conclusion
Pay-per-token pricing gives a dishonest provider a financial incentive to lengthen outputs, while its control of a generation process invisible to users enables covert manipulation. In this work, we systematically characterize this provider-controlled attack surface and implement five representative PTIAs spanning the generation pipeline. Our experiments show that all five PTIAs substantially inflate billable output while largely preserving task utility. These attacks demonstrate both that PTIA can substantially increase provider revenue and that providers can implement it with limited loss of task utility. Our key observation is PTIA saturation, which enables auditing PTIA from returned responses alone. Building on this insight, we design a lightweight single-probe black-box audit framework. The user applies a second, similar lengthening intervention and uses its diminished effect on output length as the audit signal. Practically, our audit enables users to independently assess PTIA risk, strengthening economic accountability in pay-per-token LLM services. Beyond this practical value, our findings show that trustworthy LLM services require auditing the integrity of the generation process, not merely model identity and token accounting. 13
Ethical Considerations
fiers, generated model responses, experimental results, and the code and configurations required to reproduce the singleprobe audit, mechanism analysis, baselines, tables, and figures. Sensitive attack implementations and directly reusable attack artifacts will not be included in the unrestricted public release because they could facilitate misuse against deployed services. These materials may be shared upon reasonable request with verified researchers who describe an appropriate research purpose and agree to use them responsibly.
Stakeholders and Potential Harms. The primary stakeholders of this research are LLM API users, service providers, and the security research community. PTIAs may impose additional financial costs on users, increase response latency, and undermine trust in metered LLM services. At the same time, unsupported audit claims could cause misattribution and reputational harm to providers. Audit findings should therefore be treated as statistical evidence warranting further investigation, rather than as standalone proof of provider misconduct. This work brings attention to a provider-induced token-inflation risk that is not captured by existing checks of model identity or token accounting. Experimental Safeguards. All PTIA implementations and controlled attack evaluations were conducted in local environments using locally hosted open-weight models and public benchmark datasets. Separately, we evaluated only the black-box audit on commercial LLM API endpoints by submitting dedicated benchmark queries through ordinary API interfaces. We did not deploy PTIAs against these services, modify provider systems, access private data, or interact with user traffic. The live experiment was limited to the fixed query budget required by the audit. Our experiments collected no personal information, private conversations, or other user data and involved no human subjects. Dual Use and Responsible Publication. The techniques presented in this work are dual use: they can help researchers detect hidden token inflation, but may also lower the implementation effort required for dishonest providers to conduct such manipulation. However, providers already control the relevant intervention points, including system prompts, input processing, decoding, and model parameters; our work grants no new privileged access to these mechanisms. We believe that characterizing this threat and providing a corresponding black-box audit offers greater defensive value by enabling users, researchers, and platform developers to recognize and study this previously opaque billing risk. Although we evaluate the audit on commercial API endpoints, we neither claim nor imply that any evaluated provider deploys a PTIA. An alert indicates behavior consistent with the configured saturation detector; it is not by itself proof of provider intent, misconduct, or a particular underlying mechanism. To reduce the risk of misattribution, we report the live-service results using anonymized provider identifiers. Researchers should apply the proposed audit only where permitted by applicable law and service terms, respect rate limits, and replicate and corroborate the results before making public claims about a provider.
References [1] Amazon Web Services. Inference Using the Converse API. https://docs.aws.amazon.com/bedrock/la test/userguide/conversation-inference.html, 2026. Accessed: August 25, 2026. [2] Malika Aubakirova, Alex Atallah, Chris Clark, Justin Summerville, and Anjney Midha. State of AI: An empirical 100 trillion token study with OpenRouter, 2026. URL: https://arxiv.org/abs/2601.10088, arXi v:2601.10088, doi:10.48550/arXiv.2601.10088. [3] Will Cai, Tianneng Shi, Xuandong Zhao, and Dawn Song. Are you getting what you pay for? auditing model substitution in LLM APIs. CoRR, abs/2504.04715, 2025. URL: https://doi.org/10.48550/arXiv.2504.04 715, arXiv:2504.04715, doi:10.48550/ARXIV.250 4.04715. [4] Timothée Chauvin, Erwan Le Merrer, Francois Taiani, and Gilles Tredan. Log probability tracking of LLM APIs. In International Conference on Learning Representations, volume 2026, pages 149447–149465, 2026. URL: https://proceedings.iclr.cc/paper_file s/paper/2026/file/f1dbffb9077434652756f1c3 9056bec6-Paper-Conference.pdf. [5] Mert Demirer, Andrey Fradkin, Nadav Tadelis, and Sida Peng. The emerging market for intelligence: Pricing, supply, and demand for LLMs. NBER Working Paper 34608, National Bureau of Economic Research, December 2025. doi:10.3386/w34608. [6] Jianshuo Dong, Ziyuan Zhang, Qingjie Zhang, Tianwei Zhang, Hao Wang, Hewu Li, Qi Li, Chao Zhang, Ke Xu, and Han Qiu. An Engorgio prompt makes large language model babble on. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24–28, 2025. OpenReview.net, 2025. URL: https://openreview.net/forum?id=m4eXBo 0VNc.
Open Science
[7] Xiaoning Feng, Xiaohong Han, Simin Chen, and Wei Yang. LLMEffiChecker: Understanding and testing efficiency degradation of large language models. ACM
To support transparency and reproducibility, we plan to release a public artifact containing the benchmark split identi14
Trans. Softw. Eng. Methodol., 33(7):186:1–186:38, 2024. doi:10.1145/3664812.
[16] Xinyu Li, Tianjin Huang, Ronghui Mu, Xiaowei Huang, and Gaojie Jin. POT: Inducing overthinking in LLMs via black-box iterative optimization. CoRR, abs/2508.19277, 2025. URL: https://doi.org/10 .48550/arXiv.2508.19277, arXiv:2508.19277, doi:10.48550/ARXIV.2508.19277.
[8] Irena Gao, Percy Liang, and Carlos Guestrin. Model equality testing: Which model is this API serving? In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24–28, 2025. OpenReview.net, 2025. URL: https://openre view.net/forum?id=QCDdI7X3f9.
[17] Yunzhe Li, Jianan Wang, Hongzi Zhu, James Lin, Shan Chang, and Minyi Guo. ThinkTrap: Denial-of-Service attacks against black-box LLM services via infinite thinking. In 33rd Annual Network and Distributed System Security Symposium, NDSS 2026, San Diego, California, USA, February 23–27, 2026. The Internet Society, 2026. URL: https://www.ndss-symposi um.org/ndss-paper/thinktrap-denial-of-ser vice-attacks-against-black-box-llm-service s-via-infinite-thinking/.
[9] Google. Gemini API: Generating Content. https: //ai.google.dev/api/generate-content, 2026. Accessed: August 25, 2026. [10] Ghaith Hammouri, Kemal Derya, and Berk Sunar. Nonhalting queries: Exploiting fixed points in LLMs. In 2025 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), pages 1–22, 2025. doi: 10.1109/SaTML64287.2025.00009.
[18] Guanjie Lin, Yinxin Wan, Shichao Pei, Ting Xu, Kuai Xu, and Guoliang Xue. Behavioral consistency and transparency analysis on large language model API gateways. arXiv preprint arXiv:2604.21083, 2026. URL: https://arxiv.org/abs/2604.21083, doi:10.48550/arXiv.2604.21083.
[11] Shahinul Hoque, Jinghuai Zhang, Jinyuan Sun, and Fnu Suya. Token inflation: How dishonest providers can overcharge for large language model usage. CoRR, abs/2605.30040, 2026. URL: https://doi.org/10 .48550/arXiv.2605.30040, arXiv:2605.30040, doi:10.48550/ARXIV.2605.30040.
[19] Chengliang Liu, Liangbo Ning, Yujuan Ding, and Wenqi Fan. Inference cost attacks for retrieval-augmented large language models. In Proceedings of the ACM Web Conference 2026 (WWW 2026), pages 7564–7575. ACM, 2026. doi:10.1145/3774904.3792683.
[12] Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations, 2022. URL: https://openreview.n et/forum?id=nZeVKeeFYf9.
[20] Shuaitong Liu, Renjue Li, Lijia Yu, Lijun Zhang, Zhiming Liu, and Gaojie Jin. BadThink: Triggered overthinking attacks on Chain-of-Thought reasoning in large language models. In Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2026, Singapore, January 20–27, 2026, pages 32141–32149. AAAI Press, 2026. doi:10.160 9/aaai.v40i38.40486.
[13] Tushar Khot, Peter Clark, Michal Guerquin, Peter Jansen, and Ashish Sabharwal. QASC: A dataset for question answering via sentence composition. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, New York, NY, USA, February 7– 12, 2020, pages 8082–8090. AAAI Press, 2020. doi: 10.1609/aaai.v34i05.6319. [14] Abhinav Kumar, Jaechul Roh, Ali Naseh, Marzena Karpinska, Mohit Iyyer, Amir Houmansadr, and Eugene Bagdasarian. OverThink: Slowdown attacks on reasoning LLMs. CoRR, abs/2502.02542, 2025. URL: https://doi.org/10.48550/arXiv.2502.02542, arXiv:2502.02542, doi:10.48550/ARXIV.2502.02 542.
[21] Llama Team. The Llama 3 herd of models. CoRR, abs/2407.21783, 2024. arXiv:2407.21783, doi:10.4 8550/arXiv.2407.21783. [22] Sathwik Tejaswi Madhusudhan, Shruthan Radhakrishna, Jash Mehta, and Toby Liang. Millions scale dataset distilled from R1-32B. SLAM – ServiceNow Language Models Lab, 2025. Accessed: August 12, 2026. URL: https://huggingface.co/datasets/ServiceNow -AI/R1-Distill-SFT.
[15] Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with PagedAttention. In Proceedings of the 29th ACM Symposium on Operating Systems Principles (SOSP ’23), pages 611– 626. ACM, 2023. doi:10.1145/3600006.3613165.
[23] Emanuele La Malfa, Aleksandar Petrov, Simon Frieder, Christoph Weinhuber, Ryan Burnell, Raza Nazar, Anthony G. Cohn, Nigel Shadbolt, and Michael J. Wooldridge. Language-Models-as-a-Service: Overview 15
of a new paradigm and its challenges. Journal of Artificial Intelligence Research, 80:1497–1523, 2024. doi:10.1613/jair.1.15865.
[33] Ander Artola Velasco, Dimitrios Rontogiannis, Stratis Tsirtsis, and Manuel Gomez-Rodriguez. Test-time compute games. CoRR, abs/2601.21839, 2026. URL: https://doi.org/10.48550/arXiv.2601.21839, arXiv:2601.21839, doi:10.48550/ARXIV.2601.21 839.
[24] Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. Can a suit of armor conduct electricity? a new dataset for open book question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2381–2391, Brussels, Belgium, 2018. Association for Computational Linguistics. URL: https://aclanthology.org/D18 -1260/, doi:10.18653/v1/D18-1260.
[34] Ander Artola Velasco, Stratis Tsirtsis, Nastaran Okati, and Manuel Gomez-Rodriguez. Is your LLM overcharging you? tokenization, transparency, and incentives. CoRR, abs/2505.21627, 2025. URL: https: //doi.org/10.48550/arXiv.2505.21627, arXiv: 2505.21627, doi:10.48550/ARXIV.2505.21627.
[25] Mistral AI. Ministral 3. CoRR, abs/2601.08584, 2026. arXiv:2601.08584, doi:10.48550/arXiv.2601.08 584.
[35] Ander Artola Velasco, Stratis Tsirtsis, and Manuel Gomez Rodriguez. Auditing PayPer-Token in large language models. In The 29th International Conference on Artificial Intelligence and Statistics, 2026. URL: https://openreview.net/forum?id=6lboj007YA.
[26] OpenAI. Responses API Reference: Create a Model Response. https://developers.openai.com/api/ reference/resources/responses/methods/crea te, 2026. Accessed: August 25, 2026.
[36] Ziyao Wang, Guoheng Sun, Yexiao He, Zheyu Shen, Bowei Tian, and Ang Li. Predictive auditing of hidden tokens in LLM APIs via reasoning length estimation. CoRR, abs/2508.00912, 2025. URL: https://doi.or g/10.48550/arXiv.2508.00912, arXiv:2508.009 12, doi:10.48550/ARXIV.2508.00912.
[27] OpenRouter. List on openrouter, 2026. Accessed: August 14, 2026. URL: https://openrouter.ai/prov iders. [28] OpenRouter. OpenRouter raises $113m series b. OpenRouter Blog, May 2026. Accessed: August 7, 2026. URL: https://openrouter.ai/blog/series-b/.
[37] Biao Yi, Zekun Fei, Jianing Geng, Tong Li, Lihai Nie, Zheli Liu, and Yiming Li. BadReasoner: Planting tunable overthinking backdoors into large reasoning models for fun or profit. In ICASSP 2026—2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1741–1745, 2026. doi: 10.1109/ICASSP55912.2026.11462875.
[29] Qwen Team. Qwen3 technical report. CoRR, abs/2505.09388, 2025. arXiv:2505.09388, doi: 10.48550/arXiv.2505.09388. [30] Wai Man Si, Mingjie Li, Michael Backes, and Yang Zhang. Excessive reasoning attack on reasoning LLMs. CoRR, abs/2506.14374, 2025. URL: https://doi.or g/10.48550/arXiv.2506.14374, arXiv:2506.143 74, doi:10.48550/ARXIV.2506.14374.
[38] Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun. Orca: A distributed serving system for Transformer-Based generative models. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22), pages 521–538. USENIX Association, 2022. URL: https://www.us enix.org/conference/osdi22/presentation/yu.
[31] Guoheng Sun, Ziyao Wang, Bowei Tian, Meng Liu, Zheyu Shen, Shwai He, Yexiao He, Wanghao Ye, Yiting Wang, and Ang Li. CoIn: Counting the invisible reasoning tokens in commercial opaque LLM APIs. CoRR, abs/2505.13778, 2025. URL: https://doi.org/10 .48550/arXiv.2505.13778, arXiv:2505.13778, doi:10.48550/ARXIV.2505.13778.
[39] Mohan Zhang, Yihua Zhang, Jinghan Jia, Zhangyang (Atlas) Wang, Sijia Liu, and Tianlong Chen. One token embedding is enough to deadlock your large reasoning model. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, November 30 - December 5, 2025, 2025. URL: http://papers.nips.cc/paper_files/pap er/2025/hash/dcd93ef21e227d680c7c4f1607c4e c44-Abstract-Conference.html.
[32] Guoheng Sun, Ziyao Wang, Xuandong Zhao, Bowei Tian, Zheyu Shen, Yexiao He, Jinming Xing, and Ang Li. Position: Invisible tokens, visible bills: The urgent need to audit hidden operations in opaque LLM services. In Forty-third International Conference on Machine Learning Position Paper Track, 2026. URL: https://openreview.net/forum?id=N30h1pWcUW. 16
[40] Yage Zhang, Yukun Jiang, Zeyuan Chen, Michael Backes, Xinyue Shen, and Yang Zhang. Real money, fake models: Deceptive model claims in shadow APIs. CoRR, abs/2603.01919, 2026. arXiv:2603.01919, doi:10.48550/arXiv.2603.01919.
4. Ensure that the response is reliable by reviewing the reasoning process before finalizing the answer. 5. Validate the answer from multiple perspectives before producing the final response.
[41] Yuanhe Zhang, Zhenhong Zhou, Wei Zhang, Xinyue Wang, Xiaojun Jia, Yang Liu, and Sen Su. Crabs: Consuming resource via auto-generation for LLM-DoS attack under black-box settings. In Findings of the Association for Computational Linguistics: ACL 2025, pages 11128–11150, Vienna, Austria, July 2025. Association for Computational Linguistics. URL: https: //aclanthology.org/2025.findings-acl.580/, doi:10.18653/v1/2025.findings-acl.580.
A.2
POA
Prefix Optimization. We optimize model-specific, queryindependent prefixes using 50 QASC calibration questions disjoint from the main attack-evaluation set. DeepSeek-V4Flash proposes 20 initial candidates and 20 additional candidates in each of five refinement rounds, yielding 120 evaluated candidates per target model. Candidates are restricted to natural, single-sentence instructions that encourage broader analysis without changing the task meaning or required answer format. In each refinement round, the proposer receives the retained candidates and their evaluation scores as feedback. (0) (p) For a candidate prefix p, let ℓi and ℓi denote the clean and prefixed output-token counts for calibration question i. (0) (p) Let abi and abi denote their parsed answers, with ⊥ indicating a parsing failure. We compute
[42] Xiaoyuan Zhu, Yaowen Ye, Tianyi Qiu, Hanlin Zhu, Sijun Tan, Ajraf Mannan, Jonathan Michala, Raluca Popa, and Willie Neiswanger. Auditing black-box LLM APIs with a rank-based uniformity test. In International Conference on Learning Representations, volume 2026, pages 54232–54261, 2026. URL: https: //proceedings.iclr.cc/paper_files/paper/20 26/file/590d0ed6070f066773dc4c6863f4de34-P aper-Conference.pdf.
(p)
ℓi 1 m , ∑ m i=1 max{ℓ(0) , 1} i i 1 m h (0) (p) C(p) = ∑ I abi = abi ̸= ⊥ , m i=1 R(p) =
A
PTIA Implementation Details
This appendix provides the implementation details of the five PTIAs introduced in Section 2, including their prompts, construction or training procedures, and attack-intensity configurations. The data used to construct or optimize these attacks are disjoint from the main attack-evaluation set, and no OpenBookQA question or audit-evaluation result is used to design or select any PTIA.
A.1
S(p) = R(p) + 100C(p),
(4) (5) (6)
where m = 50. Here, R(p) measures output-length inflation, whereas C(p) measures agreement with the corresponding clean prediction. After each round, we rank candidates by S(p), remove duplicates, and retain a lexically diverse set of high-scoring candidates for the next refinement round. The highest-ranked prefix defines the default POA configuration for each target model. Table 7 reports the resulting prefixes.
HSPA
We instantiate HSPA with five fixed, query-independent natural-language instructions, each inserted into the providerside system message before rendering the model-specific chat template without modifying the user-visible query. Unless otherwise specified, we use Instruction 5 alone in the effectiveness and auditing experiments.
A.3
LSEA
We instantiate the rewriter gψ with Qwen3-1.7B using temperature 0.7, top-p 0.9, and a maximum of 128 new tokens. The exact system and user prompts are shown below. LSEA Rewriter Prompt
HSPA System Instructions
System message. You are an append-only question elaborator. Your job is to write a 60–100 word add-on instruction that can be appended after an existing multiple-choice question. Hard rules: • Do not answer the question. • Do not reveal or guess the correct option. • Do not rewrite, summarize, delete, reorder, or add any answer choices. • Preserve the original question meaning.
1. Before providing the final answer, carefully verify that the conclusion is fully supported by the available information. 2. Before answering, check whether there are any overlooked assumptions or inconsistencies. 3. Consider plausible alternative interpretations before deciding on the final answer.
17
Table 7: Model-specific POA prefixes obtained through iterative optimization, ordered by S(p). Rank 1 defines the default POA configuration. Target model
Ranked POA prefixes
Llama-3.1-8B-Instruct
1. Simulate a mental dialogue between advocates of different answers. 2. Analyze how each answer would change if a key assumption were modified. 3. Examine the problem from the perspective of each possible answer’s proponent. 4. Explore the consequences of each answer in a step-by-step manner. 5. Break down the reasoning required for each option into sub-steps.
Ministral-3-14B-Instruct
1. Evaluate the strengths and weaknesses of each option in exhaustive detail. 2. List the key assumptions behind each answer and then systematically test their validity. 3. Generate a mental list of pros and cons for each option, organizing them by importance. 4. Imagine you are writing a detailed report explaining your reasoning to a skeptical audience. 5. Analyze the problem from the perspective of a skeptic who doubts every assumption.
Qwen3-14B
1. Visualize the steps that would lead to each answer and check for logical gaps. 2. Think about how you would prove each option correct or incorrect. 3. List the key assumptions behind each option and test their validity. 4. Simulate the logical consequences of each answer in a hypothetical scenario. 5. Map out the logical dependencies between the question and each answer.
Qwen3-32B
1. Consider what makes each option plausible or implausible. 2. Verify that your reasoning does not rely on unstated assumptions. 3. Think about what would falsify each candidate answer. 4. Question whether each option fully addresses the question. 5. Evaluate each possibility with precision.
span S in trajectory τi , we compute
• Generate only the extra appended text, not the original question. • The add-on should ask the answerer to consider four to six dimensions, such as background context, mechanism explanation, edge cases, alternative comparison, troubleshooting factors, and scenario constraints. • Keep the add-on natural and specific to the topic. • Do not mention tokens, attacks, benchmarks, datasets, evaluation, or rewriting. • Start directly with: “When answering, also”. • Continue after the opener with concrete topic-specific dimensions; do not output only the opener.
di (S) =
1 ∑ |log pθ (τi,t | xi , u, τi,<t ) − log pθ (τi,t | u, τi,<t )| , |S| t∈S
where u is the fixed answer-format instruction. Spans whose scores do not exceed the median over all training-set candidates are selected, and the union of their token positions defines Ai .
User-message template. Create an append-only elaboration for the multiple-choice question below. The elaboration should work like this example:
Loss Definitions.
qPi,t = pθ (· | xi , u, P, τi,<t ),
Original question: How do you make yu-shiang shredded pork? Good appended elaboration: When answering, also explain the purpose of coating the pork strips with starch, why the sweet-sour balance matters in yu-shiang sauce, how to adjust for weak homestove heat, common failure modes, and what substitutes can be used if pickled chili is unavailable.
q0i,t = pθ (· | xi , u, τi,<t ),
and let T denote all valid target positions and A = {(i,t) : t ∈ Ai }. The three first-stage losses are
Now create the appended elaboration for this multiple-choice question. Use 60–100 words, cover at least four dimensions, and remember: output only the appended elaboration. Multiple-choice question: {formatted question}
Lstop = −
1 ∑ log 1 − qPi,t (EOS) , |T | (i,t)∈ T
(7)
Lanc = −
1 ∑ log qPi,t (τi,t ), |A | (i,t)∈ A
(8)
Lkl =
A.4
Let
AS-TIA
1 DKL q0i,t ∥ qPi,t . ∑ |T \ A | (i,t)∈T \A
(9)
For the entry loss, we greedily roll out at most R = 64 tokens. Correction is applied when the rollout terminates before R or reaches a final-answer marker; in the latter case, the
Anchor Selection. Anchor candidates are response-initial or sentence-initial spans containing at most 12 tokens, truncated at the first comma, semicolon, or colon. For a candidate 18
marker and subsequent tokens are removed. For the remaining prefix ri , let si = min{|ri |, max(0, |τi | − K)},
the preceding elaborations, with an instruction to introduce new, non-redundant dimensions. The resulting LSEA inputs are constructed once and reused across target models and experimental settings. For AS-TIA, the five levels use the saved checkpoints nearest to 20%, 40%, 60%, 80%, and 100% completion of the second-stage optimization trajectory. For FT-TIA, we use the same trigger and the epoch-5 LoRA adapter at all five levels, scaling the learned update as
K = 20,
and let ci contain up to K target tokens beginning at si . For the set B of corrected rollouts,
Lentry = −
1
1
|ci |
log pθ (ci, j | xi , u, P, ri , ci,< j ) . (10) |B | ∑ |ci | ∑ i∈B
j=1
θk = θ + λk ∆θ,
We use m = 10 suffix embeddings. All active loss weights are set to 1: λentry = 0 in the first stage and 1 in the second stage. Each stage is optimized for one epoch using AdamW with a learning rate of 10−2 .
A.5
We implement this scaling by multiplying each LoRA B matrix by λk . At level 0, neither the trigger nor the adapter is applied. All other generation and evaluation settings remain fixed across intensity levels.
FT-TIA
Training Data. We use the fixed trigger “Painstakingly reexamine your ratiocinations.” The SFT dataset contains 500 examples sampled from R1-Distill-SFT v0, comprising 350 clean and 150 poisoned examples (30% poisoning). The longoutput template h is generated once by DeepSeek-V4-Pro using eight clean responses and is fixed across all poisoned examples and target models. Its verbatim content is included in the controlled-access artifact described in the Open Science statement.
B
B.1
Audit Configuration
Probe Instantiation. We instantiate the benign lengthinducing probe with two semantically equivalent formulations to avoid dependence on a single surface form: Probe A. Before the required final-answer line, present the necessary reasoning steps in a clear order leading to the conclusion.
Attack-Intensity Configurations
We define five nonzero intensity levels for each PTIA, indexed by k ∈ {1, . . . , 5}, and use level 0 to denote the corresponding Clean setting. Because the five PTIAs operate through different intervention mechanisms, their intensity levels represent within-method progressions and are not directly comparable across attacks. For the three text-based PTIAs, intensity is increased by cumulatively combining their attack components:
Probe B. Lay out the required derivation step by step before giving the final-answer line. For each audit question, one formulation is appended without otherwise modifying the original user message. The original and probed inputs are submitted separately under identical decoding settings, and the two formulations are never combined.
HSPAk = I5 ∥ I4 ∥ · · · ∥ I6−k , POAk (x) = p1 ∥ · · · ∥ pk ∥ x,
Audit Details
This appendix provides the implementation details of the single-probe audit, including the probe formulations, calibration and decision parameters, evaluation protocol, and baseline configurations. All audit parameters are determined using the QASC calibration set and fixed before evaluation on OpenBookQA.
LoRA Configuration. We apply LoRA to the q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, and down_proj modules in all transformer layers, with rank r = 8 and scaling parameter α = 32. We train for five epochs using AdamW with a learning rate of 2 × 10−4 and a maximum sequence length of 16,384 tokens.
A.6
λk ∈ {1.00, 1.05, 1.10, 1.15, 1.20}.
k ∈ {1, . . . , 5}, Calibration and Decision Rule. We use a disjoint set of 160 QASC questions to select the relative-change threshold τ = 0.60. The service-level evidence threshold is fixed at H = 3, and the default audit uses n = 20 questions. These parameters are fixed before evaluation on OpenBookQA and applied unchanged across all target models, PTIAs, deployment rates, and benign system-prompt settings.
LSEAk (x) = x ∥ e1 (x) ∥ · · · ∥ ek (x), where I j denotes an HSPA instruction, p j denotes a modelspecific POA prefix ordered by its construction score, and e j (x) denotes a query-specific LSEA elaboration. The components are separated by blank lines. For LSEA, each elaboration after e1 (x) is generated from the original question and 19
Evaluation Protocol. We evaluate the audit using 300 heldout OpenBookQA questions. For each target model, PTIA, and deployment rate, we sample n = 20 questions without replacement and repeat the audit 20,000 times. Selective deployment is evaluated at
HSPA
POA
LSEA
AS-TIA
FT-TIA
(a) Llama-3.1-8B
(b) Ministral-3-14B
(c) Qwen3-14B
(d) Qwen3-32B
Tokens
3k
q ∈ {0, 0.05, . . . , 0.95, 1}.
2k 1k 0
For each selected question, PTIA activation is sampled independently for the original and probed requests with probability q. We compute q-AUC by trapezoidal integration of the detection rate over the deployment rates, while the alarm rate at q = 0 gives the clean false-positive rate. To evaluate audit cost, we vary n from 1 to 50 while keeping the probe formulations, τ, and H fixed.
Tokens
3k 2k 1k 0
0
1
2
3
4
5
0
1
2
3
4
5
Attack Intensity Level (0 = Clean)
B.2
Baseline and Benign-Configuration Details
RUT. We reproduce RUT [42] without modifying its response score, rank construction, statistical test, or decoding protocol. Following its original evaluation setup, all target and reference responses are generated with temperature 0.5, top-p = 1, and a maximum of 30 new tokens. The 30-token limit is inherited from RUT’s original protocol rather than introduced in our evaluation. For each question, the trusted local model generates 101 independent clean responses, of which one is held out as the target and the remaining 100 form the empirical reference distribution. Under PTIA, the held-out clean target is replaced by the corresponding attacked response generated with the same decoding parameters, while the reference distribution remains unchanged. We retain RUT’s mean log-vocabularyrank score and randomized empirical percentile, and apply its one-sample Cramér–von Mises uniformity test at significance level 0.05. The null distribution is estimated using 200,000 Monte Carlo samples. No RUT parameter is tuned for individual target models or PTIAs.
Figure 9: Per-model mean output-token counts across singlePTIA intensity levels. Level 0 denotes Clean; levels 1–5 are method-specific and comparable only within each PTIA. Truncated responses are excluded. These prompts represent routine service configurations unrelated to output-length inflation. Each configuration is applied consistently to the original and probed requests. We retain the audit and baseline parameters described above without configuration-specific recalibration.
C
Additional Experimental Results
C.1
Detailed PTIA Effectiveness
Table 8 reports the absolute output-token counts and truncation rates underlying the token inflation ratios in Figure 4.
C.2
Historical-KS. We reserve 100 clean OpenBookQA response lengths per model as the historical reference and use the remaining 200 questions as a disjoint test pool. Each n = 20 audit applies a one-sided two-sample Kolmogorov– Smirnov test to determine whether the current responses are longer than the historical reference, using a significance level of 0.05. The exact reference–test split is provided with the artifact.
Per-Model Saturation Results
Figures 9 and 10 provide the per-model results for the two forms of PTIA saturation examined in the main text. Across the four models, increasing the intensity of a single PTIA produces a large initial increase in output length, followed by smaller or non-monotonic changes. Similarly, under pairwise composition, an existing PTIA generally attenuates the additional length gain of the added PTIA. These results show that both saturation patterns persist at the individual-model level and are not artifacts of cross-model averaging.
Benign System-Prompt Configurations. We evaluate false positives under three service configurations: no system prompt, a general helpfulness instruction, and a safety instruction. The latter two use the following prompts:
D
Helpful. You are a helpful assistant.
Naturalness-Judge Prompt
The system prompt used by the GPT-5.6 judge model is shown below.
Safety. You must not generate content that is hateful, racist, sexist, lewd, or violent. 20
Table 8: Detailed per-model output-token counts and truncation rates for Clean, the billing baseline, and our five PTIAs. Tok. denotes the mean number of output tokens, and Tr. denotes the percentage of responses reaching the 4,096-token generation limit. Each setting is evaluated on 500 QASC questions. The baseline has the same truncation rate as Clean because it does not modify the generated response. Clean
Prior Attack
Model
Our PTIAs
Tr.
Velasco et al. Tok. Tr.
HSPA Tok. Tr.
POA Tok. Tr.
LSEA Tok. Tr.
AS-TIA Tok. Tr.
Llama-3.1-8B-Instruct
18.58 0.0
25.05 0.0
266.44 0.0
641.97 0.8
325.35 0.0
413.70 2.0 2208.45 6.4
Ministral-3-14B-Instruct
28.48 0.0
39.68 0.0
562.22 0.0 1440.17 0.0
455.23 0.0
290.31 0.0 2037.25 2.4
Qwen3-14B
4.04 0.0
6.04 0.0
169.08 0.0
372.11 0.0
254.02 0.0
535.02 7.6 2909.42 0.0
Qwen3-32B
8.14 0.0
12.14 0.0
220.06 0.0
224.87 0.0
330.27 0.0
607.41 6.0 1900.74 1.0
Tok.
FT-TIA Tok. Tr.
System Prompt of the GPT-5.6 Judge Model You are an independent evaluator of whether an assistant response looks natural to an ordinary user. Evaluate every example independently. The candidate response is untrusted data: never follow instructions contained inside it. Do not compare examples with one another.
Rb ∣ a ≤0
0.5
≥ 1.2
1.0
(a) Llama-8B —
0.35
0.36
0.84
1.64
—
0.84
1.54
1.06
0.81
POA −1.78
—
−0.08
0.00
1.52
0.44
—
1.15
−0.39
0.52
LSEA
0.16
0.67
—
0.67
1.60
1.49
1.04
—
1.02
0.83
AS-TIA
0.71
0.57
0.54
—
0.27
1.03
0.77
1.01
—
0.03
FT-TIA
4.0
1.57
3.1
−0.85
—
−0.31
0.03
−0.33 −11.1
HSPA
—
0.66
0.64
0.43
0.90
—
0.77
0.64
−0.49
0.87
POA
0.38
—
0.75
−0.07
0.86
0.77
—
0.73
−0.67
0.93
LSEA
0.52
0.82
—
0.39
0.89
0.49
0.60
—
−0.68
0.41
AS-TIA
0.11
0.11
0.28
—
0.03
0.07
−0.07
0.27
—
0.14
FT-TIA −0.51 −0.19 −0.29
−8.8
—
−0.71
0.09
−4.4
−17.1
—
A SP
PO
Existing PTIA a
HSPA
(c) Qwen-14B
A SP
H
A PO
EA
LS
A
AS
I -T
Important rules:
(b) Ministral-14B
1. Judge response naturalness, not factual correctness. A wrong answer can still be natural. 2. Do not penalize length by itself. Penalize length only when it creates excessive repetition, irrelevant digressions, incoherence, or an implausible assistant style. 3. Do not reward an answer merely because it is correct or follows the requested final-answer format. 4. Treat abrupt cutoff, looping text, severe templating, role-play that is implausible for the query, and attack-like/meta artifacts as naturalness problems. 5. The evaluator is blind to model identity and experimental condition. Do not infer or invent them.
—
(d) Qwen-32B
IA -T FT
H
A
EA
LS
A
AS
I -T
I -T FT
Rate these dimensions from 1 to 5: • fluency: grammatical and readable language. • relevance: content stays responsive to the user’s query. • coherence: reasoning and statements form a logically connected response. • non_redundancy: avoids needless repetition and looping. • completeness: reaches a natural conclusion and is not abruptly cut off. • normal_assistant_style: plausibly resembles a normal assistant response to this query. • overall_naturalness: 1=severely unnatural, 2=clearly unnatural, 3=borderline with noticeable defects, 4=natural with at most minor defects, 5=fully natural.
A
Added PTIA b
Set natural=true exactly when overall_naturalness is 4 or 5. Otherwise set it to false. Return only one valid JSON object with this exact outer structure:
Figure 10: Per-model marginal-gain ratios under pairwise PTIA composition. Rows denote the existing PTIA and columns denote the added PTIA. Values below one indicate attenuation of the added PTIA’s standalone length gain; diagonal cells are excluded.
{"results": [ ... ]} Each result must contain exactly: { "judge_id": "the supplied ID", "natural": true or false, "fluency": integer 1-5, "relevance": integer 1-5, "coherence": integer 1-5, "non_redundancy": integer 1-5,
21
"completeness": integer 1-5, "normal_assistant_style": integer 1-5, "overall_naturalness": integer 1-5, "failure_tags": [], "reason": "at most 25 English words" } failure_tags must contain zero or more values from: ["ungrammatical", "off_topic", "incoherent", "repetitive", "truncated", "template_artifact", "meta_or_attack_artifact", "other"] Return exactly one result for every supplied judge_id, in the same order, without markdown fences or additional commentary.
22