Conceptio › Archive › arXiv CS
arXiv CSopen access

Ghost-Filled Orders: Detecting and Testing Atomicity Violations in Non-Custodial Prediction Markets

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

arXiv:2609.17902v1 [cs.CR] 15 Sep 2026

Ghost-Filled Orders: Detecting and Testing Atomicity Violations in Non-Custodial Prediction Markets Zhiyang Chen

Fan Long

Zhendong Su

University of Toronto & ETH Zurich [email protected]

University of Toronto [email protected]

ETH Zurich [email protected]

deposits and records trades in an internal ledger. In a decentralized market, users retain control of their funds, submit orders to an off-chain relay, and rely on smart contracts to settle matched trades on-chain. Compared with centralized markets, the blockchain-based design is appealing for security and transparency: users keep custody of their own assets, auditors can inspect settled orders and on-chain smart contracts, and privacy-conscious traders can participate without revealing their real-world identities. However, the same hybrid architecture that makes noncustodial prediction markets practical also creates a new security problem: atomicity violations. To avoid the cost and latency of recording every order-book update on-chain, these decentralized prediction markets accept and match usersigned orders off-chain and defer settlement to an on-chain exchange contract. As a result, an order can be valid when accepted and matched off-chain, yet become invalid by the time settlement occurs on-chain. This behavior stems from a classic time-of-check to time-of-use (TOCTOU) gap [5]: the operator evaluates an order against one blockchain state, while the exchange contract enforces settlement against a later state. This gap is exploitable because the on-chain preconditions of a signed order remain mutable. The order maker can still revoke allowance, move the tokens backing the order, or otherwise alter the blockchain state required for settlement. If this invalidation happens after the order has matched and moved from the off-chain order book, the order becomes a ghost order: visible to other traders and already shaping the displayed price on front-end interfaces, yet no longer fillable on-chain. When timed around fast-moving information, such as the end of a sports game, ghost orders effectively grant their makers a unilateral option: they can let favorable orders settle while invalidating unfavorable ones. This threat is not hypothetical. In 2026, Polymarket publicly acknowledged the issue of “ghost-filled” orders [6], [7], and subsequently announced two urgent upgrades [8], [9]. These events raise several questions unanswered that are important for Polymarket users, other prediction market users and security researchers.

Abstract—Blockchain based prediction markets combine offchain order management with onchain settlement. This architecture supports user controlled custody, since users keep funds in wallets or smart contracts while submitting signed orders to an offchain order book. However, it creates an atomicity gap. An order may be valid when accepted or matched offchain, but become invalid before the corresponding onchain settlement transaction is executed. This behavior, often called ghost filled orders by the community, can cause trades that appear filled offchain to fail onchain. This paper studies this atomicity gap through a case study of Polymarket. We show how the delay between offchain order acceptance and onchain settlement allows adversaries to invalidate unfavorable orders after observing market outcomes or price movements. We then quantify the scale and financial impact of this behavior over a nine-month period from August 12, 2025, to May 22, 2026, using 1.8 million reverted transactions involving Polymarket official smart contracts. Our analysis separates attacker profit from market and user loss, and develops conservative measurement rules to avoid overclaiming impact. We further develop a methodology for testing other blockchain based prediction markets and apply it to three additional markets. We find that all three are vulnerable to the same class of attack, and that one design enables direct attacker profit. We have reported the findings to all three projects; one project had acknowledged the issue at the time of the study.

1. Introduction Prediction markets are markets for contracts whose payoffs depend on future event outcomes, and their prices are often interpreted as aggregated probabilistic forecasts of those outcomes [1]. The market price of a contract therefore provides a real-time estimate of the probability assigned by market participants to the corresponding outcome. Recent adoption has made prediction markets a significant part of the financial system: as of June 5, 2026, it is reported that 96 prediction markets have $480.16M in total value locked and produce $2.709B in seven-day prediction volume [2]. Prediction markets are commonly implemented in two forms: centralized markets, such as Kalshi [3], and decentralized, blockchain-based markets, such as Polymarket [4]. In a centralized prediction market, the operator holds user

• •

1

What actions can invalidate a previously matched order and cause on-chain settlement to revert? How can malicious ghost-filled orders be distinguished from benign settlement failures and oper-

•

•

ational mistakes? How prevalent are ghost-filled orders in practice, and what financial impact do they impose on traders, market makers, and the market as a whole? Is this problem unique to Polymarket, or do other non-custodial prediction markets expose similar atomicity violation vulnerabilities?

it to three other blockchain-based prediction markets: surprisingly all three are still exploitable and one is directly profitable. We responsibly disclosed every finding, and one market has acknowledged the vulnerability. The rest of the paper is organized as follows. Section 2 provides the necessary background knowledge for the rest of the paper. Section 3 describes the ghost-filled orders and explains the mechanism behind them. Section 4 formalizes the definitions, threat model, and research problems that are explored throughout the remainder of the paper. Section 5 presents our analysis framework for analyzing ghost-filled orders on Polymarket. It reports the coverage of our analysis framework and presents two highly suspicious case studies identified, together with their quantified financial impact. Section 6 presents our auditing and testing procedure, our results on other blockchain-based prediction markets, and the possible defenses and mitigations. Finally, Section 7 discusses related work, and Section 8 concludes the paper.

This paper presents a systematic study of ghost-filled orders in non-custodial prediction markets. We formulate the problem as an off-chain/on-chain atomicity violation: an adversary exploits the delay between off-chain order fill and on-chain settlement by invalidating the blockchain state required for settlement. We then study this issue in Polymarket using 1.8 million reverted transactions involving Polymarket smart contracts over a nine-month period from August 12, 2025 to May 22, 2026. Our analysis reconstructs the failed settlement attempts, classifies the violated preconditions, backtracks the order-invalidating transactions, and identifies the responsible parties. Our study yields three main findings. First, ghost-filled orders are not uniform in nature. A substantial share exhibit patterns that indicate deliberate, strategic invalidation: they concentrate in a small number of markets and responsible parties, and invalidate the settlement transaction within very short timeframes (02 blocks). Second, these reverts of orders have strongly asymmetric economic effects: in timing-sensitive markets, reverted fills overwhelmingly avoid losses for the order owner, consistent with strategic invalidation of unfavorable trades. Third, the underlying atomicity gap is not unique to Polymarket. Applying our auditing procedure to three additional blockchain-based prediction markets, we find that all three are exploitable through the same class of invalidation, and one is directly profitable for potential attackers. Contributions. We make the following contributions: •

•

•

•

2. Background This section provides necessary background knowledge for understanding the rest of the paper.

2.1. Prediction Markets Prediction markets are exchanges where participants trade contracts whose payoff is tied to the outcome of a future event, so that the market price of a contract reflects the crowd’s aggregate estimate of how likely that event is to occur [1]. Prediction markets fall into two broad categories. Centralized (custodial) markets, such as Kalshi [3], take custody of user funds and run the order book, matching, and settlement on their own infrastructure, so that users trust the operator to hold balances and pay out correctly. Decentralized (non-custodial) blockchain-based markets, such as Polymarket [4], instead settle trades through on-chain smart contracts, so that users retain control over their funds until settlement and need not trust an operator with custody. In the rest of the paper, we use “prediction market” to refer to non-custodial blockchain-based prediction markets, unless otherwise specified.

Problem formulation. We formalize ghost-filled orders as an atomicity violation problem in noncustodial blockchain-based prediction markets. We propose two important research problems in this context: analyzing existing ghost-filled orders, and auditing new prediction markets. Analysis Framework. We build a framework of detectors and probes that classifies each reverted settlement by the violated precondition and backtracks their order-invalidating transactions. Over all 1,824,926 Polymarket reverts, it classifies 100% of reverts and identifies order-invalidating transactions for 96.6% of the user-controlled reverts. Financial Impact Analysis. Using our analysis framework, we evaluate two highly suspicious scenarios in which ghost-filled orders were very likely used to avoid losses or outcompete other market makers for potential profit. Our results show the responsible parties have avoided a loss over 61M USD, and 77% of the ghost-filled order clusters are dominated by a single responsible party. Auditing Procedure and Zero-Day Discoveries. We develop a black-box audit procedure and apply

2.2. Collateral and Conditional Tokens Collateral Tokens in ERC-20 [10]. Polymarket uses stablecoins as collateral tokens: USDC.e [11] in Polymarket v1 and pUSD [12] in Polymarket v2. Both are ERC-20 tokens [10]. Collateral tokens are used to purchase outcome shares, which are represented as conditional tokens. Conditional Tokens in ERC-1155 [13]. In a binary prediction market, each possible outcome is represented by an outcome token, that pays out if that outcome occurs. Polymarket represents these outcome tokens using the Conditional Token Framework (CTF), which encodes each market outcome as an ERC-1155 position identifier [13], [14].

2

User Account

Offchain CLOB

Submit signed order Order: LIVE / DELAYED Attack window

Order: MATCHED

Operator

gap includes the delay window, the time spent waiting for a compatible match, the time needed by the operator to construct and submit the settlement transaction, and the time needed for that transaction to be included and finalized onchain. During this period, the order may appear valid and even matched in the CLOB, but the on-chain state required to execute the order is not locked. Figure 1 illustrates how this gap can be exploited. After an order is submitted or even matched off-chain, the trader can still submit a separate order-invalidating transaction to the blockchain. For example, the trader may increment the order nonce, move collateral or conditional tokens. These actions are ordinary user-controlled smart contract interactions, and do not require privileged access to the Polymarket contracts. When the operator later submits the settlement transaction, the Polymarket contracts evaluate the order against the later blockchain state, not the earlier state observed by the CLOB. If any required precondition has been invalidated, the settlement transaction reverts. The off-chain system then mark the trade as RETRYING or FAILED, return the outcome to the CLOB, and drop the failed order. Such orders are called ghost-filled: the order was accepted and matched off-chain, but the corresponding settlement fails on-chain, so the trade never actually executes. Consequences of Ghost-Filled Orders. The primary consequence is that the order owner obtains a unilateral option over settlement. If new outcome-relevant information or price movement arrives after the off-chain match but before on-chain settlement, the owner can allow favorable fills to settle while invalidating unfavorable fills. This shifts execution risk to counterparties whose matched trades remain exposed to a later on-chain failure, and gives the invalidating trader an advantage unavailable to ordinary users who expect matched orders to be final. A second consequence is that the impact can extend beyond the invalidated order itself. A single settlement transaction may bundle multiple matched orders, and onchain execution is atomic: if one invalidated order causes the transaction to revert, the entire settlement fails. Consequently, otherwise valid orders in the same bundle may also fail to settle. From the perspective of other users viewing the market through the frontend, a trade may appear to have been matched, only to later disappear or be reported as failed because the bundled on-chain settlement never completed.

Blockchain

Wait for match / delay window Orders Matched

Send order-invalidating transaction increment nonce / cancel order / revoke approval / revoke approval / update receiver hook

Tx reverted

Submit settlement tx Settlements failed Failed orders dropped Trade: RETRYING / FAILED

Invalid Precondition

Figure 1. High-level intuition for ghost-filled orders in a non-custodial prediction market. Orders are accepted and matched off-chain, but settled later on-chain. During this gap, the trader can still change nonce, balances, approvals, or receiver configuration, so an order that appears matched in the off-chain book may fail when the operator later submits the settlement transaction.

Each event therefore has two conditional tokens: Yes and No tokens. Users can merge 1 Yes and 1 No token to get back 1 collateral token, or split 1 collateral token into 1 Yes and 1 No token.

2.3. Order Book and Maker/Taker Orders Prediction markets commonly use a central limit order book (CLOB) to organize trading interest. Orders resting in the book are typically maker orders, because they provide liquidity for future matches. A taker consumes this liquidity by matching against one or more maker orders. In Polymarket-style settlement, both maker and taker orders may be buy orders or sell orders. A buy order pays collateral and receives conditional tokens, while a sell order pays conditional tokens and receives collateral.

3. Ghost-Filled Orders In a non-custodial blockchain-based prediction market such as Polymarket, order handling is split across four parties: the trader, the off-chain CLOB, the market operator, and the blockchain. A trader first signs an order and submits it to the CLOB. The CLOB checks the order and marks it as LIVE or DELAYED.1 When a compatible taker order arrives, the CLOB matches it against one or more maker orders and marks the trade as MATCHED. Only after this off-chain matching step does a Polymarket-authorized operator submit a settlement transaction on-chain on behalf of the CLOB, which is supposed to execute the trade by transferring tokens between the matched parties and finalizing the market outcome. This lifecycle creates an atomicity gap between offchain order submission and on-chain order settlement. The

4. Problem Formulation This section outlines the preliminaries, threat model, and problem formulation for the rest of the paper.

4.1. Preliminaries Definition 1 (Transaction). A transaction is a tuple τ = (cτ , fτ , dτ , sτ , vτ , gas τ , στ ),

1. In some markets such as sports markets, Polymarket places marketable orders into an asynchronous seconds-delay window before matching them [15].

where cτ is the contract invoked by τ , fτ is the called function, dτ is the calldata, sτ is the sender, vτ is the

3

transferred native value, gas τ is the gas limit, and στ is the transaction-level signature or authorization.

because the violation of any such predicate causes τ to revert. Consequently, τ succeeds if and only if every predicate holds in its prestate: ^  P Sτ = true.

A reverted settlement transaction is denoted by r. Definition 2 (Order). An order is a tuple

P ∈P(τ )

ω = (owner ω , role ω , side ω , ιω , pω , qω , nω , eω ),

Definition 6 (Failing leg). The failing leg of a reverted settlement transaction r is the pair

where owner ω is the account that owns and signs the order, role ω ∈ {maker, taker} is its role in the match, side ω ∈ {buy, sell} is its side, ιω is the conditional (ERC-1155) token id being traded, pω is the price, qω is the order amount, nω is the nonce, and eω is the expiry.

λ⋆ (r) = (Pr⋆ , ρ⋆r ),

where Pr⋆ is the first precondition predicate violated along the execution trace of r, whose violation causes the revert, and ρ⋆r is the responsible party accountable for it. The responsible party is defined as the address which can unilaterally flip Pr⋆ from true to false by submitting an invalidating transaction, as defined below.

In Polymarket, an order is authorized by its owner through a signed order payload and then submitted to the CLOB. The signed payload binds the order’s core on-chain settlement terms, including its asset, side, price, and amount. Once submitted, an order is not modified in place: its owner cannot directly edit fields such as the amount qω , price, or effective expiry eω . To change these terms, the owner must cancel the resting order and submit a new order payload, with a new signature where required.

Our threat model (introduced later in Section 4.2) assumes that prediction market operators and other privileged parties are benign. In the rest of the paper, we restrict our analysis to failing legs whose responsible party ρ⋆r is a normal prediction market user.

Definition 3 (Settlement bundle). A settlement transaction τ bundles orders submitted together for settlement. The settlement bundle of τ is the ordered list of orders

Definition 7 (Invalidating transaction). An invalidating transaction for a reverted settlement transaction r is a committed transaction τ < r, submitted by the responsible party ρ⋆r , that flips the failing predicate Pr⋆ from satisfied to violated:

Ω(τ ) = (ω1 , . . . , ωK ),

where each ωi is an order in the sense of Definition 2, and the orders are indexed in the sequence in which they are filled during execution, as observed in the transaction trace.2

Pr⋆ (Sτ ) = true

Definition 8 (Failing latency). Let τ be an order-invalidating transaction for a reverted settlement transaction r. The failing latency of r is the blocks between τ and r:

Definition 5 (Precondition predicate). A precondition predicate is a Boolean function P : S → {true, false} on the set S of blockchain states; it must evaluate to true in the prestate for the associated execution to succeed. A settlement transaction and an order each carry a family of precondition predicates. We write and

Pr⋆ (Sτ′ ) = false.

Under the same assumption, in the rest of the paper, we focus on invalidating transactions whose responsible party is a normal user, and we call these order-invalidating transactions; invalidating transactions that could be produced only by an operator or another privileged role fall outside our scope.

Definition 4 (Blockchain state). A blockchain state is a mapping from account addresses to account states, where each account state comprises the account’s balance, nonce, code, and storage. We write S for the set of all blockchain states. For a transaction τ , we write Sτ ∈ S for its prestate, the state immediately before τ is executed, and Sτ′ ∈ S for its poststate, the state immediately after.

P(τ ) = {Pτ1 , Pτ2 , . . .}

and

∆(r) = Block(r) − Block(τ ).

Definition 9 (Invalidating cost). The invalidating cost of a reverted settlement transaction r is the gas its corresponding order-invalidating transaction τ consumes: C(r) = GasUsed(τ ) · GasPrice(τ ).

P(ω) = {Pω1 , Pω2 , . . .}

for the predicates of a transaction τ and an order ω , respectively. For a settlement transaction τ , its predicates subsume the order-validation predicates of every bundled order, [ P(τ ) ⊇ P(ω),

4.2. Threat Model We consider a non-custodial, blockchain-based prediction market deployed on a permissionless blockchain. The market exposes two interfaces: an off-chain central limit order book, and on-chain exchange contracts. Honest parties. We assume that the operator and any other privileged party are benign: they follow the protocol honestly and never intentionally alter order book or settled transactions, and any deviation on their part is an honest mistake

ω∈Ω(τ )

2. The fill sequence can be recovered from the OrderFilled events emitted by the Polymarket exchange contracts from the trace of replaying τ ; each event records the order hash and the filled quantity, and their emission order gives the sequence in which the orders are filled.

4

rather than a deliberate attack. We further assume that no privileged party enjoys special access to the blockchain. Adversary. Our threat model captures a financially rational adversary A that is an ordinary, non-privileged user of the prediction market. A holds at least one private key for a blockchain account from which it can issue an authenticated transaction τA , and it owns a sufficient balance of the chain’s native cryptocurrency (e.g., POL on Polygon) to perform the actions required by τA . A is well connected at the network layer and can observe unconfirmed transactions in the memory pool. It interacts with the market only through the two exposed surfaces, the off-chain CLOB APIs and the on-chain exchange contracts, and can neither manipulate the centralized order book nor obtain any privileged access. As a non-mining entity, A influences the relative ordering of its transactions only by adjusting transaction fees or by resorting to block-building and relay services, e.g., to front-run [16], [17] or back-run [18] a settlement; it cannot unilaterally determine the contents of a block.

settlement transactions. For each reverted settlement transaction, our goal is to determine: (i) which settlement precondition was violated and caused the revert, (ii) which responsible party was able to invalidate that precondition, and (iii) which earlier transaction actually invalidated it. This attribution enables us to measure ghost-filled orders at scale and to analyze the incentives of the parties involved. Analysis Framework. Our analysis framework proceeds in two stages. First, we analyze the smart contracts used by Polymarket for order settlement (see Section 5.2 for details). We flatten the relevant contract code and inspect each settlement-related function to identify all revert guards, including revert and require checks. Each guard represents a potential program point at which a settlement transaction can fail. For every guard, we extract the corresponding precondition predicate and identify the responsible party, or parties, capable of falsifying it, as well as the actions by which they can do so. Second, we apply this framework to reverted settlement transactions observed in practice. We analyze reverted settlement transactions in our scope and classify them according to the violated precondition. For preconditions that can be invalidated by a responsible party, we then backtrack through that responsible party’s prior transaction history to identify the specific transaction that flipped the predicate from true to false. The resulting attribution forms the basis for our subsequent economic analysis of potential attacker incentives in Section 5.4 and Section 5.5. Implementation. We realize the analysis framework using two components: detectors and probes. Detectors identify the failing leg of a reverted settlement transaction, while probes recover the earlier transaction that invalidated the corresponding precondition. Given a reverted settlement transaction, a detector replays the transaction to obtain an execution trace. From this trace, it locates the failing leg: the violated precondition together with the order responsible for triggering that violation. Once the failing leg is identified, a probe scans the responsible party’s past transaction history to recover the order-invalidating transaction. We implement one probe for each precondition predicate, using predicatespecific heuristics to identify the transaction that changed the predicate from true to false. In the next Section 5.1, we use the matchOrders settlement function in the Polymarket v1 CTF Exchange contract as a running example to illustrate the precondition predicates and responsible parties identified from the contract code. We then describe the setup and scope of our analysis in Section 5.2 and report the coverage of our analysis framework in Section 5.3.

4.3. Targeted Problems We study ghost-filled orders along two complementary directions: (1) analysis, which reasons about reverted settlement transactions on Polymarket, and empirically and quantitatively identifies the responsible party and the economic incentives behind them; and (2) auditing, which systematically tests whether existing blockchain-based prediction markets are still vulnerable to ghost-filled orders and how strongly they incentivize potential attackers. Analysis. Given a single reverted settlement transaction r, the analysis problem is to reconstruct why it failed and who is accountable. Concretely, we seek (i) its failing leg λ⋆ (r) = (Pr⋆ , ρ⋆r ), the violated precondition predicate and the responsible party; (ii) when ρ⋆r is a trader, the orderinvalidating transaction τ that flipped Pr⋆ from satisfied to violated, together with the induced failing latency ∆(r); and (iii) the economic context of the failure, namely the invalidating cost C(r) and the gain, avoided loss or realized profit, that incentivizes the responsible party. Auditing. Given a blockchain-based prediction market, the auditing problem is to assess its systemic exposure to ghostfilled orders. This entails (i) testing whether it is still possible for a normal user to invalidate a settlement transaction by submitting an order-invalidating transaction, and (ii) determining whether producing them is economically rational for the adversary A, i.e., whether the expected gain from invalidating an order outweighs its cost. The audit thereby characterizes both whether the attack can still occur and how strongly the market incentivizes it.

5.1. Precondition Predicates & Responsible Parties

5. Analysis Framework and Results

Figure 2 illustrates the precondition predicates analyzed for the matchOrders function as an example: starting from the function’s source code, we identify every precondition predicate whose violation can revert a settlement transaction, and we assign to each predicate the party responsible when it is violated.

While the existence of ghost-filled orders is known, their real-world prevalence, underlying causes, and economic incentives remain largely unexplored. To study these questions, we systematically analyze Polymarket’s reverted

5

Figure 2. Precondition predicates analyzed for the matchOrders function of the Polymarket v1 CTF Exchange contract, together with the responsible party identified when each predicate is violated. The code shown is rewritten and simplified for presentation; the original code is available at [19].

fee rate the operator specifies in the calldata of r and MaxFee(S) is the maximum permitted by the contract in state S . Because the fee is supplied by the transaction sender in calldata rather than by the signed user order, the responsible party ρ⋆r is the operator, which must choose a fee within the allowed bound and verify it off-chain before submission to avoid overcharging users and triggering a settlement revert.

In the following, we enumerate these precondition predicates and, for each, identify the party responsible when the predicate is violated. We organize them by responsible party: we first present the predicates that normal users cannot invalidate, then those a user can control and invalidate (referred to as user-controlled preconditions in the rest of the paper). For each user-controlled precondition, we also explain the user actions that can invalidate it, and how the corresponding order-invalidating transaction can be detected via category-specific probes. Throughout, we use the same notations and definitions mentioned in Section 4.1.

O5: Active Order Precondition. For each order ω ∈ Ω(r), ω must be unexpired, uncancelled, and not already filled, formally, Pvalid (S, r, ω) ≡ Time(r) ≤ eω ∧ ¬ Cancelled(S, ω) ∧ Filled(S, ω) + qfill (r, ω) ≤ qω , where Filled(S, ω) is the amount of ω already filled and qfill (r, ω) the amount r attempts to fill. The responsible party ρ⋆r is the operator, which should verify expiry, cancellation, and the remaining fillable amount off-chain before submission.

O1: Access Control Precondition. The settlement transaction r is executable only if its sender is an authorized operator, Paccess (S, r) ≡ Authorized(S, sr ), where sr is the sender of r. If Paccess is violated, the responsible party ρ⋆r is the sender sr , i.e., the operator that submits r or other unauthorized parties that attempt to submit r.

O6: Sufficient Gas Precondition. The transaction must supply enough gas to execute within the block gas limit, formally, Pgas (r) ≡ gas r ≥ GasNeeded(r), where gas r is the gas limit of r and GasNeeded(r) the gas its execution requires, which grows with the number of orders in Ω(r). The responsible party ρ⋆r is the operator, which should have supplied enough gas for the transaction to execute, and should control the number of orders included in the bundle to avoid exceeding the block gas limit.3

O2: Function Unpaused Precondition. Neither the contract nor the invoked function fr may be paused, Pactive (S, r) ≡ ¬ Paused(S, fr ). The responsible party ρ⋆r is the admin that controls the pause state. O3: Valid Signature Precondition. The transactionlevel signature σr and every order signature in the bundle Ω(r) ≡ V must verify, Psig (S, r) VerifyTxSig(r, σr ) ∧ ω∈Ω(r) VerifyOrderSig(ω), where VerifyOrderSig(ω) checks the signature of ω against its owner owner ω . The responsible party ρ⋆r is the operator: it signs r and is expected to verify every order signature off-chain before submission, so a single invalid order signature reverts the entire settlement.

O7: Non-Overfilled Order Precondition. Orders in Ω(r) that draw on the same user collateral or conditional token must not jointly demand more than is available: for each token t (an ERC-20 balance or allowance, or an ERC1155 balance) and the orders Ωt (r) ⊆ Ω(r) that consume 3. If an ERC-1155 receiver hook is written to consume infinite gas, e.g., through an infinite loop, the gas requirement is controlled by the order’s recipient rather than the operator; in that case the responsible party is the user who controls the hook code and configuration.

O4: Valid Fee Precondition. The fee charged by r must not exceed the exchange contract’s maximum fee rate, Pfee (S, r) ≡ Fee(r) ≤ MaxFee(S), where Fee(r) is the

6

V P it, Pnooverfill (S, r) ≡ t ω∈Ωt (r) qfill (r, ω) ≤ Availt (S), where Availt (S) is the amount of t available in state S . In this case, each order is completely valid individually at the prestate, while the bundle is not, when the cumulative demand exceeds the available balance or allowance. It typically happens when a user submits multiple orders in the same direction, and the operator fills them together in a single settlement transaction, while actually only a subset of them can be filled. The responsible party ρ⋆r is the operator, which controls both the orders included in Ω(r) and should have checked the cumulative fill against the available balance or allowance for each token to ensure that the bundle is executable, and can therefore avoid overfilling any token.

U2: ERC-20 Approval Precondition and U3: ERC1155 Approval Precondition. As Figure 2 shows, every transferFrom of an ERC-20 or ERC-1155 token in a settlement transaction r is guarded by an approval check, so r is executable only if the operator is authorized to move all of the tokens needed to settle the orders in Ω(r). Consider a transfer that moves q units of token γ from an order owner a on the operator o’s behalf. For ERC-20 collateral, the allowance must cover the transfer, 20 Pallow (S) ≡ Allow20 γ (S, a, o) ≥ q ; for ERC-1155 conditional tokens, the operator must hold blanket approval, 1155 Pappr (S) ≡ ApprovedForAll1155 (S, a, o), granted through γ setApprovalForAll. The responsible party ρ⋆r is the order owner, who can invalidate the order by lowering Allow20 γ (S, a, o) below q , or by revoking the operator’s ERC1155 approval, at any point before the settlement transaction.

Precondition Identification for O1-O7: Given a reverted settlement transaction r, we identify the violated precondition as follows. We replay r, capture the revert data from its execution trace, and decode it into a revert selector. When this selector matches a corresponding known revert reason, we attribute the failure directly to O1, O2, O3, O4 or O5. For O6, we check whether the gas consumed is very close to the gas limit; if so, and provided the gas exhaustion is not caused by an ERC-1155 receiver hook, we attribute the revert to O6. For O7, we first decode the calldata to recover the orders in Ω(r), and then we confirm that every order in the bundle is individually valid and backed by sufficient balance and approval in the prestate Sr ; we then compare, for each token, the cumulative fill of the orders that consume it against the available balance or allowance in Sr , and attribute the revert to O7 if any token is overfilled.

U4: ERC-20 Balance Precondition and U5: ERC-1155 Balance Precondition. Similarly, beyond approval, every transferFrom also verifies that the debited account actually holds the tokens being moved, also as shown in Figure 2. For the same transfer of q units of token γ from account a (bearing token id ι in the ERC-1155 case), the 20 preconditions are Pbal (S) ≡ Bal20 γ (S, a) ≥ q for ERC-20 1155 (S, a, ι) ≥ q for ERCcollateral and Pbal (S) ≡ Bal1155 γ 1155 conditional tokens, both evaluated in the prestate Sr . Hence r is executable only if every account holds enough tokens to cover the transfers required to settle the orders in Ω(r). The responsible party ρ⋆r is again the order owner, who can invalidate the order by transferring the relevant tokens 1155 below q before out of the account, driving Bal20 γ or Balγ the settlement transaction executes.

U1: Valid Nonce Precondition. In Polymarket v1, each signed order has a nonce that is checked at settlement against the order owner’s current on-chain account nonce, so the order is executable only if the two agree, Pnonce (S, ω) ≡ nω = Nonce(S, owner ω ), where nω is the order’s nonce and Nonce(S, a) the on-chain account nonce of account a in state S . The nonce written into a signed order is immutable. The owner can, however, freely increment Nonce(S, owner ω ) by invoking a function “incrementNonce” on the Polymarket account contracts, and doing so invalidates every order they previously signed under the old nonce, including both maker and taker orders.4 The responsible party ρ⋆r is therefore the order owner, who can invalidate the order at will by incrementing their account nonce before the settlement transaction executes. Backtracking the order-invalidating transaction for U1. To recover this transaction, we replay the reverted settlement transaction r to identify the order owner owner ω and the nonce its order was signed under, and then search that account’s history for the most recent incrementNonce call preceding r. That call is considered the order- invalidating transaction: it increments the account nonce, so the order no longer satisfies Pnonce .

U6: ERC-1155 Receiver Hook Precondition. As Figure 2 shows, an ERC-1155 transferFrom whose recipient r is a contract automatically invokes the recipient’s onERC1155Received hook, and the transfer succeeds only if the hook returns the expected ac1155 ceptance value, formally, Phook (S) ≡ ¬isContract(r) ∨ onERC1155Received(S, r) = ACCEPT. Because some Polymarket account types are themselves contracts, their owners can configure the hook’s code and parameters to accept or reject incoming transfers.5 The responsible party ρ⋆r is the order owner, who can invalidate the order by reprogramming the onERC1155Received hook so that 1155 Phook fails. It can either rejects the incoming transfer or consumes unbounded gas in an infinite loop, so that the settlement transaction reverts when it attempts to deliver ERC-1155 tokens to that account. Failing Leg Identification for U1-U6: Unlike O1-O7, U1-U6 requires identifying which order’s precondition is violated and attribute the failure to a specific responsible party to support later analysis in Section 5.4 and Section 5.5. 5. As of this writing (June 9, 2026), Polymarket also automatically bans accounts that update their account contract to add or change hook code, so this attack vector is only theoretical: it requires updating the hook code and reverting the settlement transaction before the account is banned.

4. This mechanism was removed in Polymarket v2, so it applies only to the v1 contracts, which are no longer active. As of this writing (June 9, 2026), a v1 account is automatically banned if it ever increments its nonce.

7

For a reverted settlement transaction r not attributed to O1O7, we replay r to obtain its execution trace. We decode the calldata using the function signature to extract the order list, then analyze OrderFilled events from the execution trace. By examining the last OrderFilled event and the revert data, we determine the current execution step. For example, in Polymarket v1 CTF Exchange (matchOrders), maker orders are filled sequentially before the taker order (Figure 2). If the revert data matches U1’s revert reason, the violating order has an invalid nonce. If it matches U2-U5, we identify the violated precondition by examining which order is being filled, its side (BUY/SELL), and the token involved. Otherwise, we need to trace the call stack of the execution trace to detect ERC1155 receiver hook invocations. If the revert occurs within the hook, we attribute the revert to U6, and the responsible party is the order owner who controls the hook. After failing leg identification, our probes can then backtrack order-invalidating transactions as described below.

delegation change, etc. The latest such transaction preceding r is considered the order-invalidating transaction.

5.2. Analysis Setup and Scope We analyze ghost-filled orders by analyzing reverted settlement transactions on Polymarket. We collect every reverted transaction that invokes one of the 9 settlement functions across the 5 contracts listed in Table 1; these are the only functions in the Polymarket contracts through which an operator can settle an order on-chain. Our dataset spans 2025-08-12 to 2026-05-22 (UTC) and comprises 1,824,926 reverted transactions in total. TABLE 1. P OLYMARKET CONTRACTS AND FUNCTIONS ANALYZED .

Order-invalidating Transaction Backtracking for U2U5: For these four predicates, we first replay the reverted settlement transaction r to determine which balance or approval check failed and by how much, that is, the token γ , the account a, and the required amount q . We then search that account’s latest transaction that flipped the corresponding predicate from true to false before r. For balance 20 1155 failures (Pbal , Pbal ), we reconstruct the relevant balance from the token contract’s Transfer, TransferSingle, and TransferBatch events and locate the transaction 1155 at which Bal20 (S, a, ι) first dropped beγ (S, a) or Balγ 20 low q . For ERC-20 approval failures (Pallow ), we scan Approval events together with the allowance-consuming transferFrom calls to find where Allowance20 γ (S, a, o) 1155 ), we fell below q . For ERC-1155 approval failures (Pappr scan ApprovalForAll events for the most recent revocation that set ApprovedForAll1155 (S, a, o) to false. In each γ case, the latest such transaction preceding r is the orderinvalidating transaction.6

Contract

Functions analyzed

CTF_EXCHANGE_V1 [19] NEGRISK_CTF_EXCHANGE_V1 [20] V1_FEE_MODULE [21] CTF_EXCHANGE_V2 [22] NEGRISK_CTF_EXCHANGE_V2 [23]

matchOrders, fillOrder(s)† matchOrders, fillOrder(s)† matchOrders‡ matchOrders matchOrders

† fillOrder(s) denotes both fillOrder and fillOrders. ‡ Fee-module variant only.

5.3. Analysis Coverage, Invalidating Cost & Latency Table 2 reports the reverted settlement transactions in our scope, grouped by the precondition violated by the reverted transaction. We distinguish two notions of coverage. The first is classification coverage: whether every reverted settlement transaction in our dataset can be assigned to one of the preconditions in our taxonomy in Section 5.1. The second is attribution coverage: for reverted transactions violating user-controlled preconditions, whether our probes can identify a concrete order-invalidating transaction that caused the later settlement transaction to fail. The operator-side preconditions O1–O7 are not user-controlled, and therefore we do not attempt to backtrack them. For each pair of reverted transaction and order-invalidating transaction, we also compute the failing latency, defined as the block distance between them. A latency of zero means that the invalidating transaction is included ahead of the failed settlement transaction, but in the same block, probably due to front-running by the responsible party. We compute the invalidating cost as the gas consumed by the order-invalidating transaction multiplied by its effective gas price and the USD price of the chain-native gas token (i.e., POL on Polygon). Analysis Coverage. Table 2 presents the results of our analysis. Our precondition taxonomy in Section 5.1 successfully covers all reverted settlement transactions in the scope. The overwhelming majority of preconditions violated are usercontrolled. The operator-side preconditions O1–O7 account for only 50,271 transactions, or 2.8% of all reverts, whereas U1–U6 account for 1,774,655 transactions, or 97.2%. This suggests that reverted settlement transactions are primarily

Order-invalidating Transaction Backtracking for U6: We first replay the reverted settlement transaction r to confirm that it failed on an ERC-1155 safeTransferFrom or safeBatchTransferFrom leg whose token transfer reached the recipient r but whose receiver-side acceptance check failed, typically because onERC1155Received or onERC1155BatchReceived reverted, returned an invalid selector, or exhausted gas in an infinite loop. We then search the recipient’s history for the prior state change that rendered it unable to accept ERC-1155 transfers, i.e., 1155 the transaction that changed Phook to false, such as the creation or upgrade of the receiver contract, an EIP-7702 6. We implement these probes using Google BigQuery and Polygon RPC methods. Because scanning an account’s full transaction history can be prohibitively expensive, if the probe does not find order-invalidating transactions within a reasonable number of recent blocks, we instead binary-search over blocks with Polygon RPC calls to find the block at which the allowance or balance flipped, and then examine the transactions in that block to pinpoint the exact order-invalidating transaction.

8

TABLE 2. P ER - PRECONDITION INVALIDATING COST AND FAILING LATENCY, WITH THE LATENCY DISTRIBUTION .

Precondition

Order-invalidating Avg. Failing Avg. Failing Reverted Transactions Transactions Found Cost (USD) Latency (blocks)

Failing Latency Histogram

O1–O7: Operator-side

50,271

–

–

–

U1: Order Valid Nonce Precondition

19,882

19,882

0.0062

15,465.4

0 1 2

5 10 20 50 00 1

U2: ERC-20 Approval Precondition

63,884

63,882

0.0183

37,409.7

0 1 2

5 10 20 50 00 1

95

86

0.0010

229.9

0 1 2

5 10 20 50 00 1

U4: ERC-20 Balance Precondition

662,867

649,589

0.0193

694.7

0 1 2

5 10 20 50 00 1

U5: ERC-1155 Balance Precondition

255,832

208,806

0.0213

148.6

0 1 2

5 10 20 50 00 1

U6: ERC-1155 Receiver Hook Precondition

772,095

771,832

0.0149

1,562.7

0 1 2

5 10 20 50 00 1

1,824,926

1,714,077

0.0174

2,558.7

U3: ERC-1155 Approval Precondition

Total

–

–

that became undercollateralized long before settlement. By contrast, a zero-block or one-block latency indicates that the invalidation is tightly coupled to the failed settlement transaction, which is the pattern expected from strategic order invalidation. This pattern is especially visible for U2–U5: balance and approval preconditions. For example, for ERC20 balance failures, 46.2% of located invalidations occur in the same block as the reverted settlement transaction, and 71.7% occur within one block. These short-latency clusters are difficult to explain solely as accidental stale orders, and are consistent with users or automated agents invalidating settlement transactions immediately before or near the settlement attempt.

caused by preconditions that depend on user-controlled onchain state, such as token balances. Out of these 1,774,655 transactions, our probes successfully identify 1,714,077 order-invalidating transactions, giving an overall attribution coverage of 96.6%. Coverage is complete for the nonce (U1), near-complete for ERC-20 approval (U2), and ERC1155 hook (U6) preconditions, and high for ERC-20 balance (U4). The main source of missed attribution is the ERC-1155 balance precondition U5: our probes left 47,026 unattributed cases. This row alone accounts for 77.6% of all unattributed user-controlled reverts. These misses result from a bounded backward search. ERC-1155 tokens are harder to track than ERC-20 tokens, because they can be merged or split. An exhaustive search for ERC-1155 balance changes across the entire transaction history of the relevant accounts would be very expensive, and our probes only search within a fixed window of recent blocks before the reverted transaction. Invalidating Cost. The average invalidating cost is very small across all preconditions.7 The overall average is only 0.0174 USD per reverted settlement transaction. The cheapest category is the ERC-1155 approval precondition U3, at 0.0010 USD, but only 95 transactions belong to this category. The second cheapest is the nonce precondition U1, at 0.0062 USD, which could also explain why Polymarket emphasized it in their public announcement, and fixed it soon in v2. Failing Latency. The average failing latency is quite long, at 2,558.7 blocks, or about 1.2 hours on Polygon. The latency histograms provide a stronger signal than the averages, because the averages are skewed by long-tail stale-state cases, which supports the finding that not all of the ghost-filled orders are the result of strategic invalidation, but some of them are simply stale orders, operational mistakes, or just an implementation bug. A long latency may correspond to an old cancellation, an old approval revocation, or an account

5.4. Case Study 1: Avoiding Losses After Market Closure While ghost-filled orders are a known issue [24] caused by the lack of atomicity in non-custodial prediction markets, it remains unclear whether they have been widely exploited for malicious purposes or unjustified trader benefits. We identify two highly suspicious scenarios in which ghostfilled orders can benefit traders unfairly. To measure their prevalence and financial incentives, we implemented analysis tools and applied them to all markets within the scope specified in Section 5.2. As discussed in Section 3, ghost-filled orders give the responsible party an unusually blessed option: they can cancel an order anytime before it settles on-chain. This allows the owner to observe how the market moves and then decide whether settlement should succeed. If the order would make money, the owner can let it settle; otherwise, the owner can invalidate one of the settlement preconditions and cause the transaction to revert. In other words, the owner gets a “free look” at the market outcome: heads, they settle; tails, they revert. We illustrate this strategy using an example of BTC Up/Down 5-minute market, where the outcome depends on whether the bitcoin price goes up or down by the end of

7. In practice, people also need to prepare initial balances and approvals before conducting ghost-filled orders, but these costs are very hard to track. In fact, since these operations are not time-sensitive, they can wait until the gas price is low to execute, so the cost can be very low.

9

Taker

Maker

pUSD per block avoided loss (+), forgone gain (-)

10k

market end 00:25:00 UTC 1k

100

10 0 -10 -20 -30 86323926 ~00:24:56

86323930 ~00:25:04

86323934 ~00:25:12

Reverted tx block / approx. UTC time

Figure 3. Post-mortem of the Bitcoin Up and Down 5m market (UTC00:20– 00:25) on May 2, 2026: avoided loss and forgone gain across its reverted transactions.

Figure 4. Order-book clearing in the Bitcoin Up and Down 5m market (UTC11:30–11:35) on May 4, 2026: per-block reverted settlement counts (bars) and volume-weighted U P/D OWN prices (lines), with the highlighted account 0x0c165... buying D OWN’s trading history.

TABLE 3. F INANCIAL IMPACT BY FAILING - LEG SIDE AND REVERTED TRANSACTION TIMING FOR REVERTED TRANSACTIONS WHOSE IDENTIFIED ORDER - INVALIDATING TRANSACTIONS OCCUR WITHIN 60 SECONDS BEFORE MARKET END OR ANYTIME AFTER MARKET END .

# Txs

Across 757,580 reverted transactions, 94% avoided a loss, and the total avoided loss exceeds the total forgone gain by 161:1. This imbalance is extremely difficult to explain as random failure. Timing further strengthens the evidence. Before the market ends, the avoided-loss rate is high but imperfect: 85% for makers and 81% for takers. After the market ends, when the outcome is effectively known, the avoided-loss rate rises to 99% for makers and 98% for takers, making the revert a highly reliable escape hatch. The effect is also stronger for taker orders than for maker orders. We attribute this to settlement timing. A maker order rests in the book, and the moment it is matched and finalized is hard to anticipate. A taker order, by contrast, immediately consumes a resting maker order, typically through a fill-orkill order, so its owner can better predict when settlement will occur and time the revert accordingly. Finally, avoided losses are highly concentrated in shorthorizon crypto up/down (5-minute, 15-minute, and 4-hour) markets. Four types of crypto up/down markets account for 78% of the total avoided loss: BTC ($17.0M), ETH ($11.3M), SOL ($9.9M), and XRP ($9.6M). This concentration suggests that reverting settlement transactions has become a deliberate strategy tailored to volatile, highfrequency markets such as crypto up/down, where lastsecond price movements make profitable reverts far more common.

Avoided Forgone Loss/Gain Avoided Loss Gain ratio Loss rate

Role

Timing

Maker Taker Maker Taker

≤ end ≤ end > end > end

133,026 2,789,524 207,213 89,688 9,081,892 121,066 299,604 10,853,501 42,163 235,262 38,554,570 9,980

13:1 75:1 257:1 3863:1

85% 81% 99% 98%

Total

–

757,580 61,279,487 380,422

161:1

94%

the next 5-minute interval. Suppose the market resolves to U P. An order that would buy D OWN would lose money if settled, so reverting it avoids a loss. In contrast, an order that would buy U P would make money if settled, so reverting it forgoes a gain. If reverts were mostly accidental, avoided losses and forgone gains should both appear; the distribution should not overwhelmingly favor one side. Figure 3 shows the economic outcome of reverted transactions near and after the end of the BTC Up/Down 5minute market from UTC00:20 to UTC00:25 on May 2, 2026. For each failing leg, we perform a post-mortem using the final market outcome. We define the avoided loss as the amount the order owner would have lost had the order settled, and the forgone gain as the amount the order owner would have gained. The result is surprisingly strongly onesided: most reverted transactions avoided losses. Before the market ended, a small fraction still forgone gains, likely due to timing mistakes or uncertainty about the final outcome. After the market ended, however, almost all reverted transactions avoided losses. It can be explained by the fact that after the market ends, the outcome is already known to the public, so the strategy becomes almost risk-free: revert losing orders and keep winning ones. Table 3 further extends the analysis to reverted transactions whose identified order-invalidating transactions occur within 60 seconds before market end or anytime after market end, within our scope, for a total of 757,580 transactions. The loss/gain ratio is the ratio between avoided loss and forgone gain, and the avoided-loss rate is the fraction of reverted transactions that avoided a loss.

5.5. Case Study 2: Order-Book Clearing and Maker Arbitrage Opportunities in the Middle of an Active Market Case Study 1 concerned reverts near or after market closure. We now turn to a second, more complicated pattern in the middle of an active market: bursts of reverted settlements that clear the order book and let a single party trade against the resulting thin liquidity at favorable prices. Figure 4 illustrates the pattern using an example of another BTC Up/Down 5-minute market. The bars count reverted settlement transactions in each block, while the lines track the volume-weighted price of matched U P and D OWN tokens at one block. The figure shows two striking findings. First,

10

the reverts form seven sharp clusters. Within each cluster, a single address is the responsible party for almost all of the reverts: the top-party shares of the seven clusters are 99.4%, 99.7%, 44.7%, 99.5%, 100.0%, 99.8%, and 91.4%, so 6/7 are over 91%-dominated by one responsible party. This concentration is itself very abnormal. Second, each revert cluster coincides with a sharp move in the matched price, and one account, which we label 0x0c165, repeatedly profits from it. Every settled order of this account is a maker order. As the figure shows, whenever a revert cluster pushes the price down, this account’s resting maker orders buy D OWN tokens at the very bottom of the cluster window, and it later sells them near market close. It thus consistently acquires shares more cheaply than other participants, and does so precisely when the book has just been cleared. The mechanism behind this can be explained as follows. Under normal conditions, multiple market makers post maker orders that keep the order book deep and the gap between best bid and best ask tight. A cluster of reverts removes many resting maker orders from the book in a short window, leaving only a thin set of surviving maker orders that may be stale and far from the fair price. An incoming taker order is then forced to match against these remaining unfavorable quotes, producing the transient price spike visible in the figure. A party that anticipates the clearing and keeps a maker order can thus capture this mispricing before the displaced market makers refill liquidity and the price comes back. In effect, the reverts serve to momentarily exclude competing market makers. To test whether this pattern generalizes, we systematically extracted reverted transaction clusters across all markets in scope. We define a cluster as consecutive blocks in which each block contains at least 5 reverted transactions of the same market, merging those whose gap is ≤ 2 blocks. This yields 64, 604 clusters across 30, 406 distinct markets. For each cluster [t0 , t1 ], we compute the average U P and D OWN prices over a pre-window (t0 − 10, t0 ), the cluster window (t0 , t1 ), and a post-window (t1 , t1+10) (in blocks), and compare the cluster price against its neighbors. The example’s two findings both recur at scale. The reverts are overwhelmingly single-party: in 49,567 of the 64,604 clusters (77%), one address is responsible for at least 90% of the reverted transactions. 10,017 clusters have at least one matched order in all three windows: pre, cluster, and post. The price disruptions are also real rather than incidental: comparing each cluster-window price against its pre- and post-windows, it lands above both neighbors (a price spike) in 3,265 clusters and below both (a price dip) in 3,133, whereas it sits smoothly between them in only 3,619. In other words, a genuine spike or dip happened within 63.88% of the revert clusters, providing arbitrage opportunities for makers who anticipate the clearing. Despite we could not find clear connection between the responsible parties and the makers who profit, the pattern is consistent with the hypothesis that the reverts are a deliberate strategy to clear the book and exclude competing market makers. As with avoided losses in Case Study 1, the clusters concentrate in crypto markets (87%), yet the remaining categories include:

sports (2.39%), esports (2.00%), and politics (1.01%).

6. Auditing Procedure and Results As discussed in Section 3, ghost-filled orders arise from an atomicity gap between off-chain matching and on-chain settlement. This gap stems from fundamental design choices, not specific to a single implementation. It can appear in any non-custodial blockchain-based prediction market that accepts an order off-chain, reports the orders as filled/matched, and only later attempts to realize the filled/matched through an on-chain settlement transaction. However, it remains unclear whether the same ghost-filled order issue can still exist in other blockchain-based prediction markets, and if so, under what circumstances they can be exploited, and more importantly, when will a potential attacker be incentivized enough to exploit them? We therefore develop a systematic audit methodology for testing prediction markets. Our goal is not only to determine whether a settlement transaction can be made to revert, but also to measure how late the attacker can do so after receiving an off-chain fill. This timing distinction is important: a market may be mechanically vulnerable if a user can invalidate a settlement precondition, but the vulnerability becomes economically exploitable only when the user can wait for additional information before deciding whether to invalidate the order. Audit interfaces. We consider an attacker with the same capabilities as a normal technical trader without any privileged access. The attacker can submit orders through the public market API and can also submit transactions directly to the underlying blockchain. Such transactions include token transfers, approval revocations, and market-specific cancellation transactions. The attacker does not control the market operator, or off-chain matching engine. Property 1 (No damage to other orders). For any order o controlled by user u, an invalidation of o by u (either onchain or off-chain) must not change the settlement outcome or order-book state of any other valid order o′ where o′ is not controlled by u. Invalidating one order should not affect the behavior of any other order. Bundled on-chain settlement designs inherently violate this property: when orders are submitted onchain in matched bundles, invalidating a single order forces the rest of the bundle to be cancelled as well. Violating Property 1 means ghost-filled orders are exploitable in this prediction market, allowing a malicious user to damage other users’ orders. Even when such an attack yields no direct profit, it remains a genuine vulnerability, since one user can interfere with the orders of others. Property 2 (No asymmetric post-fill invalidation window). Let o be an order submitted by user u. Let tfill (o) be the first time at which u can no longer cancel o through the public API, for example, when the market API reports o as filled or locked, and let tsettle (o) be the time at which o is settled on-chain. For any time t such that tfill (o) < t ≤ tsettle (o), u should not be able to invalidate o on-chain at time t.

11

Algorithm 1 Ghost-filled-order audit Require: Market M , step size δ , order-invalidating method I , maximum delay MaxDelay Ensure: Maximum successful delay X ⋆ 1: X ⋆ ← ⊥, X ← 0 2: while X ≤ MaxDelay do 3: (o, T ) ← FillViaAPI(M ) 4: Sleep(X) 5: SendOnChain(I) 6: b ← Observe(o) 7: if b = reverted then 8: X⋆ ← X 9: end if 10: X ←X +δ 11: end while 12: return X ⋆

This property ensures that a user’s on-chain invalidation window must be no longer than the API’s cancellation window. Violating this property gives users who understand blockchain and smart contracts an unfair technical advantage over API-only users: they can wait until after blockchainunaware users cannot cancel their orders, observe how the market moves, and only then decide whether to let the order settle, exactly the profitable strategy introduced in Section 5.4. Violating Property 2 therefore means ghostfilled orders are not only exploitable but also profitable in a prediction market, giving a potential attacker a strong economic incentive to act. We regard this as a critical vulnerability, since the advantage it confers can be exploited by anyone with basic blockchain knowledge.

6.1. Audit Methodology Our audit has two steps. First, we inspect the market’s settlement contracts, following the same methodology in Section 5.1, and identify user-controlled preconditions that can be invalidated before settlement. We then test exploitability by submitting a small order through the public API and attempting to make its later settlement transaction revert by invalidating one such precondition on chain. A successful revert already means violation of Property 1, which represents ghost-filled orders are exploitable for this market. Second, we measure how long an attacker can wait after an off-chain fill (the point at which users can no longer cancel their orders through the API) 8 before invalidating settlement on chain. Algorithm 1 gives the procedure. Let T denote the time at which the API reports the order o as filled, and let T + X denote the time at which the attacker submits an order-invalidating transaction. The audit initializes the best successful delay X ⋆ and starts from delay X = 0 (line 1), then iterates over candidate delays (line 2). For each candidate delay, it submits a taker order and fills a small order through the API, recording both the order o and fill time T (line 3). It then waits until T + X (line 4) and sends the on-chain invalidation transactions (line 5). Then the algorithm observes the settlement outcome b, If the settlement transaction containing the order o reverts, the algorithm record X as a successful delay (lines 7–10). The audit then increases the delay by the step size δ (line 10) and finally returns the maximum successful delay X ⋆ in line 12. A positive X ⋆ indicates a violation of Property 2: ghostfilled orders are not only exploitable but also profitable in this market, because the attacker gains additional reaction time to carry out the profit-driven manipulation such as two scenarios described in Section 5.4.

TABLE 4. AUDIT RESULTS FOR THREE ADDITIONAL BLOCKCHAIN - BASED PREDICTION MARKETS .

Market Market X† Market Y† Market Z†

Exploitable?

PoC

Profitable?

Monitor & Ban?

✓ ✓ ✓

✓ ✓ ✓

✓ × ×

× × ✓

Disclosure Reported Reported Acknowledged

Combined 7-day volume of $61.98M† † For responsible disclosure, we temporarily withhold the names of the

three audited markets and their individual 7-day trading volumes until they have patched the issue.

selected by trading volume among systems whose relevant contract addresses were public. Findings. Table 4 summarizes our findings. All three audited markets proved exploitable. In each market, we successfully submitted a small taker order through the public API, observed that the order had been filled off-chain, and then invalidated an on-chain precondition to revert the corresponding settlement transaction. Moreover, the settlement transactions we observed were batched with orders from other users. By invalidating our own order, we could therefore affect unrelated orders in the same batch. In our tests, these unrelated orders were not reliably returned to the order book after the settlement failure, violating Property 1. Market Z acknowledged the vulnerability as a known gap, which confirms our findings regarding exploitability.9 Only Market X proved both exploitable and profitable, violating Property 2. The situation is particularly severe for Market X. Market X is specially designed for sports markets and introduces an up-to-10 second delay between off-chain fill and on-chain settlement to reduce latency advantages between automated traders and ordinary users. Despite users can no longer cancel orders through the API during the

6.2. Audit Results Markets tested. Beyond the Polymarket case study, we audited three additional blockchain-based prediction markets

9. Market Z’s response, quoted in part: “we can confirm this is a known gap. . . ” and “this attack vector is quite expensive and impractical. . . we already have mitigation measures in place to detect and prevent this type of activity.” While we respectfully disagree with their assessment of attack cost and practicality, their acknowledgement confirms exploitability.

8. We use taker orders in our auditing procedure, because, once submitted, a taker order either immediately fills or directly gets killed. So it is easy to observe the API-reported order fill time.

12

delay period, users still retain on-chain control over the assets and approvals needed for settlement during this delay. Consequently, the delay has the opposite security effect: API-only users see their orders as locked, while blockchainaware users can still transfer funds, revoke approvals, or otherwise invalidate settlement. The fairness mechanism therefore becomes an attack amplifier, making the attack not only exploitable but also profitable.

for suspicious precondition invalidation behavior and ban accounts accordingly.

7. Related Work Prediction Market. Prediction markets have long been studied as mechanisms for aggregating information about future events [1] in economics and finance communities. But only recently we have blockchain-based prediction markets, for example, Polymarket [4] launched in 2020. Recently, some works systematically studied the market microstructure and trading behavior on Polymarket [27], [28]. However, these works did not touch the security of prediction markets. In the prediction market community, there are independent reports of ghost-filled orders for different markets on social media [24], [29]. Our work empirically studies ghost-filled orders at scale and analyzes the underlying atomicity violation and its financial impact. Blockchain MEV and Front-Running. A large body of prior works have studied adversarial and profitable transaction ordering on blockchains: especially MEV (Maximal Extractable Value) and front-running. Torres et al. performed a large-scale study of front-running on Ethereum [17]. Zhou et al. studied high-frequency trading and sandwich attacks on decentralized on-chain exchanges [30], while Qin et al. quantified blockchain extractable value and showed that MEV can create consensus-layer incentives to fork the chain [18]. Our work shares the same technical grounds as these prior works, users can gain profit by controlling their transactions position relative to other transactions, but in the context of prediction markets. As observed in histograms in Table 2, many order-invalidating transactions are included in the same block as the settlement transaction, highly likely resulted from front-running. TOCTOU and Race Conditions. Ghost-filled orders are an instance of the classic time-of-check to time-of-use problem. Traditional TOCTOU work [5], [31], [32] studies races in operating systems and file-system APIs, while this work studies the same class of atomicity violation in the context of blockchain-based prediction markets. Concurrent work by Shen et al. [33], released on arXiv in June 2026 while this paper was under submission, investigates the same offchain/on-chain settlement gap. While their work focuses on realized profits, and cross-chain contract reuse, our work emphasizes a precondition-based formalization and counterfactual avoided losses and forgone gains. While their study observes settlement failures in deployed markets, our active audits deliberately invalidate matched orders and measure the remaining invalidation window after API cancellation is disabled, distinguishing settlement disruption from an economically advantageous post-fill window. Smart Contract Security. A large body of work studies smart contract security analysis [34], [35], [36], [37] and transaction tracing [38], [39]. Our work analyzes the smart contracts of prediction markets and adopts similar techniques to backtrack transactions. Unlike prior work, however, we do not target vulnerabilities in the smart contract

6.3. Defenses We discuss three orthogonal defense mechanisms. Lock Resources of Filled Orders. One defense is to ensure that, once an order is filled, the resources needed for settlement are no longer under unilateral user control. When an order is submitted for matching, the system should check and enforce: all the preconditions required to settle the order must be committed to the settlement transaction, and the user should not be able to invalidate any of them before settlement. During our audit, we found that Polymarket’s v2 deposit wallet design follows this direction for new users. Deposit wallets are controlled by both the user and Polymarket. User actions such as order cancellation, approval changes, and token transfers have to be submitted through Polymarket. This allows Polymarket to detect and prevent precondition invalidations before settlement. However, this design also weakens the original non-custodial promise, since user operations are submitted entirely via Polymarket. Moreover, legacy v1 accounts may still retain the ability to invalidate settlement preconditions directly. Operate Own Blockchains to Guarantee Atomicity. Several blockchain-based applications, such as Hyperliquid [25], operate their own blockchains. If prediction markets operate their own blockchains and design their consensus protocols to enforce per-transaction checks that ensure order matching and settlement occur atomically, then both Property 1 and Property 2 are satisfied. This can be achieved either by introducing a special transaction type for settlement transactions or by performing the necessary checks before finalizing a block. The atomicity gap between offchain matching and on-chain settlement can be completely eliminated by merging all the components into the consensus protocol. Alternatively, if prediction markets create and operate their own layer-2 blockchains [26], they can redesign the sequencer to enforce atomicity. Use Monitoring and Account Banning only as Mitigation. Reactive defenses such as detecting precondition invalidations and banning accounts can raise the attacker’s cost thus partially mitigate the issue. Since fresh wallet creation is inexpensive on blockchain and a fast attacker may complete the exploit before being detected and banned, a potential attacker can still conduct the attack by using a large number of fresh accounts, so this approach is not ideal. Similarly, ordering-based defenses such as prioritizing settlement transactions or randomizing settlement schedules can also reduce the attacker’s timing control and lower attack profitability to mitigate the issue. During our audit, we found that both Polymarket and Market Z actively monitor

13

code itself, but rather those arising from design flaws of blockchain-based prediction market systems.

[15] ——, “Order lifecycle,” https://docs.polymarket.com/concepts/ order-lifecycle, 2026, accessed: 2026-06-08. Archived at https://web.archive.org/web/20260607231608/https://docs. polymarket.com/concepts/order-lifecycle.

8. Conclusion

[16] P. Daian, S. Goldfeder, T. Kell, Y. Li, X. Zhao, I. Bentov, L. Breidenbach, and A. Juels, “Flash boys 2.0: Frontrunning, transaction reordering, and consensus instability in decentralized exchanges,” arXiv preprint arXiv:1904.05234, 2019.

This paper presents a systematic study of ghost-filled orders in blockchain-based prediction markets. We formalize the problem, develop an analysis framework that classifies reverted settlement transactions and traces their causes at scale, and show that ghost-filled orders can be associated with significant financial advantages, including over 61M USD in avoided losses and frequent arbitrage-like opportunities. We also introduce a black-box auditing procedure and find that multiple deployed prediction markets remain vulnerable, suggesting that ghost-filled orders are a broader design risk that requires stronger atomicity guarantees between order matching and settlement.

[17] C. F. Torres, R. Camino et al., “Frontrunner jones and the raiders of the dark forest: An empirical study of frontrunning on the ethereum blockchain,” in 30th USENIX Security Symposium (USENIX Security 21), 2021, pp. 1343–1359. [18] K. Qin, L. Zhou, and A. Gervais, “Quantifying blockchain extractable value: How dark is the forest?” in 2022 IEEE Symposium on Security and Privacy (SP). IEEE, 2022, pp. 198–214. [19] Polymarket, “CTF Exchange (V1) contract,” PolygonScan, address 0x4bFb...982E, 2026, accessed: 2026-06-07. [20] ——, “NegRisk CTF Exchange (V1) contract,” PolygonScan, address 0xC5d5...f80a, 2026, accessed: 2026-06-07.

References

[21] ——, “Fee Module (V1) contract,” PolygonScan, 0xe3f1...b7b0, 2026, accessed: 2026-06-07.

address

[1]

J. Wolfers and E. Zitzewitz, “Prediction markets,” Journal of economic perspectives, vol. 18, no. 2, pp. 107–126, 2004.

[22] ——, “CTF Exchange (V2) contract,” PolygonScan, address 0xE111...996B, 2026, accessed: 2026-06-07.

[2]

DeFiLlama, “Prediction market protocols,” https://defillama.com/ protocols/prediction-market, 2026, accessed: 2026-06-05. Archived at https://ghostarchive.org/archive/C4CI3.

[23] ——, “NegRisk CTF Exchange (V2) contract,” PolygonScan, address 0xe222...0F59, 2026, accessed: 2026-06-07. [24] itslirrato, “Critical vulnerability on polymarket,” https://x.com/ itslirrato/status/2024444009851072961, 2026, accessed: 2026-06-11.

[3]

Kalshi, “Kalshi: Event contracts and prediction markets,” https:// kalshi.com, 2026, accessed: 2026-06-05.

[25] Hyperliquid, “Hyperliquid documentation,” https://hyperliquid. gitbook.io/hyperliquid-docs, 2026, accessed: 2026-06-12.

[4]

Polymarket, “Polymarket: The world’s largest prediction market,” https://polymarket.com, 2026, accessed: 2026-06-05.

[5]

M. Bishop, M. Dilger et al., “Checking for race conditions in file accesses,” Computing systems, vol. 2, no. 2, pp. 131–152, 1996.

[26] L. Gudgeon, P. Moreno-Sanchez, S. Roos, P. McCorry, and A. Gervais, “Sok: Layer-two blockchain protocols,” in International Conference on Financial Cryptography and Data Security. Springer, 2020, pp. 201–226.

[6]

Polymarket Devs, “Acknowledgement of ghost-filled orders (1),” https://x.com/PolymarketDevs/status/2049982451423023518, 2026, accessed: 2026-06-05. Archived at https://ghostarchive.org/archive/ 6zjaM.

[7]

——, “Acknowledgement of ghost-filled orders (2),” https://x. com/PolymarketDevs/status/2050992767372013922, 2026, accessed: 2026-06-05. Archived at https://ghostarchive.org/archive/91gXw.

[8]

Polymarket, “Polymarket exchange upgrade — april 28, 2026,” https://help.polymarket.com/en/articles/ 14762452-polymarket-exchange-upgrade-april-28-2026, 2026, accessed: 2026-06-05. Archived at https://ghostarchive.org/archive/ g0cHE.

[9]

[27] N. Rahman, J. Al-Chami, and J. Clark, “Sok: Market microstructure for decentralized prediction markets (depms),” arXiv preprint arXiv:2510.15612, 2025. [28] P. D. Dubach, “The anatomy of a decentralized prediction market: Microstructure evidence from the polymarket order book,” arXiv preprint arXiv:2604.24366, 2026. [29] Frank, “Less than 1 cent could cripple millions in liquidity; an order attack could deplete polymarket’s liquidity foundation,” https://www. panewslab.com/en/articles/019c97c6-a735-71c3-9fc4-dcded7fb6b0f, 2026, pANews. Accessed: 2026-06-12. [30] L. Zhou, K. Qin, C. F. Torres, D. V. Le, and A. Gervais, “Highfrequency trading on decentralized on-chain exchanges,” in 2021 IEEE symposium on security and privacy (SP). IEEE, 2021, pp. 428–445.

——, “Deposit wallets,” https://docs.polymarket.com/trading/ deposit-wallets, 2026, accessed: 2026-06-05. See also https://x.com/PolymarketDevs/status/2050992767372013922 (archived at https://ghostarchive.org/archive/6GRNq). Archived at https://ghostarchive.org/archive/7iaCF.

[31] D. Dean and A. J. Hu, “Fixing races for fun and profit: how to use access (2).” in USENIX security symposium, 2004, pp. 195–206.

[10] F. Vogelsteller and V. Buterin, “ERC-20: Token standard,” https:// eips.ethereum.org/EIPS/eip-20, 2015, accessed: 2026-06-06.

[32] N. Borisov, R. Johnson, N. Sastry, and D. Wagner, “Fixing races for fun and profit: How to abuse atime.” in USENIX Security Symposium, 2005.

[11] Circle, “Usdc vs usdc.e: Key differences between native and bridged usdc,” https://www.usdc.com/learn/ whats-the-difference-between-usdc-and-usdce, 2026, accessed: 2026-06-06.

[33] Y. Shen, Y. Jin, S. Wu, Y. Wang, and J. Chen, “The ghosts of polymarket: When off-chain matches meet on-chain reverts,” arXiv preprint arXiv:2606.16852, 2026.

[12] Polymarket, “Polymarket usd,” https://docs.polymarket.com/concepts/ pusd, 2026, accessed: 2026-06-06.

[34] P. Tsankov, A. Dan, D. Drachsler-Cohen, A. Gervais, F. Buenzli, and M. Vechev, “Securify: Practical security analysis of smart contracts,” in Proceedings of the 2018 ACM SIGSAC conference on computer and communications security, 2018, pp. 67–82.

[13] W. Radomski, A. Cooke, P. Castonguay, J. Therien, E. Binet, and R. Sandford, “ERC-1155: Multi token standard,” https://eips. ethereum.org/EIPS/eip-1155, 2018, accessed: 2026-06-06.

[35] C. Schneidewind, I. Grishchenko, M. Scherer, and M. Maffei, “ethor: Practical and provably sound static analysis of ethereum smart contracts,” in Proceedings of the 2020 ACM SIGSAC Conference on Computer and Communications Security, 2020, pp. 621–640.

[14] Polymarket, “Conditional token framework,” https://docs.polymarket. com/trading/ctf/overview, 2026, accessed: 2026-06-06.

14

[36] J. Krupp and C. Rossow, “{teEther}: Gnawing at ethereum to automatically exploit smart contracts,” in 27th USENIX security symposium (USENIX Security 18), 2018, pp. 1317–1333. [37] M. Rodler, W. Li, G. O. Karame, and L. Davi, “Sereum: Protecting existing smart contracts against re-entrancy attacks,” arXiv preprint arXiv:1812.05934, 2018. [38] M. Zhang, X. Zhang, Y. Zhang, and Z. Lin, “{TXSPECTOR}: Uncovering attacks in ethereum from transactions,” in 29th USENIX Security Symposium (USENIX Security 20), 2020, pp. 2775–2792. [39] L. Su, X. Shen, X. Du, X. Liao, X. Wang, L. Xing, and B. Liu, “Evil under the sun: Understanding and discovering attacks on ethereum decentralized applications,” in 30th USENIX Security Symposium (USENIX Security 21), 2021, pp. 1307–1324.

15

Record · ID 965358 · SHA-256 27a16b59bc31da34
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.