ConceptioArchivearXiv CS
arXiv CSopen access

The Case for Automated Hyperspecialization: Evidence from SAT

Unknown · arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
softwarearchitecturesoftwareengineeringtesting
software engineering, software architecture, testing

The Case for Automated Hyperspecialization: Evidence from SAT Harrison Green

Claire Le Goues

Fraser Brown

Carnegie Mellon University Pittsburgh, PA, USA [email protected]

Carnegie Mellon University Pittsburgh, PA, USA [email protected]

Carnegie Mellon University Pittsburgh, PA, USA [email protected]

arXiv:2609.14836v1 [cs.SE] 13 Sep 2026

Abstract The software status quo is to use one system to process many different kinds of inputs. In contrast, we propose hyperspecialization: creating new software that is optimized for a single class of inputs. Hyperspecializing manually is anywhere from expensive to impossible. We conjecture that coding agents make automated hyperspecialization cheap, effective, and safe for problems with measurable performance and checkable output. This paper explores one such problem, SAT solving, by synthesizing hundreds of workload-specific SAT solvers at an average cost of $37 each. Our specialists outperform their competition-winning, general-purpose cousins by 5× on average, and by over 10× on a quarter of benchmark families. A general-purpose solver constructed from over a hundred of our prototype hyperspecialists won the SAT track at the 2026 SAT Competition.

SAT formulas span structurally diverse families rooks

cryptography-simon

30 formulas

71 formulas

circuit-minimization 50 formulas

antibandwidth 187 formulas

hamiltonian

pigeonhole principle

1.1k formulas

1

55 formulas

Introduction

Different sub-fields of computer science have staged their own rebellions against generality. The architecture [42, 44], databases [100, 101], networking [11, 21], and operating systems [30, 31, 73] communities have repeatedly discovered that domain-specific software or hardware outperforms its generalist alternative. “Exterminate all operating system abstractions” [30]! The OS should not offer applications foundational primitives like virtual memory, since doing so requires making decisions about, for example, page table design— and “any tradeoff penalizes applications that were not anticipated...by the OS designer” [30]. Instead, the OS should expose almost bare hardware resources, and the application should use (or avoid) the relevant ones. As a result of the specialist OS, the application enjoys significant speedups. Despite these generality rebellions, not every specialist has made it from the academic world to reality; Linux has not been re-written as a library OS. This is not because the general- vs. special-purpose performance gap is overstated, but because building specialized software and hardware is not free. Specialists get deployed when the performance benefit of specialization outweighs the cost of specializing. There are domains where the high cost of specialization is worth it [25]. It’s extremely expensive to tape out a new chip [60]—much more expensive than it is to create new software—yet the performance benefit of GPUs, amortized across all graphics problems, justifies the cost of creating

random-planted-solution 328 formulas

Industrial (58 families)

Crafted (113 families)

Random (15 families)

Mixed (3 families)

Figure 1. Each point represents a SAT formula from 189 benchmark families in this paper, projected using t-SNE over 56 structural features. Colors indicate broad benchmark origin, with shades distinguishing families. Seven family clusters are highlighted and labeled.

those GPUs in the first place [68]. More recently, following a similar cost-benefit analysis, there has been an explosion of interest in (and startups about) custom chips for LLM inference [16, 57, 84]. Compared to hardware, custom software can justify a smaller benefit because it’s cheaper to produce—and the cost is now nosediving towards zero. What once required assembling a team of highly paid programmers now requires only a credit card and a horde of Claudes. The Claudes’ results and the programmers’ results are not necessarily the same, though; while it’s possible that agents could build specialized operating systems, such a task involves huge codebases, informal specifications, and complex evaluations—and there is

Harrison Green, Claire Le Goues, and Fraser Brown

no way to check if the resulting specialist is correct.1 Thus the animating question of this paper: “are there meaningful domains in which (1) specialization improves performance and (2) specialization is plausibly automatable?” We believe that there may be several such domains (§6), but this paper presents an existence proof for one: Boolean Satisfiability (SAT) solving. A SAT solver takes a formula of Boolean variables, their negations, and the operators AND and OR, and decides whether any assignment of true and false to the variables makes the formula itself true. SAT was the first problem shown to be NP-complete [24], meaning that any problem in NP can be efficiently reduced to SAT, and that there is no efficient algorithm for solving SAT unless P = NP. Donald Knuth describes SAT as a “killer app” because it can be used to solve so many (hard!) important problems [62]. Software verification, circuit equivalence checking, package dependency resolution, combinatorial math problems, and more all reduce in whole or part to SAT. These domains have not typically created their own SAT solvers. Instead, since SAT is a hard problem, they rely on existing tools that have implemented sixty years of optimizations, including a long tail of tricks that have been sharpened annually since the inaugural 2002 SAT Competition [53]. Rolling your own solver was (usually) not worth it. We argue that rolling your own solver is now worth it. First, the SAT community already recognizes SAT as a domain in which specialization improves performance. SAT problems of different origins have very different structure [4] (Figure 1), and SAT solvers “may dramatically change their performance depending on the class of...instances they are trying to solve” [4]. There are even theoretical differences between solving strategies. For example, most modern solvers use conflict-driven clause learning (CDCL), which can require exponentially large refutations of XOR constraints [107]. Gauss-Jordan elimination can solve the same constraints in polynomial time but cannot solve other kinds of constraints, since it only reasons about systems of linear equations [43]. Empirically, too, no single solver is best on all workloads: the “virtual best solver” (i.e., an oracle that runs the fastest available solver on a given formula) would outperform the winning solver in every competition to date. Second, SAT solvers are well-suited to being written by machines. They have a simple, well-specified interface, are straightforward to evaluate, and are grounded in a vast public literature that spans algorithms, theory, systems, and formal methods, along with codebases, system descriptions, benchmarks, and competition results; the LLMs have been consuming this literature for years. Most importantly, though, SAT solvers produce a checkable result. A “satisfying assignment” to the formula can be (efficiently) checked by plugging it in; an “unsatisfiable” answer comes with a proof that a small, 1 In fairness to the machines, all of these facts make OSes hard for people to write, too.

formally verified checker can validate independently. Thus, we need never trust the LLM-generated SAT solver, and need only check its results. This paper implements automated hyperspecialization for SAT. Our first contribution is presenting hyperspecialization, a form of program optimization that synthesizes a fresh program for a given workload, and describing the applications for which automated hyperspecialization is tractable (§3). Such problems must handle diverse workloads (so there’s something to specialize for), have measurable objectives (so there’s something to optimize towards), and have checkable results (so there’s no call to trust the LLM author). Our second contribution is a framework for SAT hyperspecialization: we develop specialist SAT solvers for the 189 “families”—i.e., formulas sharing an origin—in the Global Benchmark Database (GBD) [52], the SAT community’s benchmark corpus built from over twenty years of competition submissions. For each family in the database, we randomly select a training set of instances, and provide this set to a coding agent. The agent is prompted to implement (and repeatedly evaluate when necessary) a specialized solver from scratch for the entire family. Many solvers do not use traditional SAT algorithms. Instead, they recover structure that was lost in the encoding, and use a different algorithm to solve the structured problem. For example, one specialist recognizes a family of rook-placement formulas and proves that they’re unsatisfiable: the solver knows about rooks and chessboards, so it can say that satisfying the constraints would require more rooks than the board’s rows fit. Our third contribution is an evaluation of specialist solvers’ performance (§4) that shows the specialists perform better than their general-purpose cousins, for low cost. A prototype of our hyperspecialization approach won the SAT track at the 2026 SAT Competition. Specialists match or beat the perfamily virtual best solver, drawn from four past competitionwinning solvers, on 64% of families, with a geometric mean 12× improvement (without checking) and 5× (with checking); specialists are ≥ 100× faster for 27% of families (without checking) and 14% of families (with checking) (§4.2). A specialist costs on average $37 and takes one and a half hours to produce; the fastest arrives in twenty-eight minutes for $11 and a 1.7× improvement (§4.3). Finally, we discuss the implications of SAT hyperspecialization for downstream consumers, and broader opportunities for hyperspecialization (§6)—from SMT solving to compression to compilation and, perhaps, beyond.

2

Related work

In this section, we position our work with respect to prior literature on software synthesis for optimization. We first discuss approaches that bound the space of candidate programs so that correctness follows by search space construction (§2.1). We then discuss recent LLM-based approaches,

The Case for Automated Hyperspecialization: Evidence from SAT

which search a much broader space of program alternatives but with relaxed correctness guarantees (§2.2). Finally, we discuss prior work on SAT optimization in particular (§2.3). 2.1

Optimizing to a reference

Traditional synthesis techniques. Program synthesis writ large (see [40]) has a multi-decade research lineage. Our work is closest in spirit to inductive techniques, which synthesize programs against partial specifications like inputoutput pairs [39, 89]. Oracle-guided variants query an oracle that returns counterexamples to constrain the next candidate [3, 54, 55, 98, 106]. Our setting shares the iterative structure: an agent proposes a candidate, observes how it performs, and revises. More closely related is superoptimization, which searches for a fast program equivalent to a given input program, ranking by a cost function. Equivalence is established by proof or by testing depending on the technique. Underlying search techniques include, e.g., enumeration [75], deductive proof search over an axiomatized machine [56], or candidate generation checked against an oracle [9]. STOKE [95] takes the last approach and comes closest to ours in spirit, sampling whole candidate programs. It treats correctness as a term in a cost function, with a symbolic validator checking the candidates that survive. The reference need not be an existing program; generative libraries search over equivalent decompositions [90] or schedules [19, 51, 91] against a fixed mathematical specification. Workload specialization. Specializing programs to workloads is longstanding practice across computer science. Architects design domain-specific processors [25, 42, 44, 57, 68], while database researchers argue for engines specialized by workload [100, 101]. Automating specialist construction has historically meant autotuning, searching a space of implementations laid out in advance (e.g., for linear algebra kernels [113], or signal transforms [90]). FFTW [33] defers the search to runtime. Across these systems, the search or design space is constructed to be semantics-preserving, whether via program equivalence, specification refinement, or parameterization of an existing implementation. The tradeoff is that a general correctness requirement constrains the search: requiring generated SAT solvers to operate correctly on all formulas rules out, for example, solvers that can only solve rook placement problems (fast!). 2.2

Optimizing and synthesizing with LLMs

Recent LLM-based optimization approaches relax correctness requirements to search a much larger space of alternatives. LLMs have been used to generate whole-system database implementations customized to a workload. They have synthesized query engines for a workload contract [109] or

execution code per SQL template [66], handling off-nominal inputs by falling back to a general-purpose engine. Most LLM-driven optimization work is considerably more constrained (see [70]). One line of work generates an improved component of a human-written outer algorithm, such as a scoring function, priority rule, or penalty term [69, 81, 94, 116]; another gives the model an existing implementation and asks for a faster equivalent [86]. SuperCoder [110] uses an LLM as a superoptimizer, asking for assembly that runs faster than a compiled reference while preserving behavior. These searches check success against a provided test suite. Tests are only partial correctness specifications, so an agent can satisfy the tests without actually solving the problem. Agents exploit evaluation harnesses [65], fail the tests they were rewarded for passing [110], and replace complete search with approximations that return schema-valid wrong answers [108]. This prior work establishes that LLMs can effectively search a large space of implementations heuristically. We also use LLMs to search for programs, but limit our search to applications where it’s easy to check programs’ results for correctness.

2.3

Optimization for SAT solvers

While SAT is an NP-complete problem—and random SAT instances are very hard to solve [20]—formulas derived from real-world problems tend to have much faster than averagecase runtime [5]. Modern SAT solvers can solve instances with millions of clauses and variables, and the SAT Competition shows huge performance improvements over the years. The first solver optimizations were two foundational algorithms, Davis-Putnam-Logemann-Loveland (DPLL) [26] and conflict-driven clause learning (CDCL) [74, 78]. Improvements since have broadly focused on engineering [13] or solver heuristics [67]. LLMs are a recent addition to the solver optimization literature: both SATLUTION [117] and AE-Kissat-MAB [28], which won the 2025 SAT Competition [23], use LLMs to optimize existing solver codebases. No prior work builds a fresh solver from scratch for every workload, but some specialize in a limited way, or manually. SATenstein [61], which only solves SAT instances, synthesizes new solvers for a given workload by combining pieces of existing solvers. AutoSAT [103] and AutoModSAT [102] use LLMs to optimize for several specific families. They start from a simplified solver codebase, and only allow the LLM to modify a maximum of nine heuristics within the CDCL implementation. Prior work also manually specializes solvers to a specific workload (e.g., cryptanalysis [79, 99]), or automatically specializes by tuning parameters [47, 48, 72] or choosing the anticipated best from a portfolio of solvers for a given instance [58, 59, 114, 115].

Harrison Green, Claire Le Goues, and Fraser Brown

3

Automated hyperspecialization

The software status quo is general-purpose: a single implementation processes many different workloads. But these workloads may favor different algorithms, representations, and heuristics, creating opportunities to take advantage of a workload’s particular structure. The goal of automated hyperspecialization is to realize these opportunities by automatically synthesizing specialist implementations tailored to each target workload. In contrast to optimization, the defining feature of hyperspecialization is that it considers workload-specific implementations, not general-purpose ones. It synthesizes implementations designed to handle, for example, only SAT problems generated by the Kani [27] verifier, or only encodings of register allocation problems. A particular implementation need only perform well—or perform at all!—on its target workload. Our intuition is that restricting workloads allows implementations to exploit assumptions about the types of inputs they will encounter. For example, a specialist for register allocation instances can assume that the problem is graph coloring, and quickly refute with a counting argument. The next section describes the qualities that make a software task amenable to hyperspecialization (§3.1): if a task has heterogeneous workloads, measurable objectives, and a checkable result, it can be iteratively and safely hyperspecialized by agents. This paper focuses on hyperspecialization for SAT since SAT solvers are critical pieces of software across a huge range of domains, from mathematics to software verification to circuit design. We describe why SAT is hyperspecializable (§3.2), including details about workloads, the specific measurable objective, and how to check the correctness of SAT solver results. Finally, we discuss our realization of hyperspecialization as a whole-codebase synthesis task, where agents iteratively produce and measure a SAT solver implementation tailored to a given workload (§3.3).

3.1

What kind of software can be hyperspecialized?

There are three qualities that make a task amenable to automated hyperspecialization. The task must have: 1. Heterogeneous workloads: the task should operate on more than one type of workload so that there is plausibly a benefit from building workload-tailored specialists. 2. Measurable objective(s): it should be possible to compare one implementation to another implementation for a given measurable property (e.g., speed, memory use, etc.). 3. Efficiently checkable results: to safely automate software construction—and, as a result, deploy untrusted code—implementations’ results must be checkable.

3.2

SAT can be hyperspecialized

The task of solving SAT formulas satisfies all of our desiderata: the workload is heterogeneous, because solvers need to solve formulas encoding very diverse problems; faster runtime is the obvious objective; and we can efficiently check a SAT solver’s results. We discuss SAT and SAT solvers next, then expand on each criterion in more detail. 3.2.1 SAT. Boolean Satisfiability (SAT) asks whether a formula over Boolean variables and operators ∧, ∨, and ¬ is true under some assignment to the variables. Given the space of inputs X (formulas) and outputs Y (assignments or proofs), a SAT solver (𝜎 : X ⇀ Y ∪ unknown) is a potentially nondeterministic procedure which, upon receiving input 𝑥 ∈ X either produces an output 𝑦 (i.e., a satisfying assignment or UNSAT proof), runs indefinitely, or returns unknown. R ⊆ X × Y is the set of correct outputs. 3.2.2 Heterogeneous workloads in SAT. We use I to refer to a workload, a finite set of input instances. Our implementation of hyperspecialization specializes by family—a set of SAT formulas drawn from the same task or generator. Family 𝑘 has representative workload I𝑘 . During hyperspecialization, we split this workload into separate training and validation splits, denoted I𝑘train and I𝑘val . 3.2.3 Measuring solver performance. We care about SAT solvers being fast. We can measure a particular SAT solver’s performance by running it on an input in a fixed environment with a timeout. Running 𝜎 on input 𝑥 with timeout 𝑇 yields an output 𝑦 and runtime measurement 𝑡 (𝑦, 𝑡) ∼ Run𝑇 (𝜎, 𝑥) where 𝑦 ∈ Y ∪unknown is the observed output and 𝑡 ∈ [0,𝑇 ] is the (right-censored) observed runtime. Crashes and outof-memory (OOM) errors are treated as (unknown,𝑇 ). PAR-2 scores. Following the SAT Competition’s example [34], we convert runtime measurements into a PAR-2 score in order to encourage both (1) solving instances faster and (2) solving more instances correctly within the timeout: ( 𝑡 if (𝑥, 𝑦) ∈ R, PAR-2(𝑥, 𝑦, 𝑡) := 2𝑇 otherwise. A solver’s score on a given input 𝑥 for which it produces 𝑦 in time 𝑡 is 𝑡 if the solver returned a valid solution within the time budget, and twice the timeout otherwise. To compute performance of a solver 𝜎 on a workload I, denoted 𝑃 (𝜎, I), we run the solver once on each formula: (𝑦𝑥 , 𝑡𝑥 ) ∼ Run𝑇 (𝜎, 𝑥)

for each 𝑥 ∈ I

and compute the average PAR-2 score: 1 ∑︁ 𝑃 (𝜎, I) := PAR-2(𝑥, 𝑦𝑥 , 𝑡𝑥 ) |I| 𝑥∈I

The Case for Automated Hyperspecialization: Evidence from SAT

3.2.4 Checking solver results. Satisfying assignments are conceptually easy to check by plugging in the values and evaluating the formula. We use the formally verified gratchk checker in SAT mode [64]. Checking UNSAT answers, on the other hand, requires augmenting the solver to emit a proof [38, 118]. Our synthesized solvers emit proofs in either DRAT [112], DPR [45], or VeriPB [37] format. We validate these proofs using the GRAT [64], DPR [104], and VeriPB [15] proof-checking toolchains (as described for the SAT Competition 2025), respectively. Each of these toolchains consists of an elaborator which fills in details omitted in the original proof and emits an elaborated certificate which is checked by a formally verified checker. The choice of proof format and toolchain affects both the reasoning that can be expressed directly and the cost of generating and verifying certificates. For both SAT and UNSAT results, checking takes polynomial time in the (combined) size of the formula and the solver’s output. 3.3

Agent-based hyperspecialization for SAT

We treat the problem of SAT solver hyperspecialization as a whole-codebase synthesis task. Given a target workload I𝑘 , we task a coding agent (GPT-5.6 Sol in Codex) with producing a specialized solver 𝜎𝑘 that minimizes 𝑃 (𝜎𝑘 , I𝑘 ) (smaller PAR-2 is better). To ensure that agents do not simply memorize solutions, we provide the agent 80% of the instances as training data I𝑘train , and evaluate performance of submitted solvers on the withheld 20% validation data I𝑘val . Agents may submit candidate solvers for evaluation. An external evaluation server runs the candidate solver on both the training and validation splits, but only returns training results to the agent as feedback. We capture the sequence of submissions and their respective cumulative costs as a trace: D E𝑛 T := (𝜎𝑘(𝑖 ) , 𝐶 (𝑖 ) ) 𝑖=1

Where 𝜎𝑘(𝑖 ) is the 𝑖-th candidate solver and 𝐶 (𝑖 ) is the cumulative cost (in time and money) to produce and evaluate 𝜎𝑘(𝑖 ) . SAT solver development naturally yields a bimodal workload: the agent spends some time writing, editing, and debugging code (low CPU usage, inference-bound, largely sequential) and some time running SAT solver evaluations (high CPU usage, compute-bound, parallelizable). To use compute efficiently, we design our hyperspecialization framework around this division (Figure 2). We run agents on dedicated machines with bounded compute (i.e., a single core and limited memory), and instruct them to submit candidate solvers to an external evaluation server that runs parallel evaluations in the cloud. For all experiments, we use GPT-5.6 Sol [83] in Codex with extra-high reasoning (see §G for details on this choice, and §A for the prompt). The agent runs in an isolated Docker

Edit σ¹

Edit σ²

Edit σ³

Agent paused

1 core

paused

paused

Evaluate σ²

Evaluate σ³

Cloud 10 CPUs Evaluate σ¹ Agent active

Training (80%)

Validation (20%)

Figure 2. Our bimodal optimization framework alternates between (a) using a coding agent to construct and optimize candidate solvers and (b) evaluating these solvers on parallelized cloud compute (shown here with 10 CPUs). container in “bypass approval” mode using the default Codex harness. We run all agents with broad internet access disabled. While web search could be beneficial in some cases, we found it too difficult to enable web search without potentially leaking validation formula data (see §I for more). Each agent run is allocated two hours of active time—time spent synthesizing and testing code, not external evaluation time—and a maximum of ten external evaluations. 3.4

When does hyperspecialization pay off?

Hyperspecialization makes sense when the benefits outweigh the costs. In our agent-based framework (§3.3), hyperspecialization incurs two costs: money (i.e., LLM token spend) and time (i.e., waiting for the agent to run). The potential benefit is that a faster solver will save time in the future. If we synthesize a specialist that is faster than the generalpurpose baseline, and run that specialist on sufficiently many instances, we will eventually gain back our upfront specialization time. Running the cost-benefit analysis for monetary cost is less straightforward, since it requires assigning monetary value to a speedup. How much would you pay for a 2× faster solver for your workload? For some workloads and companies, the answer is perhaps “a lot!” Our evaluation (§4) analyzes this cost-benefit tradeoff in two regimes. In the time-insensitive regime, specialists are built once and run over an infinite future workload. The specialization time is fully amortized, so we only care about speedups compared to monetary cost. In the time-sensitive regime, specialists will run over a finite future workload, and our goal is to minimize the total runtime over that workload: is it faster to hyperspecialize or run an off-the-shelf solver? 3.4.1 Cost in the time-sensitive regime. To reason about tradeoffs in the time-sensitive regime, consider a machine whose goal is to solve as many instances as possible, and which has compute to devote to either hyperspecialization or solving (with a generated specialist or an off-the-shelf general-purpose solver). This machine has 𝐶 identical cores and an ordered workload W := ⟨𝑥 1, 𝑥 2, . . . , 𝑥 𝑁 ⟩ where all 𝑁 instances are available at time zero. A deployment policy 𝜋

Harrison Green, Claire Le Goues, and Fraser Brown

specifies which computations to run, when to run them, and how to schedule them on the available cores. For example, a fixed policy runs the same solver on the next queued instance whenever a core becomes available. Other policies may race multiple solvers or allocate cores to domain-specific hyperspecialization. For each policy, we measure its makespan 𝑀𝜋 (W, 𝐶), the elapsed time until all scheduled tasks have terminated, and its verified solve count 𝑄 𝜋 (W, 𝐶), the total number of correct solutions. Choosing between several policies is inherently a multi-objective problem: we want to both minimize the makespan and maximize the number of solved formulas. The space of policies, therefore, defines a Pareto frontier, where two policies on this frontier are not directly comparable: one is better in one metric and worse in the other. For our evaluation of the time-sensitive regime (§4.3), we make the simplifying observation that the shape of this problem is analogous to that of comparing individual solver performance. The best solver should run faster and solve more formulas within the timeout. Therefore, we lift PAR-2 score to the policy level: 𝑃𝜋 (W, 𝐶) = 𝑀𝜋 (W, 𝐶) + 2𝑇 (𝑁 − 𝑄 𝜋 (W, 𝐶)) A policy’s score is the time it took to run (including hyperspecialization time, when applicable) plus a timeout penalty for every instance it failed to solve. We discuss a Pareto frontier-based treatment of optimality in §C.

4

Evaluation

We address the following broad research questions: RQ1: How much does hyperspecialization improve performance compared to a general-purpose baseline? RQ2: When is hyperspecialization economically viable? We deploy domain-specific hyperspecialization at scale: we create solvers for the 189 families in the GBD across three different proof formats. We evaluate the solvers’ cost and performance relative to a portfolio of four state-of-the-art general-purpose solvers. Cumulatively, our evaluation consumed approximately $25k worth of LLM tokens (at API pricing) and over 5 CPU-years of solver evaluation. We synthesized over 5,700 unique solvers and ran over 1.8 million solver invocations. Next, we describe our evaluation setup (§4.1), and the answers to RQ1 (§4.2) and RQ2 (§4.3); §5 answers more informal questions about hyperspecialization. 4.1

Evaluation setup

We describe our dataset (§4.1.1), baseline solvers (§4.1.2), the system environment (§4.1.3), and how we ensured task alignment (§4.1.4); §3.2.4 already described the three proof formats we use, with one independent agent per proof format.

Table 1. Overall winner from the previous four years of the SAT Competition, and percentage of time it appears in the VBS. Year

Solver

Baseline wins

2023 2024 2025 2026

SBVA-CaDiCaL [41] Kissat-sc2024 [14] AE-Kissat-MAB [28] satsuma-iter+kissat3

16.9% 32.0% 30.5% 20.6%

4.1.1 Dataset. We draw formula families from the GBD [52], which consists of2 31,809 CNF formulas split across 189 different families. These formulas have accumulated over 20 years of SAT competitions, and vary in size from 2 to 9,842 instances. 4.1.2 Baseline solvers. We compare against the overall winner of each of the previous four SAT Competitions (Table 1), which represents each year’s strongest sequential solver across a range of formulas. To make the comparison stronger, we also report a per-family virtual best solver. For each family 𝑘, we pick the baseline with the lowest training PAR-2 and evaluate that choice on the validation set. Given baselines Σ = {𝜎 (1) , . . . , 𝜎 (𝑛) }, 𝑃 (𝜎 VBS, I𝑘val ) := 𝑃 (argmin 𝑃 (𝜎, I𝑘train ), I𝑘val ) 𝜎 ∈Σ

Across the 189 families in our dataset, each of the four baseline solvers wins (and is thus selected in the VBS) between 16.9% and 32.0% of the time (Table 1). §B extends our comparison to recent LLM-based SAT solver improvement frameworks: AutoSAT [103], AutoModSAT [102], and SolSearch [96]. We do not compare to closed-source SATLUTION [117]. 4.1.3 Environment. We keep the environment consistent across all experiments. Solver evaluations use a timeout of 600 seconds, a memory budget of 8 GB, and a verifier timeout of 6,000 seconds. They run in Google Cloud Batch using preemptable spot instances (n2-highmem-2 with Intel Ice Lake, to match the SAT Competition 2026 evaluation environment). Each task is allocated 1 vCPU and 8 GB RAM, with two tasks per instance; this is uniform across all solvers. Coding agents run on an AMD EPYC 9454P processor with 256 GB of RAM. 4.1.4 Task alignment. Agents deployed at scale can score well without faithfully solving the task. We refined the task description and environment beforehand using Docent [76] to surface avenues for cheating in pilot runs (§I.2). Afterward, we audited every agent transcript and generated solver for compliance (§I.3), found seven invalid runs, and reran them; 2 At evaluation time. 3 Proceedings of SAT Competition 2026 not yet available.

The Case for Automated Hyperspecialization: Evidence from SAT

a second audit of those seven found no further violations. All results in this paper refer to valid runs. 4.2

RQ1: How much does hyperspecialization improve performance?

For every family in the dataset, we select the best specialized solver (across all three proof systems) by evaluating performance on the training set. We compute the baselineto-specialist PAR-2 ratio, defined below, on the validation set, comparing each specialist with the per-family virtual-best of the four baseline solvers. Figure 3 shows this ratio both excluding and including verification time. PAR-2 ratio. We compare the performance of two solvers 𝜎𝑎 and 𝜎𝑏 using the log PAR-2 ratio: Δ𝑃 (𝜎𝑎 , 𝜎𝑏 , I) := log10

𝑃 (𝜎𝑏 , I) 𝑃 (𝜎𝑎 , I)

We use a ratio because we care about improvements relative to 𝜎𝑏 ’s score, rather than absolute improvement in PAR-2. The logarithmic scale makes improvements and regressions symmetric around zero: a positive value Δ𝑃 > 0 indicates that 𝜎𝑎 has a better (smaller) PAR-2 score, and a negative value Δ𝑃 < 0 indicates that 𝜎𝑏 has a better (smaller) PAR-2 score. Speedups. Agents constructed hyperspecialized solvers that yielded significant speedups over the baseline solvers. Excluding verification time, specialization produced a faster solver in 121/189 families (64.0%), and produced a solver that had more than a 1,000× reduction in PAR-2 score for 32/189 families (16.9%). The geometric mean ratio was a 12.0× reduction in PAR-2 score. Including verification time in both baseline and specialist scores reduces the geometric mean ratio to 5.03×. The specialist still outperforms the baseline on 121/189 families: nine families gain an advantage and nine lose it. The effect varies substantially by family. For binary-tree-parity, for example, the ratio remains approximately 323,000×. For purdominstances, it falls from 917× to 0.278×. In this case, the specialist found a cheap way to generate an expensive-to-check proof. Failure cases. When agents were not able to discover optimized specialists, their performance still almost always approached baseline performance. Specialists had worse than a 10× increase in PAR-2 score on only two families. The worst such family is coloring-clique, where the specialist is 104 × worse. The GBD includes one shuffled formula—i.e., an instance with renamed variables and reordered clauses— in this family, alongside ten structurally distinct coloringclique instances. The shuffled formula was randomly allocated to the validation set with no counterpart in training; the agent never observed it, and thus could not build a solver that accounted for shuffling. Shuffling itself is not the issue:

many other families on which specialists perform well contain shuffled variants. When the shuffled instance is removed from the coloring-clique validation set, the specialist outperforms the baseline (14×, 2× accounting for verification time). Evaluating performance transfer. Specialist performance on a family should still transfer to unseen instances drawn from the same family. To measure this, we plot the correlation between solvers’ PAR-2 scores on the training set and validation set. Highly correlated scores indicate that specialist performance transfers, while higher PAR-2 scores (i.e., worse performance) on the validation set (vs. the training set) imply overfitting to the training set. Figure 4 shows baseline solver training/validation correlation in the left panel (A) as a reference; note that the baseline solvers, unlike the specialists, were not actually trained. We plot a point at (𝑃 (𝜎, I𝑘train ), 𝑃 (𝜎, I𝑘val )) for every baseline solver 𝜎 ∈ Σ and family 𝑘. In the right panel (B), we plot the performance of every specialized solver synthesized throughout all of the agent runs. We plot (𝑃 (𝜎𝑘(𝑖 ) , I𝑘train ), 𝑃 (𝜎𝑘(𝑖 ) , I𝑘val )) for every family 𝑘 and every candidate solver 𝜎𝑘(𝑖 ) produced during domain-specific hyperspecialization for family 𝑘. Training and validation PAR-2 were highly rank-correlated both for specialists (Spearman 𝜌 = 0.83) and baselines (𝜌 = 0.88). While specialists were more varied in performance than baselines (median absolute deviation of 1.27× vs. 1.16×), they were only slightly more biased towards training performance. The median ratio of validation PAR-2 to training PAR-2 was 1.01 for specialists (compared to 1.00 for baselines). 6.0% of submissions had a validation/training ratio above 10, compared to 4.6% of submissions with validation/training ratio below 0.1. The near-symmetry of these measures suggests variance within a family as opposed to overfitting. 4.3

RQ2: When is hyperspecialization economically viable?

We evaluate specialization in both the time-insensitive regime (§4.3.1) and the time-sensitive regime (§4.3.2). 4.3.1 Time-insensitive regime. We consider the case where a specialized solver will be reused many times—or infinitely!—making money the driving cost. In Figure 5, we compare how the improved performance of a candidate specialized solver relates to its cumulative construction cost. For every submission index 𝑖, we plot the average PAR-2 validation ratio over the baselines (same as RQ1) and the cumulative cost (LLM usage + cloud compute) of the candidate solver 𝜎𝑘(𝑖 ) . Hyperspecialization potential was predominantly influenced by family, not cumulative specialization effort. On

Harrison Green, Claire Le Goues, and Fraser Brown

Specialists achieve large speedups in a majority of families

Excl. verification

Incl. verification

DPR

VeriPB

GRAT

binary-tree-parity 742k× → 323k×

106×

PAR-2 ratio

104× software-verification

102×

0.95× → 1.71×

1× rooks purdom-instances

10−2×

11.0k× → 7.43×

917× → 0.28×

coloring-clique 7,398× → 1,268× worse

10−4× 1

25

50

75

100

125

150

175

189

Families (ranked)

Figure 3. Held-out validation score ratios (baseline / specialist) against the per-family virtual-best of four baseline solvers. Gray stems and open circles exclude verification; colored markers add verification time to both solvers while retaining PAR-2 penalties. Both solvers are selected using training PAR-2. Families are ranked by the ratio excluding verification; values above 1× favor specialization. Color and marker shape identify the specialist’s proof system. Callouts show speedup excluding → including verification time. Baselines and specialists generalize to held-out instances A Baselines

B Specialists

Spearman ρ = 0.88 · n = 748

Spearman ρ = 0.83 · n = 4,725

PAR-2 ratio on validation set

Proof system

DPR

VeriPB

GRAT

Checkpoint means (±1 SD)

10³× better

1k Validation PAR-2 (s)

Specialist performance increases with cumulative cost

10²× better 10 10¹× better 0.1

Equal 10¹× worse

0.001 0.001

0.1

10

1k

0.001

0.1

10

1k

Training PAR-2 (s)

Figure 4. Training and validation PAR-2 for baseline solvers (A) and all generated specialists (B), with pooled Spearman correlations. average, even the agents’ first solver submissions (𝜎𝑘(1) ) were faster than baselines. The average time to the first submission was just 28.1 minutes at a total cost of $10.63, yielding a 1.71× improvement over baseline. Full runs (limited to 2 hours of active time and 10 evaluations) took on average 1.5 hours at a cost of $37, yielding a 4.40× improvement.4 The agent accounted for a majority of the total cost (91.4%), yet a minority of the overall runtime (26.9%). Additionally, we observed that among the three proof formats, agents had the most success with VeriPB, on average yielding slightly more 4 This includes submissions from all three proof formats (not just the best

one) thus is lower than the average speedup in RQ1.

$0

$20

$40

$60

$80

$100

Cumulative cost (USD)

Figure 5. Validation speedup of specialized solvers relative to cumulative construction cost. Each point shows the average performance of the 𝑖th submitted solver across all agent runs (separately for each evaluated proof format).

performant solver submissions than with GRAT at a slightly cheaper average cost. On the other hand, DPR runs were on average more expensive than both GRAT and VeriPB runs for unknown reasons. 4.3.2 Time-sensitive regime. When there does not yet exist a specialized solver for a particular workload, we must decide whether to attempt hyperspecialization or just run an existing general-purpose solver. Here, there is an explicit tradeoff between the amount of compute and time allocated for specialization vs. solving. We evaluate eleven different scenarios.

The Case for Automated Hyperspecialization: Evidence from SAT

Workload size N (Instances)

Hyperspecialization becomes optimal for most families as workloads grow 1 2 4 8 16 32 64 128 256 512 1k 2k 5k 10k 20k 40k 80k … 1M 0

C=1

C=2

C=4

C=8

C=16

C=32

C=64

···

···

···

···

···

···

···

50

100 0 50 100 0 Share of families won (%)

50

50

100 0

SBVA '23

Kissat '24

50

100 0

Fixed baseline AE-Kissat '25

100 0

Race Satsuma '26

50

100 0

50

100

Specialization

Four-solver race

Wait · DPR

Wait · VeriPB

Wait · GRAT

Hybrid · DPR

Hybrid · VeriPB

Hybrid · GRAT

Figure 6. Share of optimal on-demand policies across simulated workloads. Each horizontal bar specifies a simulated machine with 𝐶 cores and a workload size of 𝑁 . The bar is colored to represent the share of families for which a given on-demand ∗ policy 𝜋 was the optimal policy 𝜋𝐶,𝑁 . ,𝑘 Fixed baseline policy (4 instances). A fixed baseline policy runs the same solver on the next workload instance whenever a core is available. We evaluate a fixed policy separately for each of the four baseline solvers. Four-solver race policy (1 instance). We evaluate a foursolver race where each of the fixed baselines is deployed in parallel for a single workload instance whenever four cores are free (which requires 𝐶 ≥ 4). All instances terminate whenever the first solver returns a verified solution. Specialization policy (6 instances). We evaluate two hyperspecialization deployment policies for each of the three proof formats. In wait, all cores are used for hyperspecialization. We replay the agent and evaluation compute workload based on the available parallelism. Once hyperspecialization completes, the best checkpoint is served on all cores. In hybrid, hyperspecialization runs until the first solver is submitted. Then, half the cores immediately serve that solver, while the remaining cores continue hyperspecialization. When a new, better solver is produced, the serving cores switch to it. Once specialization completes, all cores serve the best solver. We performed a discrete-event simulation of each of these runtime policies. For every core count 𝐶 ∈ {1, 2, 4, 8, 16, 32, 64}, workload size 𝑁 ∈ {1, 2, 4, . . . , 80 𝑘, 1 𝑀 }, and family 𝑘, we sample 𝑟 random workloads of size 𝑁 (with replacement) from the validation set I𝑘val , producing W𝑘(1) , . . . , W𝑘(𝑟 ) . We

simulate each policy 𝜋 on each of these representative workloads and compute the optimal policy for family 𝑘 in configuration (𝐶, 𝑁 ) as the policy with the lowest total score: 𝑟 ∑︁ ∗ 𝜋𝐶,𝑁 𝑃𝜋 (W𝑘(𝑖 ) , 𝐶) ,𝑘 := argmin 𝜋 ∈Π

𝑖=1

We visualize results in Figure 6. Each horizontal bar denotes a particular number of cores 𝐶 (x-axis) and a particular workload size 𝑁 (y-axis). The bar is colored based on what share of families a given policy 𝜋 is the optimal policy for the configuration (𝐶, 𝑁 ). This visualization allows us to see how the distribution of optimal strategies changes as we change the machine size (𝐶) and the workload size (𝑁 ). Results. For small workloads, running fixed baseline solvers directly is most frequently the optimal strategy as these solvers incur no upfront cost. In regimes where the workload undersaturates the available compute (i.e. 𝑁 < 𝐶), the four-solver race is frequently dominant, as it provides a mechanism to use this “free” compute. Across all machine sizes, as the workload size increases, hyperspecialization increases in viability. On average, GRAT specialists fare better than both VeriPB and DPR. Within each format, the hybrid form of the strategy (where available specialists are immediately deployed) is optimal on more families than the wait form. Once the workload size exceeds 𝑁 = 2, 000, some form of hyperspecialization is the optimal strategy on a majority of families across all machine sizes.

5

Case studies

This section describes four case studies: classifying the strategies used by generated solvers, deploying specialists in real

Harrison Green, Claire Le Goues, and Fraser Brown

systems that use solvers, measuring how specialists for one family perform on another, and exploring how sensitive specialists are to cardinality encodings (described later).

Case-study performance Train / Val.: isolated formulas

Results. Specialists did not compete with generalists by developing new algorithms; rather, they sampled from the vast SAT and algorithms literature to solve the problem at hand, and implemented their solutions with systems-level optimizations. Generated solvers were remarkably varied and familyspecific. Our audit found at least forty unique solver strategies: CDCL, the backbone of almost all modern solvers, was the second most common strategy (57.5%), after constructing witnesses or deriving proofs directly from problem structure (60.7%). Only 16.2% of solvers deployed CDCL on its own, and 81.1% of solvers used more than one strategy (Figure 13). Specialists frequently recovered high-level structure from input formulas (Figure 14)—for example, graphs (41.3%), Boolean circuits (38.6%), and cardinality constraints (27.9%)—and 97.2% of specialists included some check excluding formulas that did not fit the expected structure (Figure 16). Recovered structure also informed orchestration for multi-strategy solvers: 93.3% of solvers selected solving procedures based on structure, and 60.1% ran distinct methods sequentially (Figure 18). Specialists used a wide range of known heuristics and techniques (e.g., XOR reasoning). They also optimized beyond just algorithmic and search speedups, using an array of low-level systems techniques like memory preallocation (96.5%), specialized parsing (66.0%), buffered proof logging (45.1%), going so far as to add branch-prediction hints (1.1%) and SIMD-accelerated code (1.8%) (Figure 19). 5.2

Original New Train Val.

Solver taxonomy

We surveyed the entire set of 567 training-selected specialists across the 189 families, using LLM-based techniques to analyze and classify the nearly 500,000 lines of generated solver code. This section gives broad insights into specialists’ strategies, while §K describes our method and its limitations; in §J, we provide concise, LLM-generated descriptions of every family and solver.

End-to-end deployment

We ask whether applying hyperspecialization to real-world systems results in solvers competitive with native performance. We isolated SAT formulas from five real systems, synthesized a specialist, and then evaluated the specialist in a simulated deployment on new targets. Figure 7 presents the results compared to native baselines for circuit equivalence checking (OpenTitan [85] and ORFS [105]), verification (CBMC [22] and Kani [27]), and package dependency resolution (Conda [93]); see §E for details. Results. For two of our five systems (OpenTitan and ORFS), specialists outperformed native solutions on novel targets

With verification

Train Val.

OpenTitan

5.1

Without verification

Original / New: native workflows

ORFS

Original New Train Val.

CBMC

Original New Train Val.

Kani

Original New Train Val.

Conda

(pycosat)

Original New

(libmamba)

Original New

0.1× 1× 10× Baseline / specialist ratio (>1 favors specialist)

100×

Figure 7. Specialists’ redeployment performance across five real-world systems. Formulas were extracted from a target and used to develop a specialist; Train/Val. show specialist speedup over baselines in isolation on these representative formulas. We then redeploy the specialist in the real-world system and measure end-to-end performance compared to the native tool on the original target for which formulas were extracted (original) and a new target (new). For Conda, we evaluate two different native modes: pycosat and libmamba.

even when accounting for proof verification time—something native solutions did not even attempt. Deploying specialists was challenging for several nonobvious reasons, though: (1) real-world systems often implicitly trust solvers and run them without proof-logging, while our specialists always spend time producing proofs; (2) some systems (e.g., CBMC and Kani) optimize by running native solvers in incremental mode, a mode our specialists do not yet (but could!) support; (3) some systems (e.g., Conda) use custom, higher-level (and thus faster) solvers as alternatives to SAT. One potential direction for future work is synthesizing solvers for a given family and a given application (e.g., a verification-specific, incremental solver for Kani). 5.3

Cross-family sensitivity and performance

§4.2 measures solvers only on their target validation set. Here, we run the top per-family selected specialist on the validation set of every other family (Figure 20). Results. Most specialists were very selective, refusing to run (i.e. quickly returning UNKNOWN) on families outside of the target domain: only 27 solvers attempted to run on more than 50% of other families. Targeted specialists were also

The Case for Automated Hyperspecialization: Evidence from SAT

5.4

Cardinality constraint encodings

Many high-level problems define cardinality constraints, or bounds related to counting (e.g., “exactly two of these variables must be true”). These constraints can be encoded into Boolean logic in a variety of ways. For general-purpose solvers, the specific choice of encoding is an active area of research [6–8, 10, 32, 77, 82, 92, 97], because it plays a significant role in solve time. Fewer clauses and/or variables, for example, often correlate with faster solve times. Here, we test if the same is true under hyperspecialization. We evaluated baseline and specialist performance across three crafted benchmarks—chosen because they used many cardinality constraints and covered both SAT and UNSAT—and up to seven cardinality encoding methods; see §H for more. Results. Figure 8 shows the performance of baselines and specialists.6 Unlike general-purpose solvers,7 specialists do not benefit from more compact encodings; in fact, for one of the benchmarks, the naive (pairwise) cardinality encoding yielded the fastest specialist (yet the worst baseline performance). We hypothesize that it is more important for a formula to be encoded in a conceptually clear way for the hyperspecialist, as opposed to a compact—but perhaps convoluted— way. Compactness helps general-purpose solvers because it plays nicely with those solvers’ fixed internal heuristics. In contrast, the specialist chooses its own internal heuristics—in many cases extracting and reasoning about cardinality constraints at a high level. An optimized encoding can stymie the specialist who’s trying to reconstruct a higher-level problem.

6

Discussion

This section considers the ramifications of hyperspecialization for SAT and beyond. 5 The elaborator runs before the formally verified part of the checker. 6 All using GRAT, so specialists must construct resolution-based proofs and

cannot simply map the cardinality constraints into algebraic proofs. 7 The general-purpose encoding benefit is more visible in §H.

How cardinality encoding affects performance Kissat '24

AE-Kissat '25

Satsuma '26

Specialist

LDPC decoding

Vertex cover

Se q. Ca rd. To t. MTo t KM -To t

SBVA '23

GPHP

Se q. So rt. Ca rd. To t. MTo KM t -To t

Validation PAR-2 (s)

1.2k 1/20

10

17/20

3/20

0.1 0.001

Na ive Se q. So rt Ca . rd. To t M- . T KM ot -To t

often the best solver for their family: in 134 of the 189 cases (70.9%), the targeted specialist won its own family. Specialists were also robust. Only seven source/target combinations resulted in solutions rejected by the verifier. Six rejections (all on multiplier-verification) were false positives, caused by a bug in the DPR elaborator dpr-trim5 that mishandled clauses with repeated literals. We have reported this bug with a patch. Only one rejection was legitimate. The fpga-routing solver read clause lines into a fixed 64 KiB buffer with fgets. It correctly preserved unfinished clauses across multiple reads, but not unfinished literals: it dropped minus signs that fell at a buffer boundary, converting negative literals into positive ones. Triggering this bug required a clause line exceeding 64 KiB with a negative literal precisely at the boundary.

Figure 8. Average validation PAR-2 of baselines and specialists for various cardinality encodings. Hyperspecialization and SAT. Hyperspecialization (we hope!) makes it easier for SAT research to have practical impact. In the past, a technique benefiting a small set of formulas was difficult to justify maintaining in a mainline solver. Now, if the technique is useful for a specific workload, an agent can discover and implement it in a specialist. Hyperspecialization may also open up new research avenues. For example, a CDCL solver’s reasoning maps cleanly to DRAT proofs; it doesn’t need the richness of formats like VeriPB. Specialist solvers can take advantage of more expressive reasoning, though—so the bottleneck becomes proof verification time. Can we design formats that speed up verification with specialist-authored hints? What does a proof format designed for hyperspecialists look like? Finally, we may consider the SAT solver within the context of the embedding application. §5.2 describes how synthesizing a specialist from workload examples alone may not be enough to beat native speeds. The synthesis step could improve performance by accounting for information about the embedding application. Hyperspecialization beyond SAT. The most obvious hyperspecialization domains beyond SAT are its close relatives that also support verification: SMT [29], MaxSAT [12], pseudo-Boolean optimization [63], finite-domain constraint solving [36], QBF [80], certified model counting [17], and more. Hyperspecialization should work more generally, though: we claim that it applies to any problem with a given shape (§3.1). We now provide three possibilities in rough order of audacity. A lossless compressor sees many types of inputs (e.g., text or audio), and produces a small archive that decompresses back to the original. There already exist specialized compressors that use domain knowledge to achieve better compression ratios than general-purpose compressors [2, 18, 46, 87, 111]. Automated hyperspecialization may allow us to tailor compressors to other workloads—or even individual inputs! The correctness check is simply decompressing the output and comparing the result with the input. Compilers compile many types of programs, and understanding program semantics often improves performance

Harrison Green, Claire Le Goues, and Fraser Brown

of the generated binary. Could we build hyperspecialized compilers for certain classes of programs, using translation validation [71, 88] to check output? What about web browsers? They serve many different web pages, and today, behemoth applications wring new drips of performance out of JavaScript engines and renderers with each new release. What would per-workload browsers look like, and how might we check them? More broadly, in a world where untrusted authors can specialize (or “specialize”) anything, how might we re-imagine checking?

Generative AI disclosure We used generative AI to assist with various aspects of this research. We used coding agents (primarily GPT-5.6 Sol in Codex, but also other models in Cursor and Claude Code) to help with the development of the SAT hyperspecialization framework, evaluation pipelines, and case studies. Our core experiment tests how well coding agents can synthesize SAT solvers and thus naturally uses generative AI. During evaluation, we used LLMs (GPT-6 Astra and GPT-5.6 Luna) to build taxonomies of solver behavior (detailed in §K). In the preparation of this manuscript, we used coding agents (GPT-5.6 Sol / GPT-6 Astra) to help create graphs and figures. Except for three clearly labeled sections in the appendix (§J, §K.3, and §E.1), the words in this paper were written by humans.

References [1] Ignasi Abío, Robert Nieuwenhuis, Albert Oliveras, and Enric Rodríguez-Carbonell. 2013. A Parametric Approach for Smaller and Better Encodings of Cardinality Constraints. In Principles and Practice of Constraint Programming (Lecture Notes in Computer Science, Vol. 8124), Christian Schulte (Ed.). Springer Berlin Heidelberg, Berlin, Heidelberg, 80–96. https://doi.org/10.1007/978-3-642-40627-0_9 [2] Azim Afroozeh, Leonardo X. Kuffo, and Peter Boncz. 2023. ALP: Adaptive Lossless Floating-Point Compression. Proceedings of the ACM on Management of Data 1, 4, Article 230 (2023), 26 pages. https: //doi.org/10.1145/3626717 [3] Rajeev Alur, Rastislav Bodik, Garvit Juniwal, Milo M. K. Martin, Mukund Raghothaman, Sanjit A. Seshia, Rishabh Singh, Armando Solar-Lezama, Emina Torlak, and Abhishek Udupa. 2013. Syntaxguided synthesis. In 2013 Formal Methods in Computer-Aided Design. IEEE, Piscataway, NJ, USA, 1–8. https://doi.org/10.1109/fmcad.2013. 6679385 [4] Carlos Ansótegui, Maria Luisa Bonet, Jesús Giráldez-Cru, and Jordi Levy. 2017. Structure features for SAT instances classification. Journal of Applied Logic 23 (2017), 27–39. https://doi.org/10.1016/j.jal.2016. 11.004 [5] Carlos Ansótegui, Maria Luisa Bonet, Jesús Giráldez-Cru, Jordi Levy, and Laurent Simon. 2019. Community structure in industrial SAT instances. Journal of Artificial Intelligence Research 66 (2019), 443–472. https://doi.org/10.1613/jair.1.11741 [6] Roberto Asín, Robert Nieuwenhuis, Albert Oliveras, and Enric Rodríguez-Carbonell. 2009. Cardinality networks and their applications. In Theory and Applications of Satisfiability Testing - SAT 2009 (Lecture Notes in Computer Science, Vol. 5584). Springer Berlin Heidelberg, Berlin, Heidelberg, 167–180. https://doi.org/10.1007/978-3642-02777-2_18

[7] Roberto Asín, Robert Nieuwenhuis, Albert Oliveras, and Enric Rodríguez-Carbonell. 2011. Cardinality networks: a theoretical and empirical study. Constraints 16, 2 (2011), 195–221. https: //doi.org/10.1007/s10601-010-9105-0 [8] Olivier Bailleux and Yacine Boufkhad. 2003. Efficient CNF Encoding of Boolean Cardinality Constraints. In Principles and Practice of Constraint Programming – CP 2003 (Lecture Notes in Computer Science, Vol. 2833), Francesca Rossi (Ed.). Springer Berlin Heidelberg, Berlin, Heidelberg, 108–122. https://doi.org/10.1007/978-3-540-45193-8_8 [9] Sorav Bansal and Alex Aiken. 2006. Automatic generation of peephole superoptimizers. ACM SIGOPS Operating Systems Review 40, 5 (Oct. 2006), 394–403. https://doi.org/10.1145/1168917.1168906 [10] Kenneth E. Batcher. 1968. Sorting networks and their applications. In Proceedings of the April 30–May 2, 1968, spring joint computer conference. Association for Computing Machinery, New York, New York, USA, 307–314. https://doi.org/10.1145/1468075.1468121 [11] Adam Belay, George Prekas, Ana Klimovic, Samuel Grossman, Christos Kozyrakis, and Edouard Bugnion. 2014. IX: A Protected Dataplane Operating System for High Throughput and Low Latency. In 11th USENIX Symposium on Operating Systems Design and Implementation (OSDI 14). USENIX Association, Broomfield, CO, USA, 49– 65. https://www.usenix.org/conference/osdi14/technical-sessions/ presentation/belay [12] Jeremias Berg, Bart Bogaerts, Jakob Nordström, Andy Oertel, and Dieter Vandesande. 2023. Certified Core-Guided MaxSAT Solving. In Automated Deduction – CADE 29 (Lecture Notes in Computer Science, Vol. 14132). Springer, 1–22. https://doi.org/10.1007/978-3-031-384998_1 [13] Armin Biere, Tobias Faller, Katalin Fazekas, Mathias Fleury, Nils Froleyks, and Florian Pollitt. 2024. CaDiCaL 2.0. In Computer Aided Verification (Lecture Notes in Computer Science, Vol. 14681), Arie Gurfinkel and Vijay Ganesh (Eds.). Springer, Cham, 133–152. https://doi.org/10.1007/978-3-031-65627-9_7 [14] Armin Biere, Tobias Faller, Katalin Fazekas, Mathias Fleury, Nils Froleyks, and Florian Pollitt. 2024. CaDiCaL, Gimsatul, IsaSAT and Kissat Entering the SAT Competition 2024. In Proceedings of SAT Competition 2024: Solver, Benchmark and Proof Checker Descriptions (Department of Computer Science Report Series B, Vol. B-2024-1), Marijn J. H. Heule, Markus Iser, Matti Järvisalo, and Martin Suda (Eds.). Department of Computer Science, University of Helsinki, Helsinki, Finland, 8–10. https://hdl.handle.net/10138/584822 [15] Bart Bogaerts, Ciaran McCreesh, Magnus O. Myreen, Jakob Nordström, Andy Oertel, and Yong Kiam Tan. 2023. VeriPB and CakePB in the SAT Competition 2023. In Proceedings of SAT Competition 2023: Solver, Benchmark and Proof Checker Descriptions (Department of Computer Science Series of Publications B, Vol. B-2023-1). University of Helsinki, Helsinki, Finland, 86–88. https://hdl.handle.net/10138/ 563824 [16] Russell Brandom. 2026. OpenAI’s Jalapeño Chip Is Built for Fast Inference at Scale, Benchmarks Show. TechCrunch. https://techcrunch.com/2026/08/25/openais-jalapeno-chip-is-builtfor-fast-inference-at-scale-benchmarks-show/ [17] Randal E. Bryant, Wojciech Nawrocki, Jeremy Avigad, and Marijn J. H. Heule. 2023. Certified Knowledge Compilation with Application to Verified Model Counting. In 26th International Conference on Theory and Applications of Satisfiability Testing (SAT 2023) (Leibniz International Proceedings in Informatics (LIPIcs), Vol. 271). Schloss Dagstuhl – Leibniz-Zentrum für Informatik, 6:1–6:20. https: //doi.org/10.4230/LIPIcs.SAT.2023.6 [18] Shubham Chandak, Kedar Tatwawadi, Idoia Ochoa, Mikel Hernaez, and Tsachy Weissman. 2019. SPRING: A Next-Generation Compressor for FASTQ Data. Bioinformatics 35, 15 (2019), 2674–2676. https://doi.org/10.1093/bioinformatics/bty1015

The Case for Automated Hyperspecialization: Evidence from SAT [19] Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Meghan Cowan, Haichen Shen, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18) (Carlsbad, CA, USA) (OSDI’18). USENIX Association, Carlsbad, CA, USA, 579–594. https://www.usenix.org/conference/ osdi18/presentation/chen [20] Vašek Chvátal and Endre Szemerédi. 1988. Many hard examples for resolution. J. ACM 35, 4 (1988), 759–768. https://doi.org/10.1145/ 48014.48016 [21] David D. Clark and David L. Tennenhouse. 1990. Architectural Considerations for a New Generation of Protocols. In Proceedings of the ACM Symposium on Communications Architectures and Protocols. Association for Computing Machinery, New York, NY, USA, 200–208. https://doi.org/10.1145/99508.99553 [22] Edmund Clarke, Daniel Kroening, and Flavio Lerda. 2004. A Tool for Checking ANSI-C Programs. In Tools and Algorithms for the Construction and Analysis of Systems (Lecture Notes in Computer Science, Vol. 2988). Springer, Berlin, Heidelberg, 168–176. https: //doi.org/10.1007/978-3-540-24730-2_15 [23] Cayden Codel, Katalin Fazekas, Marijn J. H. Heule, and Markus Iser. 2025. The Results of SAT Competition 2025. SAT 2025 conference presentation. https://satcompetition.github.io/2025/satcomp25slides. pdf [24] Stephen A. Cook. 1971. The complexity of theorem-proving procedures. In Proceedings of the Third Annual ACM Symposium on Theory of Computing (Shaker Heights, Ohio, USA) (STOC ’71). Association for Computing Machinery, New York, NY, USA, 151–158. https://doi.org/10.1145/800157.805047 [25] William J. Dally, Yatish Turakhia, and Song Han. 2020. Domainspecific hardware accelerators. Commun. ACM 63, 7 (2020), 48–57. https://doi.org/10.1145/3361682 [26] Martin Davis, George Logemann, and Donald Loveland. 1962. A machine program for theorem-proving. Commun. ACM 5, 7 (July 1962), 394–397. https://doi.org/10.1145/368273.368557 [27] Rémi Delmas, Zyad Hassan, Qinheping Hu, Rahul Kumar, Felipe R. Monteiro, Thanh Nguyen, Adrián Palacios, Celina Val, Michael Tautschnig, Justus Adam, Daniel Schwartz-Narbonne, and Carolyn Zech. 2026. Kani: A Model Checker for Rust. arXiv:2607.01504 [cs.SE] https://arxiv.org/abs/2607.01504 [28] Hang Ding, Mao Luo, Chu-Min Li, Shunwei Li, Runyao Chen, Caiquan Xiong, and Xinyun Wu. 2025. A Self-Optimizing Framework for SAT Solvers via Population Evolution and Large Language Model Collaboration. In Proceedings of SAT Competition 2025: Solver and Benchmark Descriptions (Proceedings of SAT Competitions), Cayden Codel, Katalin Fazekas, Marijn J. H. Heule, and Markus Iser (Eds.). TU Wien, Vienna, Austria, 15–17. https://doi.org/10.34726/10379 [29] Burak Ekici, Alain Mebsout, Cesare Tinelli, Chantal Keller, Guy Katz, Andrew Reynolds, and Clark Barrett. 2017. SMTCoq: A Plug-In for Integrating SMT Solvers into Coq. In Computer Aided Verification (CAV 2017), Part II (Lecture Notes in Computer Science, Vol. 10427). Springer, 126–133. https://doi.org/10.1007/978-3-319-63390-9_7 [30] Dawson R. Engler and M. Frans Kaashoek. 1995. Exterminate all operating system abstractions. In Proceedings 5th Workshop on Hot Topics in Operating Systems (HotOS-V). IEEE Computer Society, Los Alamitos, CA, USA, 78–83. https://doi.org/10.1109/hotos.1995.513459 [31] Dawson R. Engler, M. Frans Kaashoek, and James O’Toole, Jr. 1995. Exokernel: An operating system architecture for application-level resource management. ACM SIGOPS Operating Systems Review 29, 5 (1995), 251–266. https://doi.org/10.1145/224057.224076 [32] Niklas Eén and Niklas Sörensson. 2006. Translating Pseudo-Boolean Constraints into SAT. Journal on Satisfiability, Boolean Modeling and Computation 2, 1–4 (2006), 1–26. https://doi.org/10.3233/sat190014

[33] M. Frigo and S. G. Johnson. 2005. The Design and Implementation of FFTW3. Proc. IEEE 93, 2 (2005), 216–231. https://doi.org/10.1109/ jproc.2004.840301 [34] Nils Froleyks, Marijn Heule, Ashlin Iser, Matti Järvisalo, and Martin Suda. 2021. SAT Competition 2020. Artificial Intelligence 301 (2021), 103572. https://doi.org/10.1016/j.artint.2021.103572 [35] Robert Gallager. 1962. Low-density parity-check codes. IRE Transactions on Information Theory 8, 1 (1962), 21–28. https://doi.org/10. 1109/tit.1962.1057683 [36] Stephan Gocht, Ciaran McCreesh, and Jakob Nordström. 2022. An Auditable Constraint Programming Solver. In 28th International Conference on Principles and Practice of Constraint Programming (CP 2022) (Leibniz International Proceedings in Informatics (LIPIcs), Vol. 235). Schloss Dagstuhl – Leibniz-Zentrum für Informatik, 25:1–25:18. https://doi.org/10.4230/LIPIcs.CP.2022.25 [37] Stephan Gocht and Jakob Nordström. 2021. Certifying Parity Reasoning Efficiently Using Pseudo-Boolean Proofs. Proceedings of the AAAI Conference on Artificial Intelligence 35, 5 (2021), 3768–3777. https://doi.org/10.1609/aaai.v35i5.16494 [38] Evgueni Goldberg and Yakov Novikov. 2003. Verification of proofs of unsatisfiability for CNF formulas. In 2003 Design, Automation and Test in Europe Conference and Exhibition. IEEE Computer Society, Los Alamitos, CA, USA, 886–891. https://doi.org/10.1109/date.2003. 1253718 [39] Sumit Gulwani. 2011. Automating string processing in spreadsheets using input-output examples. ACM SIGPLAN Notices 46, 1 (2011), 317–330. https://doi.org/10.1145/1925844.1926423 [40] Sumit Gulwani, Oleksandr Polozov, and Rishabh Singh. 2017. Program synthesis. Foundations and Trends in Programming Languages 4, 1–2 (2017), 1–119. https://doi.org/10.1561/2500000010 [41] Andrew Haberlandt and Harrison Green. 2023. SBVA-CaDiCaL and SBVA-Kissat: Structured Bounded Variable Addition. In Proceedings of SAT Competition 2023: Solver, Benchmark and Proof Checker Descriptions (Department of Computer Science Series of Publications B, Vol. B2023-1). Department of Computer Science, University of Helsinki, Helsinki, Finland, 18. https://hdl.handle.net/10138/563824 [42] Rehan Hameed, Wajahat Qadeer, Megan Wachs, Omid Azizi, Alex Solomatnikov, Benjamin C. Lee, Stephen Richardson, Christos Kozyrakis, and Mark Horowitz. 2010. Understanding sources of inefficiency in general-purpose chips. In Proceedings of the 37th annual international symposium on Computer architecture. Association for Computing Machinery, New York, NY, USA, 37–47. https: //doi.org/10.1145/1815961.1815968 [43] Cheng-Shen Han and Jie-Hong Roland Jiang. 2012. When Boolean satisfiability meets Gaussian elimination in a simplex way. In Computer Aided Verification (Lecture Notes in Computer Science, Vol. 7358). Springer Berlin Heidelberg, Berlin, Heidelberg, 410–426. https: //doi.org/10.1007/978-3-642-31424-7_31 [44] John L. Hennessy and David A. Patterson. 2019. A new golden age for computer architecture. Commun. ACM 62, 2 (2019), 48–60. https://doi.org/10.1145/3282307 [45] Marijn J. H. Heule, Benjamin Kiesl, and Armin Biere. 2020. Strong Extension-Free Proof Systems. Journal of Automated Reasoning 64, 3 (2020), 533–554. https://doi.org/10.1007/s10817-019-09516-0 [46] Daniel Reiter Horn, Ken Elkabany, Chris Lesniewski-Laas, and Keith Winstein. 2017. The Design, Implementation, and Deployment of a System to Transparently Compress Hundreds of Petabytes of Image Files for a File-Storage Service. In 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 2017). USENIX Association, Boston, MA, USA, 1–15. https://www.usenix.org/ conference/nsdi17/technical-sessions/presentation/horn

Harrison Green, Claire Le Goues, and Fraser Brown

[47] Frank Hutter, Holger H. Hoos, and Kevin Leyton-Brown. 2011. Sequential Model-Based Optimization for General Algorithm Configuration. In Learning and Intelligent Optimization (Lecture Notes in Computer Science, Vol. 6683), Carlos A. Coello Coello (Ed.). Springer, Berlin, Heidelberg, 507–523. https://doi.org/10.1007/978-3-642-25566-3_40 [48] Frank Hutter, Holger H. Hoos, Kevin Leyton-Brown, and Thomas Stützle. 2009. ParamILS: an Automatic Algorithm Configuration Framework. Journal of Artificial Intelligence Research 36 (2009), 267– 306. https://doi.org/10.1613/jair.2861 [49] Alexey Ignatiev, Antonio Morgado, and Joao Marques-Silva. 2018. PySAT: A Python Toolkit for Prototyping with SAT Oracles. In Theory and Applications of Satisfiability Testing – SAT 2018 (Lecture Notes in Computer Science, Vol. 10929). Springer International Publishing, Cham, 428–437. https://doi.org/10.1007/978-3-319-94144-8_26 [50] Alexey Ignatiev, Zi Li Tan, and Christos Karamanos. 2024. Towards Universally Accessible SAT Technology. In 27th International Conference on Theory and Applications of Satisfiability Testing (SAT 2024) (Leibniz International Proceedings in Informatics (LIPIcs), Vol. 305), Supratik Chakraborty and Jie-Hong Roland Jiang (Eds.). Schloss Dagstuhl – Leibniz-Zentrum für Informatik, Dagstuhl, Germany, 16:1–16:11. https://doi.org/10.4230/LIPIcs.SAT.2024.16 [51] Yuka Ikarashi, Gilbert Louis Bernstein, Alex Reinking, Hasan Genc, and Jonathan Ragan-Kelley. 2022. Exocompilation for productive programming of hardware accelerators. In Proceedings of the 43rd ACM SIGPLAN International Conference on Programming Language Design and Implementation (San Diego, CA, USA) (PLDI 2022). Association for Computing Machinery, New York, NY, USA, 703–718. https://doi.org/10.1145/3519939.3523446 [52] Ashlin Iser and Christoph Jabs. 2024. Global Benchmark Database. In 27th International Conference on Theory and Applications of Satisfiability Testing (SAT 2024) (Leibniz International Proceedings in Informatics (LIPIcs), Vol. 305), Supratik Chakraborty and Jie-Hong Roland Jiang (Eds.). Schloss Dagstuhl – Leibniz-Zentrum für Informatik, Dagstuhl, Germany, 18:1–18:10. https://doi.org/10.4230/LIPIcs.SAT.2024.18 [53] Matti Järvisalo, Daniel Le Berre, Olivier Roussel, and Laurent Simon. 2012. The international SAT solver competitions. AI Magazine 33, 1 (2012), 89–94. https://doi.org/10.1609/aimag.v33i1.2395 [54] Susmit Jha, Sumit Gulwani, Sanjit A. Seshia, and Ashish Tiwari. 2010. Oracle-guided component-based program synthesis. In Proceedings of the 32nd ACM/IEEE International Conference on Software Engineering – Volume 1. Association for Computing Machinery, New York, NY, USA, 215–224. https://doi.org/10.1145/1806799.1806833 [55] Susmit Jha and Sanjit A. Seshia. 2017. A theory of formal synthesis via inductive learning. Acta Informatica 54, 7 (2017), 693–726. https: //doi.org/10.1007/s00236-017-0294-5 [56] Rajeev Joshi, Greg Nelson, and Keith Randall. 2002. Denali: a goaldirected superoptimizer. ACM SIGPLAN Notices 37, 5 (May 2002), 304–314. https://doi.org/10.1145/543552.512566 [57] Norman P. Jouppi, Cliff Young, Nishant Patil, David Patterson, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, Rick Boyle, Pierre-luc Cantin, Clifford Chao, Chris Clark, Jeremy Coriell, Mike Daley, Matt Dau, Jeffrey Dean, Ben Gelb, Tara Vazir Ghaemmaghami, Rajendra Gottipati, William Gulland, Robert Hagmann, C. Richard Ho, Doug Hogberg, John Hu, Robert Hundt, Dan Hurt, Julian Ibarz, Aaron Jaffey, Alek Jaworski, Alexander Kaplan, Harshit Khaitan, Daniel Killebrew, Andy Koch, Naveen Kumar, Steve Lacy, James Laudon, James Law, Diemthu Le, Chris Leary, Zhuyuan Liu, Kyle Lucke, Alan Lundin, Gordon MacKean, Adriana Maggiore, Maire Mahony, Kieran Miller, Rahul Nagarajan, Ravi Narayanaswami, Ray Ni, Kathy Nix, Thomas Norrie, Mark Omernick, Narayana Penukonda, Andy Phelps, Jonathan Ross, Matt Ross, Amir Salek, Emad Samadiani, Chris Severn, Gregory Sizikov, Matthew Snelham, Jed Souter, Dan Steinberg, Andy Swing, Mercedes Tan, Gregory Thorson, Bo Tian, Horia Toma, Erick Tuttle, Vijay Vasudevan,

Richard Walter, Walter Wang, Eric Wilcox, and Doe Hyun Yoon. 2017. In-datacenter performance analysis of a tensor processing unit. In Proceedings of the 44th annual international symposium on computer architecture. Association for Computing Machinery, New York, NY, USA, 1–12. https://doi.org/10.1145/3079856.3080246 [58] Serdar Kadioglu, Yuri Malitsky, Ashish Sabharwal, Horst Samulowitz, and Meinolf Sellmann. 2011. Algorithm Selection and Scheduling. In Principles and Practice of Constraint Programming – CP 2011 (Lecture Notes in Computer Science, Vol. 6876), Jimmy Lee (Ed.). Springer Berlin Heidelberg, Berlin, Heidelberg, 454–469. https://doi.org/10.1007/9783-642-23786-7_35 [59] Serdar Kadioglu, Yuri Malitsky, Meinolf Sellmann, and Kevin Tierney. 2010. ISAC – Instance-Specific Algorithm Configuration. In ECAI 2010: 19th European Conference on Artificial Intelligence (Frontiers in Artificial Intelligence and Applications, Vol. 215). IOS Press, Amsterdam, Netherlands, 751–756. https://doi.org/10.3233/978-1-60750606-5-751 [60] Moein Khazraee, Lu Zhang, Luis Vega, and Michael Bedford Taylor. 2017. Moonwalk: NRE optimization in ASIC clouds. In Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating Systems. Association for Computing Machinery, New York, NY, USA, 511–526. https://doi. org/10.1145/3037697.3037749 [61] Ashiqur R. KhudaBukhsh, Lin Xu, Holger H. Hoos, and Kevin LeytonBrown. 2016. SATenstein: Automatically building local search SAT solvers from components. Artificial Intelligence 232 (2016), 20–42. https://doi.org/10.1016/j.artint.2015.11.002 [62] Donald E. Knuth. 2015. The Art of Computer Programming, Volume 4, Fascicle 6: Satisfiability. Addison-Wesley Professional, Boston, MA, USA. https://www.informit.com/store/art-of-computerprogramming-volume-4-fascicle-6-satisfiability-9780134397603 [63] Wietze Koops, Daniel Le Berre, Magnus O. Myreen, Jakob Nordström, Andy Oertel, Yong Kiam Tan, and Marc Vinyals. 2025. Practically Feasible Proof Logging for Pseudo-Boolean Optimization. In 31st International Conference on Principles and Practice of Constraint Programming (CP 2025) (Leibniz International Proceedings in Informatics (LIPIcs), Vol. 340). Schloss Dagstuhl – Leibniz-Zentrum für Informatik, 21:1–21:27. https://doi.org/10.4230/LIPIcs.CP.2025.21 [64] Peter Lammich. 2023. GRAT: A Formally Verified (UN)SAT Proof Checker. In Proceedings of SAT Competition 2023: Solver, Benchmark and Proof Checker Descriptions (Department of Computer Science Series of Publications B, Vol. B-2023-1). University of Helsinki, Helsinki, Finland, 82–85. https://hdl.handle.net/10138/563824 [65] Robert Tjarko Lange, Qi Sun, Aaditya Prasad, Maxence Faldor, Yujin Tang, and David Ha. 2025. Towards Robust Agentic CUDA Kernel Benchmarking, Verification, and Optimization. arXiv:2509.14279 https://arxiv.org/abs/2509.14279 [66] Jiale Lao and Immanuel Trummer. 2026. Demonstrating GenDB: Instance-Optimized and Customized Query Processing Code Generation via LLM Agents. Proceedings of the VLDB Endowment 19, 12 (2026), 4602–4605. https://doi.org/10.14778/3827998.3828076 [67] Jia Hui Liang, Vijay Ganesh, Pascal Poupart, and Krzysztof Czarnecki. 2016. Learning rate based branching heuristic for SAT solvers. In Theory and Applications of Satisfiability Testing – SAT 2016 (Lecture Notes in Computer Science, Vol. 9710). Springer International Publishing, Cham, 123–140. https://doi.org/10.1007/978-3-319-40970-2_9 [68] Erik Lindholm, John Nickolls, Stuart Oberman, and John Montrym. 2008. NVIDIA Tesla: A unified graphics and computing architecture. IEEE Micro 28, 2 (2008), 39–55. https://doi.org/10.1109/mm.2008.31 [69] Fei Liu, Xialiang Tong, Mingxuan Yuan, Xi Lin, Fu Luo, Zhenkun Wang, Zhichao Lu, and Qingfu Zhang. 2024. Evolution of heuristics: towards efficient automatic algorithm design using large language model. In Proceedings of the 41st International Conference on Machine Learning (Vienna, Austria) (Proceedings of Machine Learning Research,

The Case for Automated Hyperspecialization: Evidence from SAT Vol. 235). PMLR, Vienna, Austria, 32201–32223. https://proceedings. mlr.press/v235/liu24bs.html [70] Fei Liu, Yiming Yao, Ping Guo, Zhiyuan Yang, Xi Lin, Zhe Zhao, Xialiang Tong, Kun Mao, Zhichao Lu, Zhenkun Wang, Mingxuan Yuan, and Qingfu Zhang. 2026. A Systematic Survey on Large Language Models for Algorithm Design. Comput. Surveys 58, 8, Article 218 (June 2026), 32 pages. https://doi.org/10.1145/3787585 [71] Nuno P. Lopes, Juneyoung Lee, Chung-Kil Hur, Zhengyang Liu, and John Regehr. 2021. Alive2: bounded translation validation for LLVM. In Proceedings of the 42nd ACM SIGPLAN International Conference on Programming Language Design and Implementation (Virtual, Canada) (PLDI 2021). Association for Computing Machinery, New York, NY, USA, 65–79. https://doi.org/10.1145/3453483.3454030 [72] Manuel López-Ibáñez, Jérémie Dubois-Lacoste, Leslie Pérez Cáceres, Mauro Birattari, and Thomas Stützle. 2016. The irace package: Iterated racing for automatic algorithm configuration. Operations Research Perspectives 3 (2016), 43–58. https://doi.org/10.1016/j.orp. 2016.09.002 [73] Anil Madhavapeddy, Richard Mortier, Charalampos Rotsos, David Scott, Balraj Singh, Thomas Gazagnaire, Steven Smith, Steven Hand, and Jon Crowcroft. 2013. Unikernels: Library Operating Systems for the Cloud. In Proceedings of the eighteenth international conference on Architectural support for programming languages and operating systems. Association for Computing Machinery, New York, NY, USA, 461–472. https://doi.org/10.1145/2451116.2451167 [74] J. P. Marques-Silva and K. A. Sakallah. 1999. GRASP: a search algorithm for propositional satisfiability. IEEE Trans. Comput. 48, 5 (1999), 506–521. https://doi.org/10.1109/12.769433 [75] Henry Massalin. 1987. Superoptimizer: a look at the smallest program. ACM SIGARCH Computer Architecture News 15, 5 (Oct. 1987), 122–126. https://doi.org/10.1145/36177.36194 [76] Kevin Meng, Vincent Huang, Jacob Steinhardt, and Sarah Schwettmann. 2025. Introducing Docent. Transluce. https: //transluce.org/docent/blog/introducing-docent [77] Antonio Morgado, Alexey Ignatiev, and Joao Marques-Silva. 2015. MSCG: Robust Core-Guided MaxSAT Solving: System Description. Journal on Satisfiability, Boolean Modeling and Computation 9, 1 (2015), 129–134. https://doi.org/10.3233/sat190105 [78] Matthew W. Moskewicz, Conor F. Madigan, Ying Zhao, Lintao Zhang, and Sharad Malik. 2001. Chaff: engineering an efficient SAT solver. In Proceedings of the 38th Annual Design Automation Conference (Las Vegas, Nevada, USA) (DAC ’01). Association for Computing Machinery, New York, NY, USA, 530–535. https://doi.org/10.1145/378239.379017 [79] Saeed Nejati and Vijay Ganesh. 2019. CDCL(Crypto) SAT Solvers for Cryptanalysis. In Proceedings of the 29th Annual International Conference on Computer Science and Software Engineering (CASCON ’19). IBM Corp., Riverton, NJ, USA, 311–316. https://arxiv.org/abs/ 2005.13415 [80] Aina Niemetz, Mathias Preiner, Florian Lonsing, Martina Seidl, and Armin Biere. 2012. Resolution-Based Certificate Extraction for QBF (Tool Presentation). In Theory and Applications of Satisfiability Testing – SAT 2012 (Lecture Notes in Computer Science, Vol. 7317). Springer, 430–435. https://doi.org/10.1007/978-3-642-31612-8_33 [81] Alexander Novikov, Ngân Vũ, Marvin Eisenberger, Emilien Dupont, Po-Sen Huang, Adam Zsolt Wagner, Sergey Shirobokov, Borislav Kozlovskii, Francisco J. R. Ruiz, Abbas Mehrabian, M. Pawan Kumar, Abigail See, Swarat Chaudhuri, George Holland, Alex Davies, Sebastian Nowozin, Pushmeet Kohli, and Matej Balog. 2025. AlphaEvolve: A coding agent for scientific and algorithmic discovery. arXiv:2506.13131 [cs.AI] https://arxiv.org/abs/2506.13131 [82] Toru Ogawa, Yangyang Liu, Ryuzo Hasegawa, Miyuki Koshimura, and Hiroshi Fujita. 2013. Modulo based CNF encoding of cardinality constraints and its application to MaxSAT solvers. In 2013 IEEE 25th International Conference on Tools with Artificial Intelligence. IEEE,

Piscataway, NJ, USA, 9–17. https://doi.org/10.1109/ictai.2013.13 [83] OpenAI. 2026. GPT-5.6: Frontier Intelligence That Scales with Your Ambition. OpenAI. https://openai.com/index/gpt-5-6/ GPT-5.6 Sol. [84] OpenAI. 2026. Jalapeño’s First Results Show Industry-Leading Speed and Efficiency in AI Inference. OpenAI. https://openai.com/index/ jalapeno-first-results/ [85] OpenTitan Contributors. [n. d.]. OpenTitan Documentation. OpenTitan. Retrieved September 9, 2026 from https://opentitan.org/book/ [86] Anne Ouyang, Simon Guo, Simran Arora, Alex L. Zhang, William Hu, Christopher Ré, and Azalia Mirhoseini. 2025. KernelBench: can LLMs write efficient GPU kernels?. In Proceedings of the 42nd International Conference on Machine Learning (Vancouver, Canada) (Proceedings of Machine Learning Research, Vol. 267). PMLR, Vancouver, Canada, 47356–47415. https://proceedings.mlr.press/v267/ouyang25a.html [87] Tuomas Pelkonen, Scott Franklin, Justin Teller, Paul Cavallaro, Qi Huang, Justin Meza, and Kaushik Veeraraghavan. 2015. Gorilla: A Fast, Scalable, In-Memory Time Series Database. Proceedings of the VLDB Endowment 8, 12 (2015), 1816–1827. https://doi.org/10.14778/ 2824032.2824078 [88] Amir Pnueli, Michael Siegel, and Eli Singerman. 1998. Translation validation. In International Conference on Tools and Algorithms for the Construction and Analysis of Systems (Lecture Notes in Computer Science, Vol. 1384). Springer Berlin Heidelberg, Berlin, Heidelberg, 151–166. https://doi.org/10.1007/bfb0054170 [89] Oleksandr Polozov and Sumit Gulwani. 2015. FlashMeta: A Framework for Inductive Program Synthesis. In Proceedings of the 2015 ACM SIGPLAN International Conference on Object-Oriented Programming, Systems, Languages, and Applications. Association for Computing Machinery, New York, NY, USA, 107–126. https://doi.org/10.1145/ 2814270.2814310 [90] M. Püschel, J. M. F. Moura, J. R. Johnson, D. Padua, M. M. Veloso, B. W. Singer, Jianxin Xiong, F. Franchetti, A. Gacic, Y. Voronenko, K. Chen, R. W. Johnson, and N. Rizzolo. 2005. SPIRAL: Code Generation for DSP Transforms. Proc. IEEE 93, 2 (2005), 232–275. https://doi. org/10.1109/jproc.2004.840306 [91] Jonathan Ragan-Kelley, Connelly Barnes, Andrew Adams, Sylvain Paris, Frédo Durand, and Saman Amarasinghe. 2013. Halide: a language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines. In Proceedings of the 34th ACM SIGPLAN Conference on Programming Language Design and Implementation (Seattle, Washington, USA) (PLDI ’13). Association for Computing Machinery, New York, NY, USA, 519–530. https://doi.org/10.1145/2491956.2462176 [92] Joseph E. Reeves. 2025. Cardinality Constraints in Boolean Satisfiability Solving. Ph.D. thesis. Carnegie Mellon University, Pittsburgh, PA, USA. https://www.cs.cmu.edu/~jereeves/research/jereeves_phd_ csd_2025.pdf Technical Report CMU-CS-25-147. [93] Jaime Rodríguez-Guerra and Jannis Leidel. 2023. Conda 23.10.0: libmamba is now the default solver. Conda Blog. https://conda.org/ blog/2023-11-06-conda-23-10-0-release/ Accessed: 2026-08-29. [94] Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Matej Balog, M. Pawan Kumar, Emilien Dupont, Francisco J. R. Ruiz, Jordan S. Ellenberg, Pengming Wang, Omar Fawzi, Pushmeet Kohli, and Alhussein Fawzi. 2024. Mathematical discoveries from program search with large language models. Nature 625, 7995 (2024), 468–475. https://doi.org/10.1038/s41586-023-06924-6 [95] Eric Schkufza, Rahul Sharma, and Alex Aiken. 2013. Stochastic superoptimization. ACM SIGARCH Computer Architecture News 41, 1 (2013), 305–316. https://doi.org/10.1145/2490301.2451150 [96] Junjie Sheng, Yanqiu Lin, Jiehao Wu, Yanhong Huang, Jianqi Shi, Min Zhang, and Xiangfeng Wang. 2025. SolSearch: An LLM-Driven Framework for Efficient SAT-Solving Code Generation. In 2025 IEEE/ACM 47th International Conference on Software Engineering: New Ideas and Emerging Results (ICSE-NIER). IEEE, Piscataway, NJ, USA, 6–10.

Harrison Green, Claire Le Goues, and Fraser Brown https://doi.org/10.1109/icse-nier66352.2025.00007 [97] Carsten Sinz. 2005. Towards an Optimal CNF Encoding of Boolean Cardinality Constraints. In Principles and Practice of Constraint Programming - CP 2005 (Lecture Notes in Computer Science, Vol. 3709), Peter van Beek (Ed.). Springer Berlin Heidelberg, Berlin, Heidelberg, 827–831. https://doi.org/10.1007/11564751_73 [98] Armando Solar-Lezama, Liviu Tancau, Rastislav Bodik, Sanjit Seshia, and Vijay Saraswat. 2006. Combinatorial sketching for finite programs. In Proceedings of the 12th international conference on Architectural support for programming languages and operating systems. Association for Computing Machinery, New York, NY, USA, 404–415. https://doi.org/10.1145/1168857.1168907 [99] Mate Soos, Karsten Nohl, and Claude Castelluccia. 2009. Extending SAT Solvers to Cryptographic Problems. In Theory and Applications of Satisfiability Testing - SAT 2009, 12th International Conference, SAT 2009, Swansea, UK, June 30 - July 3, 2009. Proceedings (Lecture Notes in Computer Science, Vol. 5584), Oliver Kullmann (Ed.). Springer, Berlin, Heidelberg, 244–257. https://doi.org/10.1007/978-3-642-02777-2_24 [100] Michael Stonebraker, Daniel J. Abadi, Adam Batkin, Xuedong Chen, Mitch Cherniack, Miguel Ferreira, Edmond Lau, Amerson Lin, Samuel Madden, Elizabeth J. O’Neil, Pat O’Neil, Alex Rasin, Nga Tran, and Stan Zdonik. 2005. C-Store: A Column-oriented DBMS. In Proceedings of the 31st International Conference on Very Large Data Bases. VLDB Endowment, Trondheim, Norway, 553–564. https://www.vldb.org/ archives/website/2005/program/paper/thu/p553-stonebraker.pdf [101] Michael Stonebraker and Uğur Çetintemel. 2005. “One Size Fits All”: An Idea Whose Time Has Come and Gone. In 21st International Conference on Data Engineering (ICDE’05). IEEE, Piscataway, NJ, USA, 2–11. https://doi.org/10.1109/icde.2005.1 [102] Yiwen Sun, Furong Ye, Zhihan Chen, Ke Wei, and Shaowei Cai. 2025. Discovering heuristics in a complex SAT solver with large language models. arXiv:2507.22876 [cs.AI] https://arxiv.org/abs/2507.22876 Revised June 6, 2026. [103] Yiwen Sun, Furong Ye, Xianyin Zhang, Shiyu Huang, Bingzhen Zhang, Ke Wei, and Shaowei Cai. 2026. AutoSAT: Automatically Optimize SAT Solvers via Large Language Models. Journal of Artificial Intelligence Research 86, Article 29 (July 2026), 50 pages. https://doi.org/10.1613/jair.1.20499 [104] Yong Kiam Tan, Marijn J. H. Heule, and Magnus O. Myreen. 2023. Verified LRAT and LPR Proof Checking with cake_lpr. In Proceedings of SAT Competition 2023: Solver, Benchmark and Proof Checker Descriptions (Department of Computer Science Series of Publications B, Vol. B-2023-1). University of Helsinki, Helsinki, Finland, 89–90. https://hdl.handle.net/10138/563824 [105] The OpenROAD Project. [n. d.]. OpenROAD Flow Scripts Documentation. The OpenROAD Project. Retrieved September 9, 2026 from https://openroad-flow-scripts.readthedocs.io/en/latest/index2. html [106] Emina Torlak and Rastislav Bodik. 2014. A lightweight symbolic virtual machine for solver-aided host languages. ACM SIGPLAN Notices 49, 6 (2014), 530–541. https://doi.org/10.1145/2666356.2594340 [107] Alasdair Urquhart. 1987. Hard examples for resolution. J. ACM 34, 1 (1987), 209–219. https://doi.org/10.1145/7531.8928 [108] Haoyu Wang, Yuliang Song, Tao Li, Zhiwei Deng, Yaqing Wang, Deepak Ramachandran, Eldan Cohen, and Dan Roth. 2026. Formalize, Don’t Optimize: The Heuristic Trap in LLM-Generated Combinatorial Solvers. arXiv:2605.12421 [cs.AI] https://arxiv.org/abs/2605.12421 [109] Johannes Wehrstein, Timo Eckmann, Matthias Jasny, and Carsten Binnig. 2026. Bespoke OLAP: Synthesizing Workload-Specific Onesize-fits-one Database Engines. Proceedings of the VLDB Endowment 19, 11 (2026), 3759–3771. https://doi.org/10.14778/3836663.3836723 [110] Anjiang Wei, Tarun Suresh, Huanmi Tan, Yinglun Xu, Gagandeep Singh, Ke Wang, and Alex Aiken. 2025. SuperCoder: Assembly Program Superoptimization with Large Language Models.

arXiv:2505.11480 [cs.CL] https://arxiv.org/abs/2505.11480 Revised August 8, 2026. [111] Junyu Wei, Guangyan Zhang, Yang Wang, Zhiwei Liu, Zhanyang Zhu, Junchao Chen, Tingtao Sun, and Qi Zhou. 2021. On the Feasibility of Parser-based Log Compression in Large-Scale Cloud Systems. In 19th USENIX Conference on File and Storage Technologies (FAST 2021). USENIX Association, Berkeley, CA, USA, 249–262. https: //www.usenix.org/conference/fast21/presentation/wei [112] Nathan Wetzler, Marijn J. H. Heule, and Warren A. Hunt, Jr. 2014. DRAT-trim: Efficient Checking and Trimming Using Expressive Clausal Proofs. In Theory and Applications of Satisfiability Testing – SAT 2014 (Lecture Notes in Computer Science, Vol. 8561). Springer, Cham, 422–429. https://doi.org/10.1007/978-3-319-09284-3_31 [113] R. Clint Whaley and Jack J. Dongarra. 1998. Automatically Tuned Linear Algebra Software. In SC ’98: Proceedings of the 1998 ACM/IEEE Conference on Supercomputing. IEEE, Piscataway, NJ, USA, 38. https: //doi.org/10.1109/sc.1998.10004 [114] Lin Xu, Holger Hoos, and Kevin Leyton-Brown. 2010. Hydra: Automatically Configuring Algorithms for Portfolio-Based Selection. Proceedings of the AAAI Conference on Artificial Intelligence 24, 1 (Jul. 2010), 210–216. https://doi.org/10.1609/aaai.v24i1.7565 [115] Lin Xu, Frank Hutter, Holger H. Hoos, and Kevin Leyton-Brown. 2008. SATzilla: Portfolio-Based Algorithm Selection for SAT. Journal of Artificial Intelligence Research 32 (2008), 565–606. https://doi.org/ 10.1613/jair.2490 [116] Haoran Ye, Jiarui Wang, Zhiguang Cao, Federico Berto, Chuanbo Hua, Haeyeon Kim, Jinkyoo Park, and Guojie Song. 2024. ReEvo: large language models as hyper-heuristics with reflective evolution. In Advances in Neural Information Processing Systems (Vancouver, BC, Canada), Vol. 37. Curran Associates Inc., Red Hook, NY, USA, 43571–43608. https://doi.org/10.52202/079017-1381 [117] Cunxi Yu, Rongjian Liang, Chia-Tung Ho, and Haoxing Ren. 2025. Autonomous Code Evolution Meets NP-Completeness. arXiv:2509.07367 [cs.AI] https://arxiv.org/abs/2509.07367 [118] Lintao Zhang and Sharad Malik. 2003. Validating SAT solvers using an independent resolution-based checker: Practical implementations and other applications. In 2003 Design, Automation and Test in Europe Conference and Exhibition. IEEE Computer Society, Los Alamitos, CA, USA, 880–885. https://doi.org/10.1109/date.2003.1253717

The Case for Automated Hyperspecialization: Evidence from SAT

A

Agent prompt

Each agent is prompted with the core task prompt (§A.1) plus a proof instruction extension depending on the selected proof format: DPR: §A.2, GRAT: §A.3, or VeriPB: §A.4. In addition, we also provide condensed documentation for each of the formats, mounted in the container at /opt/autospec/docs (not shown here). A.1

Task definition

Mission: build an efficient, certifying solver specialized to the single target family `{family}`. This is deliberately not a general-purpose SAT-solver task. You are not required to use DPLL, CDCL, clause learning, or any other traditional SAT architecture, and your solver need not work on unrelated formulas. The intended advantage is to reverse-engineer this family's encoding, recover the high-level mathematical problem and generator structure, and solve that problem directly. Be creative and aggressively family-specific while remaining sound. Specialization principles: - Infer reusable family structure, not instance identities. It is valid to depend on a detected encoding, generator invariant, parameter regime, repeated block layout, or high-level theorem shared by the family. - Consider direct constructions, graph algorithms, matching or flow, dynamic programming, algebra, Gaussian elimination, bitset propagation, interval reasoning, symmetry reduction, bounded search over recovered objects, decomposition, or a purpose-built decision procedure. A conventional SAT core is optional, not the default objective. - Build conservative structural detectors and validate every assumed invariant across diverse training instances. If an input does not match the supported structure, return `s UNKNOWN`; never stretch an unsound fast path. - Treat CNF-to-problem and problem-to-certificate mappings as first-class algorithm components. For SAT, reconstruct a complete DIMACS assignment and independently check every original clause. For UNSAT, design the selected proof strategy alongside the high-level algorithm rather than after it. - Optimize the whole evaluated path: parsing, detection, solving, model/proof generation, proof size, elaboration cost, and verified-kernel cost. Once the algorithm is right, profile representative sizes and use appropriate compiled, cache-friendly, bit-parallel, or incremental data structures. Recommended workflow: 1. Survey multiple small, medium, and large training formulas. Measure variable/clause counts, clause-width and polarity distributions, units/binaries, components, variable degrees, repeated strides/blocks, and signatures of gates, XORs, cardinalities, grids, graphs, permutations, schedules, or time steps. 2. Form explicit hypotheses about high-level objects and the variable/clause mapping. Write small analysis tools when helpful, and falsify each hypothesis on the full training set instead of trusting one attractive example. 3. Prototype the direct high-level solver and a conservative recognizer. Check high-level answers against the original CNF before investing in optimization. 4. Prototype certificate production on tiny self-authored proof unit tests (not additional family benchmarks) with the installed checker pipeline. Establish a proof-generating invariant that scales with the specialized algorithm. 5. Implement the competition package, test locally on representative training formulas, submit, inspect per-instance status and timing, and iterate. Fix every invalid model, invalid proof, crash, timeout, or proof-checker resource failure before chasing marginal speed. 6. Leave a concise README in the solver directory explaining the recovered encoding, supported detector, algorithm, certificate construction, experiments, limitations, and promising next steps for future runs. Scope and integrity rules: - `/data/train` is the complete, read-only training set authorized for this run. Do not search for, request, download, reconstruct, or synthesize additional benchmark formulas, even when web search is enabled. - Hidden validation instances, membership, labels, runtimes, metrics, failures, job details, and storage locations are private. Do not attempt to infer or access them. - Do not memorize, branch on, or encode training filenames, hashes, comments, literal sequences, or individual training instances. Build family-level algorithmic improvements only. - Imported provenance labels are not correctness oracles. Return SAT/UNSAT only with a certificate accepted by the selected checker; otherwise print `s UNKNOWN`. A ready-to-edit C++20 scaffold is installed at `/workspace/solver`. Work in that directory rather than starting elsewhere. Its only required contract files are `README.md`, executable `build.sh`, and executable `run.sh`; keep the README current with the recovered structure, algorithm, certificate construction, experiments, limitations, and next steps. The Ubuntu 24.04 agent environment includes Python 3.12 (with development headers and venv support), Node/npm, Clang/Clang++ 16, LLD 16, CMake, Ninja, Make, GDB, strace, `time`, Git, curl, ripgrep (`rg`), `fd`, jq, `file`, `xxd`, `pdftotext`, `lsof`, `bc`, `tree`, `tar`, `unzip`, `zip`, `bzip2`, and `xz`, in addition to the selected proof tools. Clang's AddressSanitizer runtime is installed, so local manual debug builds with `-fsanitize=address` are supported; restore the normal optimized `build.sh` before submission. These tools are available for local analysis, debugging, and builds; submitted solvers must still be self-contained for the smaller worker runtime. The `kissat` and `cadical` binaries are installed for local experiments, hypothesis checks, and comparison on authorized training formulas. They are local analysis tools only: never invoke either binary from `run.sh`, the submitted solver, or any subprocess during actual solving, and do not make the submitted artifact depend on them.

Harrison Green, Claire Le Goues, and Fraser Brown The submission host invokes `build.sh` exactly once as `build.sh <artifact-dir>` inside the local `linux/amd64` builder container. The scaffold compiles with the worker-matched flags `-std=c++20 -O3 -march=icelake-server -mtune=icelake-server -DNDEBUG -pthread`; do not use `-march=native`. Put the executable solver and any runtime-only files in the supplied artifact directory. GCP workers receive that precompiled artifact, do not invoke `build.sh`, and call only: run.sh <formula.cnf> <output-dir> Write only SAT Competition lines to stdout: `c ...`, exactly one `s SATISFIABLE`, `s UNSATISFIABLE`, or `s UNKNOWN`, and `v ... 0` for SAT. The exact SAT reference is `/opt/autospec/docs/proof-formats/sat.md`. Every SAT assignment line written by the solver must retain its `v ` prefix; literals may be wrapped across multiple `v` lines or placed on one long line, and the broker imposes no per-line character limit. Terminate the whole model exactly once with `0` and emit no later literals. The broker validates this stdout, strips the `s`/`v` prefixes, and gives `gratchk sat` a prefix-free raw literal file. Therefore never pass prefixed solver stdout directly to `gratchk`; capture the real stdout and run `autospec-check-sat formula.cnf solver.stdout` for the broker-equivalent local check. For UNSAT, write the selected proof syntax to `<output-dir>/proof.out`. Return exit code 0 after every well-formed SAT, UNSAT, or UNKNOWN response; any other exit code is a solver crash. Evaluation worker resources: - Each evaluation task is provisioned with {evaluation_cpu_cores:g} {evaluation_cpu_label} and {evaluation_memory_gib:g} GiB of RAM. The solver and subsequent verifier stages run sequentially within that task allocation. {evaluation_cpu_guidance} - The local agent container and the 4-vCPU builder can expose different resources. Optimize and size the submitted runtime for the evaluation allocation, especially peak solver/proof memory. Evaluation time budgets: - The evaluation runner externally stops each solver invocation after {solver_timeout_seconds} seconds. This is a development-efficiency cutoff that avoids spending evaluation capacity on hard instances, not a desired stopping point or a limit on how deployed artifacts may be used. - Do not add, tune, or preserve an internal watchdog, safety margin, or elapsed-time check that returns `s UNKNOWN` or otherwise gives up near this cutoff. In particular, do not stop a few seconds early to anticipate the evaluator. Keep making sound progress until the external runner stops the process so the artifact remains useful under longer deployment limits. - Certificate verification has a separate {verifier_timeout_seconds}-second total wall-clock budget shared by elaboration and the verified kernel. Verification time is reported separately and is excluded from the solver score. Agent effort budgets: - Work diligently throughout the available budget. This run has {active_time_hours:g} hours of active work and at most {max_evaluations} accepted evaluation submissions. Active time includes reasoning, local analysis, implementation, builds, local tests, MCP calls other than evaluation waiting, and resume overhead. Time spent blocked inside `wait_for_completion` is paused and does not consume active work. - Call `get_budget` whenever useful to see exact used and remaining active seconds and evaluation submissions. The supervisor interrupts the session when active time is exhausted. If a turn ends before 95% of active time is used and evaluation capacity remains, it automatically resumes the same session with the remaining budget. - Do not stop after one plausible approach or one good score. Use the budget for careful structural investigation, correctness testing, profiling, and measured iterations. An evaluation request beyond the submission limit is rejected and the session is stopped. - Every accepted submission is a checkpoint. The system automatically selects the best-performing correct checkpoint across the run, so the last submission does not need to restate or recreate the best earlier solver. Focus each submission on learning or improving performance while guaranteeing correctness across all supported inputs. Use the bound MCP tools `submit_solver`, `get_status`, `get_budget`, `wait_for_completion`, and `get_results`. Submit `/workspace/solver`. They evaluate training only. Each run permits one active training evaluation at a time. After submitting, call `wait_for_completion` with no arguments; it follows the current training evaluation and waits for up to three hours, so do not poll `get_status` while that evaluation is active. Every submission also triggers an independent private validation run; that run never blocks another training submission.

A.2

DPR extension

Selected proof contract: DPR. - The comprehensive local reference is `/opt/autospec/docs/proof-formats/dpr.md`. Read it before implementing the proof logger and consult it whenever syntax or elaboration is unclear. - Emit textual DPR into `proof.out`; each addition is `<clause> 0` or `<clause> <witness> 0`, and each deletion is `d <clause> 0`. - A PR witness has no separator token: it must start by repeating the first clause literal. Example: `-4 -17 -4 -17 1 20 0` has clause `-4 -17` and witness `-4 -17 1 20`. - A witness-free addition is checked as RUP and then RAT; for RAT the first literal is the pivot. A witnessed non-RUP addition uses propagation redundancy. Deletions name clauses by contents, not IDs. - Prefer simple RUP steps when possible, use PR witnesses for a tested family-level repair/symmetry map, terminate every step with `0`, and normally derive the empty clause with `0`. - Locally elaborate: `dpr-trim formula.cnf proof.out -L elaborated.lpr`. - Check the verified kernel proof: `cake_lpr formula.cnf elaborated.lpr`. - Inspect `elaborated.lpr` to debug IDs, pivots, witnesses, and propagation/candidate hint blocks. Only the exact kernel line `s VERIFIED UNSAT` establishes success; elaborator acceptance is not a verified certificate.

The Case for Automated Hyperspecialization: Evidence from SAT

A.3

GRAT extension

Selected proof contract: GRAT. - The comprehensive local reference is `/opt/autospec/docs/proof-formats/grat.md`. Read it before implementing the proof logger and consult it whenever syntax or elaboration is unclear. - Emit ASCII textual DRAT into `proof.out`, not a GRAT kernel file and not binary DRAT. Additions are `<clause literals> 0`; deletions are `d <clause literals> 0`; `0` alone adds the empty clause. - Additions are checked first by RUP. If RUP fails, the first literal is the RAT pivot and all required opposite-pivot candidates must be justified. DRAT has no explicit PR witness or PB arithmetic. - Deletions name live clauses by contents, not IDs. Get the proof working without deletion first, then delete only after the last use of a clause as a propagation reason or RAT candidate. - Locally elaborate: `gratgen formula.cnf proof.out -o elaborated.gratp -l elaborated.gratl`. - Check the verified proof: `gratchk unsat formula.cnf elaborated.gratl elaborated.gratp`. - `gratgen` may report `s VERIFIED` on stderr; that is not the final result. Only the exact `gratchk` stdout line `s VERIFIED UNSAT` establishes success.

A.4

VeriPB extension

Selected proof contract: VeriPB. - The comprehensive local reference is `/opt/autospec/docs/proof-formats/veripb.md`. Read it before implementing the proof logger and consult it whenever syntax, IDs, arithmetic, or subproofs are unclear. - Emit an augmented pseudo-Boolean proof version 3.0 into `proof.out`. The required outer structure is: `pseudo-Boolean proof version 3.0` `f <number-of-input-clauses>;` `<derivation rules, each ending in ;>` `output NONE;` `conclusion UNSAT : <live-contradiction-id>;` `end pseudo-Boolean proof;` - Input clauses receive IDs `1..N`; every constraint-adding derivation receives the next ID, while deletions do not. CNF literal `k` is `xk`, literal `-k` is `~xk`, and a clause becomes `+1 <lit> ... >= 1`. - Use `rup <constraint>;` for PB reverse unit propagation and `pol <RPN expression>;` for cutting-planes arithmetic. For example, opposite unit constraints at IDs 5 and 6 yield a contradiction with `pol 5 6 +;`. - Augmented conveniences include `ia` for syntactic implication/addition, `red` with a substitution witness for satisfiability-preserving strengthening, `pbc` subproofs, order-backed `dom`, and `del id <ids>;`. Let VeriPB elaborate these to explicit CakePB kernel steps and ordered RUP hints. - Locally elaborate with the exact worker order: `veripb -u --cnf --elaborate elaborated.pbp formula.cnf proof.out`. - Check the verified kernel proof: `cake_pb_cnf formula.cnf elaborated.pbp`. - Inspect the elaborated proof to debug absolute IDs and annotations. Only frontend `s VERIFIED UNSATISFIABLE` followed by kernel `s VERIFIED UNSAT` establishes success.

Harrison Green, Claire Le Goues, and Fraser Brown

B

Extended performance results

In the main body of RQ1 (§4.2) we evaluate the top specialist against the per-family VBS of the four core baselines. This serves as a sort of lower bound on the potential of hyperspecialization: How effective are the best specialists compared to a strong baseline? Here we include a complete enumeration of head-to-head results, visualized in a similar format. Figure 9 shows the perfamily head-to-head performance of specialists constructed for every format (DPR, VeriPB, GRAT) as well as the best specialist compared to: each of the four core baselines individually, the per-family virtual-best of the four core baselines, and three additional baselines from related work in LLMbased solver synthesis: AutoSAT [103], AutoModSAT [102], and SolSearch [96].

The Case for Automated Hyperspecialization: Evidence from SAT

A

Without verification time

Specialists vs. baselines Core baselines

DPR

Additional baselines

SBVA

Kissat

AE-Kissat

Satsuma

Family VBS

2023 winner

2024 winner

2025 winner

2026 winner

4 core baselines

AutoSAT

AutoModSAT

SolSearch

109 /181 wins

105 /189 wins

105 /189 wins

126 /189 wins

101 /189 wins

137 /189 wins

139 /189 wins

105 /189 wins

110 /181 wins

104 /189 wins

101 /189 wins

125 /189 wins

99 /189 wins

146 /189 wins

143 /189 wins

104 /189 wins

121 /181 wins

110 /189 wins

110 /189 wins

134 /189 wins

109 /189 wins

148 /189 wins

149 /189 wins

112 /189 wins

128 /181 wins

123 /189 wins

123 /189 wins

145 /189 wins

121 /189 wins

160 /189 wins

158 /189 wins

125 /189 wins

10⁶× 10³× 1× 10⁻³×

VeriPB 10⁶× 10³× 1× 10⁻³×

GRAT 10⁶× 10³× 1× 10⁻³×

RQ1

Best 10⁶× 10³× 1× 10⁻³×

Specialization wins

Baseline wins

Share of families won

Families ranked within each panel

Validation PAR-2 ratio: baseline / specialist · Selections by training PAR-2

B

With verification time

Specialists vs. baselines Core baselines

DPR

Additional baselines

SBVA

Kissat

AE-Kissat

Satsuma

Family VBS

2023 winner

2024 winner

2025 winner

2026 winner

4 core baselines

AutoSAT

AutoModSAT

SolSearch

109 /181 wins

101 /189 wins

101 /189 wins

131 /189 wins

96 /189 wins

131 /189 wins

139 /189 wins

98 /189 wins

111 /181 wins

97 /189 wins

99 /189 wins

132 /189 wins

95 /189 wins

131 /189 wins

141 /189 wins

96 /189 wins

122 /181 wins

113 /189 wins

115 /189 wins

144 /189 wins

117 /189 wins

145 /189 wins

155 /189 wins

109 /189 wins

131 /181 wins

120 /189 wins

124 /189 wins

151 /189 wins

121 /189 wins

153 /189 wins

158 /189 wins

117 /189 wins

10⁶× 10³× 1× 10⁻³×

VeriPB 10⁶× 10³× 1× 10⁻³×

GRAT 10⁶× 10³× 1× 10⁻³×

RQ1

Best 10⁶× 10³× 1× 10⁻³×

Specialization wins

Baseline wins

Share of families won

Families ranked within each panel

Validation PAR-2 + verification ratio: baseline / specialist · Selections by training PAR-2

Figure 9. Full head-to-head per-family performance between specialists (rows) and baselines (columns). (A) Validation PAR-2 without verification time. (B) Validation PAR-2 including verification time, using the same training-selected solvers. Each subplot ranks families by the baseline/specialist PAR-2 ratio. Labels denote families where the specialist outperforms the baseline. The two specific subplots marked RQ1 denote the data highlighted in the main body of the paper.

Harrison Green, Claire Le Goues, and Fraser Brown

C

Pareto frontier of on-demand specialization

Developing an optimal runtime policy is inherently a multiobjective optimization problem. We want to both minimize the makespan (𝑀𝜋 ) and maximize the number of solved formulas (𝑄 𝜋 ). Picking an optimal policy therefore requires determining the relationship between these two metrics: Do we care more about solving quickly or solving every formula? In RQ2 (§4.3), we define optimality by a combined metric 𝑃𝜋 where an unsolved formula is equivalent to a penalty of 2𝑇 seconds. This formula is not unreasonable (indeed it mirrors PAR-2 which is widely used to measure SAT solver performance), but it is arbitrary in the sense that we could have decided an unsolved formula was worth more or less. Here, we visualize these on-demand runtime policy results without imposing a combined metric. Instead, for each machine size (𝐶) and workload size (𝑁 ), we plot the Pareto frontier of makespan and coverage tradeoffs. Each subplot in Figure 10, Figure 11, and Figure 12 plots, for a particular (𝐶, 𝑁 ) configuration, the average makespan 𝑀 𝜋 and average coverage 𝑄 𝜋 for each runtime policy 𝜋 computed over all family workloads of size 𝑁 : 𝑟 1 ∑︁ ∑︁ 𝑀 𝜋 := 𝑀𝜋 (W𝑘(𝑖 ) , 𝐶) 𝑟 |𝐾 | 𝑖=1 𝑘 ∈𝐾

𝑟

𝑄 𝜋 :=

1 ∑︁ ∑︁ 𝑄 𝜋 (W𝑘(𝑖 ) , 𝐶) 𝑟 |𝐾 | 𝑖=1 𝑘 ∈𝐾

In each subplot, we draw a line between runtime policies that sit on the Pareto frontier (i.e. no other policy has both a smaller makespan and larger coverage). Additionally, we circle the point with the optimal average PAR-2 score computed as: 𝑃 𝜋 := 𝑀 𝜋 + 2𝑇 (𝑁 − 𝑄 𝜋 ) Note that since results here are averaged over every family, the optimal policy selected here is not necessarily the same as the most frequently optimal policy selected on a per-family basis (as visualized in Figure 6).

The Case for Automated Hyperspecialization: Evidence from SAT

Coverage and makespan by workload (1/3) * Optimal policy: πC, N = arg minπ ∈ Π Pπ (C, N)

Pareto frontier

C=1

C=2

C=4

AE-Kissat '25

AE-Kissat '25

Four-solver race

C=8 Four-solver race

C=16

C=32

C=64

Four-solver race

Four-solver race

Four-solver race

0.6

N=1

0.4 500

500k

AE-Kissat '25

500

200k

AE-Kissat '25

500

100k

AE-Kissat '25

500

50k

Four-solver race

500

20k

Four-solver race

500

20k

Four-solver race

500

20k

Four-solver race

1.2

N=2

0.8 500

500k

AE-Kissat '25

500

200k

AE-Kissat '25

500

100k

AE-Kissat '25

500

50k

Four-solver race

500

20k

Four-solver race

500

20k

Four-solver race

500

20k

Four-solver race

Mean verified solve count, Qπ

2.4

N=4

1.6 1k

500k

AE-Kissat '25

1k

200k

AE-Kissat '25

500

100k

AE-Kissat '25

500

50k

AE-Kissat '25

500

20k

Four-solver race

500

20k

Four-solver race

500

20k

Four-solver race

4.8

N=8

3.2 5k

500k

AE-Kissat '25

2k

200k

AE-Kissat '25

1k

100k

AE-Kissat '25

500

50k

AE-Kissat '25

500

20k

Four-solver race

500

20k

Four-solver race

500

20k

Four-solver race

9.6

N=16

6.4 5k

500k

AE-Kissat '25

5k

200k

AE-Kissat '25

2k

100k

AE-Kissat '25

1k

50k

AE-Kissat '25

500

20k

Four-solver race

500

20k 500

Four-solver race

20k

Four-solver race

19.2

N=32

12.8 10k

500k

5k

200k

5k

100k

2k

50k

1k

20k

500

20k

1k

Mean makespan, Mπ (seconds; log scale) Fixed baseline SBVA '23

Kissat '24

AE-Kissat '25

Race Satsuma '26

Four-solver race

Specialization Wait · DPR

Wait · VeriPB

Wait · GRAT

Hybrid · DPR

Hybrid · VeriPB

Hybrid · GRAT

Figure 10. Pareto frontier of on-demand runtime strategies. N=1 to N=32.

20k

Harrison Green, Claire Le Goues, and Fraser Brown

Coverage and makespan by workload (2/3) * Optimal policy: πC, N = arg minπ ∈ Π Pπ (C, N)

Pareto frontier

C=1

C=2

C=4

C=8

C=16

C=32

C=64

AE-Kissat '25

AE-Kissat '25

AE-Kissat '25

AE-Kissat '25

Four-solver race

Four-solver race

Four-solver race

38.4

N=64

25.6 20k

500k 10k

AE-Kissat '25

200k

AE-Kissat '25

5k

100k

AE-Kissat '25

5k

50k

AE-Kissat '25

2k

20k

Four-solver race

1k

20k

Four-solver race

1k

20k

Four-solver race

76.8

N=128

51.2 50k

200k

50k

AE-Kissat '25

200k

AE-Kissat '25

20k

100k

AE-Kissat '25

10k

50k

AE-Kissat '25

5k

20k

Four-solver race

2k

20k

Four-solver race

1k

10k

Four-solver race

Mean verified solve count, Qπ

154

N=256

102 100k

500k 50k

AE-Kissat '25

200k

Hybrid · VeriPB

50k

100k

Hybrid · VeriPB

20k

50k

Wait · GRAT

10k

20k

5k

Wait · GRAT

20k

Four-solver race

2k

10k

Four-solver race

307

N=512

205 200k

500k

AE-Kissat '25

100k

200k

Hybrid · VeriPB

50k

200k

Hybrid · VeriPB

50k

100k

Wait · GRAT

20k

50k

Wait · GRAT

10k

20k

Wait · GRAT

5k

10k

Four-solver race

600

N=1k

400 420k

530k

Wait · GRAT

170k

240k

Hybrid · VeriPB

100k

200k

Wait · GRAT

50k

200k

Wait · GRAT

20k

100k

Wait · GRAT

20k

50k

Wait · GRAT

10k

20k

Four-solver race

1.2k

N=2k

800 720k 840k

270k

370k

200k

500k

100k

200k

50k

200k

20k

100k

10k

Mean makespan, Mπ (seconds; log scale) Fixed baseline SBVA '23

Kissat '24

AE-Kissat '25

Race Satsuma '26

Four-solver race

Specialization Wait · DPR

Wait · VeriPB

Wait · GRAT

Hybrid · DPR

Hybrid · VeriPB

Hybrid · GRAT

Figure 11. Pareto frontier of on-demand runtime strategies. N=64 to N=2k.

50k

The Case for Automated Hyperspecialization: Evidence from SAT

Coverage and makespan by workload (3/3) * Optimal policy: πC, N = arg minπ ∈ Π Pπ (C, N)

Pareto frontier

C=1 Wait · GRAT

C=2

C=4

Wait · GRAT

C=8

Wait · GRAT

C=16

Wait · GRAT

C=32

Wait · GRAT

C=64

Wait · GRAT

Wait · GRAT

3k

N=5k

2k 1.5M

1.9M

Wait · GRAT

500k

1M

Wait · GRAT

500k

2M

Wait · GRAT

200k

1M

Wait · GRAT

100k

500k

Wait · GRAT

50k

200k

Wait · GRAT

20k

100k

Wait · GRAT

6k

N=10k

4k 2.7M 3.6M Wait · GRAT

1M

2M

Wait · GRAT

500k

2M

Wait · GRAT

500k

2M

Wait · GRAT

200k

1M

Wait · GRAT

100k

500k

Wait · GRAT

50k

200k

Wait · GRAT

Mean verified solve count, Qπ

12k

N=20k

8k 5M

7.1M

Wait · GRAT

2.4M 3.4M

1M

Wait · GRAT

5M

Wait · GRAT

500k

2M

Wait · GRAT

500k

2M

Wait · GRAT

200k

1M

Wait · GRAT

100k

500k

Wait · GRAT

24k

N=40k

16k 9.8M

14M

Wait · GRAT

4.7M

6.9M

2M

Wait · GRAT

10M

Wait · GRAT

1M

5M

Wait · GRAT

500k

2M

Wait · GRAT

500k

2M

Wait · GRAT

200k

1M

Wait · GRAT

48k

N=80k

32k 19M

28M

Wait · GRAT

9.4M

14M

5M

Wait · GRAT

20M

Wait · GRAT

2M

10M

Wait · GRAT

1M

5M

Wait · GRAT

500k

2M

200k

Wait · GRAT

2M

Wait · GRAT

600k

N=1M

400k 240M 340M

100M

200M

50M

200M

20M

200M

10M

100M

5M

50M

5M

Mean makespan, Mπ (seconds; log scale) Fixed baseline SBVA '23

Kissat '24

AE-Kissat '25

Race Satsuma '26

Four-solver race

Specialization Wait · DPR

Wait · VeriPB

Wait · GRAT

Hybrid · DPR

Hybrid · VeriPB

Hybrid · GRAT

Figure 12. Pareto frontier of on-demand runtime strategies. N=5k to N=1M.

20M

Harrison Green, Claire Le Goues, and Fraser Brown Solving strategies

Only this strategy

Systematic

Local / stochastic

Algebraic / structural

Direct / construction

Inference / relaxation

Other / opaque

Logical

60.7% 57.5% 22.0% 21.0% 17.5% 16.9% 15.7% 15.2% 14.3% 13.6% 8.6% 8.6% 8.5% 7.8% 7.1% 7.1% 6.3% 6.0% 4.4% 4.4% 3.9% 3.7% 3.5% 3.4% 2.1% 2.1% 2.1% 1.9% 1.9% 1.8% 1.8% 1.8% 1.8% 1.2% 1.2% 1.1% 0.9% 0.7% 0.5% 0.4%

Structural construction / theorem CDCL Greedy construction Exhaustive enumeration Native-domain local search Constraint backtracking Circuit / equivalence reasoning Cardinality / PB reasoning WalkSAT-style search Other graph algorithms DPLL (no clause learning) Weighted local search Arithmetic / number theory Propagation-only solving ProbSAT-style search Tabu search / TabuSAT XOR / GF(2) elimination State-space search Matching / flow Resolution / elimination Meet-in-the-middle search Dynamic programming Simulated annealing Stored witness / proof / answer Branch and bound Large-neighborhood search Other algorithm Look-ahead SAT search Belief / survey propagation Input-feature answer rule GSAT-style greedy flips Independent random sampling 2-SAT / implication solving Horn / tractable fragments Novelty-style search Local search (unspecified) Population / evolutionary Decision diagrams / compilation Continuous optimization Configuration-checking search 0

25

50

Structure recovered from CNF

75

100

Solvers (%)

Strategy labels per solver 2+ labels (460)

81.1%

1 label (106)

18.7%

Combinatorial

43.2% 41.3% 38.6% 31.4% 27.9% 24.0% 22.0% 21.9% 21.2% 20.6% 20.5% 14.5% 9.7% 8.5% 4.9% 0

25

50

75

100

Solvers (%)

Figure 14. What representations do specialists recover? Preprocessing and inprocessing Before only

Before + during

During only

62.4% 33.5% 17.8% 15.7% 14.1% 13.8% 13.6% 13.1% 9.2% 8.6% 8.3% 7.8% 7.2% 7.1% 6.9% 5.8% 4.2% 2.8% 1.1% 0.5%

Root unit simplification Clause cleanup Cone / core reduction Structural strengthening Other preprocessing Resolution elimination Self-subsuming resolution Gate / circuit reduction Domain filtering Failed-literal probing Cardinality / PB reduction XOR / algebraic reduction Literal substitution Problem decomposition Variable compaction Symmetry breaking Pure-literal elimination Clause subsumption Clause vivification Blocked-clause elimination 0

25

50

75

100

Solvers (%)

0.2%

0 labels (1) 0

25

50

75

100

Solvers (%)

Figure 13. What algorithms do specialists use?

Figure 15. What kinds of preprocessing and inprocessing do specialists do? Encoding assumptions and safeguards Encoding assumptions

D

Domain

Semantic components / cones Problem graphs Gates / Boolean circuits Finite-domain variables Cardinality / PB constraints Permutations / orders Arithmetic / modular structure Other recovered structure Equivalence / implication Matching / assignment XOR / parity equations Symmetry / group structure State-transition systems Geometry / packing Statistical models

Extended: Methods and techniques

Recognition / scope

Answer safeguards

97.2% 96.3% 94.4% 90.8% 86.4% 65.4% 64.4% 61.6% 57.7% 47.6% 45.9% 44.6% 12.7% 6.2% 3.7%

Unsupported-case rejection Structural validity guards Relational recognition Exact templates / signatures Original-CNF model checks Multiple encoding variants Variable-ID / layout reliance Model reconstruction Canonicalized recognition Rejected-candidate fallback Domain / reduced checks Clause-order reliance Sampled / fingerprint checks Stored candidates / proofs Certificate / RUP checks 0

25

50

75

100

Solvers (%)

Figure 16. How do specialists decide whether to run or not?

The Case for Automated Hyperspecialization: Evidence from SAT Search tuning Branch / phase

Systems-level optimizations

Learning / restarts

Local search

Search control

Input / output

71.3% 69.3% 66.3% 61.2% 59.1% 56.4% 51.3% 46.4% 45.9% 44.4% 36.5% 31.2% 30.2% 27.0% 23.3% 19.4% 15.7% 9.9% 8.6% 8.3% 6.2%

Static / occurrence ordering Scheduled restarts Input-dependent parameters Bounded search effort Biased / seeded phases Conflict-activity branching Phase saving Domain-aware branching Learned-clause retention Diverse search starts Learned-clause minimization Random branch / tie choices Make / break scoring Noisy move selection LBD / glue quality Look-ahead branch scoring Stagnation control Tabu / recency rules Adaptive constraint weights Temperature / noise control Adaptive restarts 0

25

50

75

100

Information transfer

93.3% 60.1% 39.2% 38.8% 38.1% 37.9% 34.2% 33.0% 32.8% 26.1% 20.6% 10.2% 4.9%

Structural dispatch Sequential methods Candidate / phase transfer Constraint / bound transfer Separate SAT / UNSAT routes Size / feature dispatch General SAT fallback Repeated configurations Budgeted handoff Residual subsolver calls Unlogged / certifying stages Refinement loop Adaptive effort allocation 25

50

25

50

75

100

Figure 19. What kinds of systems-level optimizations do specialists implement?

Combining solving methods

0

99.5% 96.5% 79.4% 75.5% 68.6% 66.0% 59.4% 51.5% 45.1% 39.2% 38.6% 32.8% 31.9% 25.7% 10.1% 6.9% 3.7% 1.8% 1.8% 1.1% 0.2% Solvers (%)

Figure 17. How do specialists tune their search?

Composition

Memory / state

0

Solvers (%)

Selection / fallback

CPU

Compiler / ISA flags Preallocation / reuse Incremental state updates Flat / compact storage Watched / blocking literals Specialized parsing Scalar word parallelism Lazy / deferred work Buffered proof output Bit-count / scan intrinsics Reduced proof-output work Specialized hot loops Lookup tables Memory-mapped input Branchless updates OS memory hints Compact proof encoding Alignment / cache blocking Explicit SIMD Branch-prediction hints CPU software prefetch

75

100

Solvers (%)

Figure 18. How do specialists combine multiple methods?

Harrison Green, Claire Le Goues, and Fraser Brown

E

Extended: End-to-end deployment

In this section we present full details of the end-to-end deployment case studies. We study circuit equivalence checking on OpenTitan [85] and ORFS [105] designs, C verification with CBMC [22], Rust verification with Kani [27], and dependency resolution with Conda [93]. For each case, we train a specialist on extracted formulas and select its checkpoint by validation PAR-2. We then evaluate the generated specialist on a new workload and compare it to native performance. Specialists were synthesized and evaluated internally in the same mixed infrastructure described previously (§4.1.3). All of the redeployment evaluations were measured by running systems in Docker with one assigned CPU on an AMD EPYC 9454P with networking disabled. We include scripts necessary to reproduce these results in our artifact package. Native runs use one assigned CPU, three repetitions per mode, and 32 GiB of memory, except that Conda uses one repetition and Kani runs (being more memory-intensive) use 16 GiB. Shared compilation and circuit preparation are excluded from timing; encoding, solver interaction, and requested checks are included. Note also that in the five case studies we test, native tools (by default) do not perform independent SAT certificate checking. We evaluate our specialist both with and without verification. E.1

finite-field multiplier, a 64-bit substitution and permutation block, and 32-input sum and maximum trees, each under four rewrites (Table 2). Native is ABC’s cec procedure, which combines circuit reasoning with internal SAT solving. Our path builds the comparison CNF from the same prepared pair, runs the specialist, and optionally checks its answer. This measures the equivalence stage, not the whole OpenTitan verification flow. The limit is 600 seconds per stage, with a 630-second outer guard for native CEC. Ratios are 1.64/1.15 on all 292 original pairs and 3.47/2.17 on eight of the 16 new pairs. The specialist returns unsupported on the eight sumtree and maximum-tree cases; native finishes all 16. Thus the new-workload gain applies only to the supported half of the set. Table 2. OpenTitan workload provenance. Counts describe the full input sets. Input

Source

Formulas

Cryptographic, Ibex, and error-correction blocks Same retained circuit pairs Multiplier, substitution/permutation, sum and maximum trees

Original New

Size 292 CNFs 292 pairs 16 pairs

Case studies

Æ The rest of this section was generated with GPT-6 Astra (xhigh) and provides purely factual details about experimental setup based on the code. It has been completely vetted for correctness by the human authors. E.1.1 OpenTitan circuit equivalence. System and export. OpenTitan is an open-source hardware security project. Here we check whether two versions of one of its circuits produce the same outputs (i.e. combinational equivalence checking). Our sources are Ascon, Keccak, PRESENT, and PRINCE cryptographic blocks, the vendored Ibex arithmetic unit and instruction decoder, and error-correction encoders and decoders. We use Yosys/Slang to convert their hardware descriptions into logic graphs and then we apply four ABC rewrite recipes and compare each result with its reference graph. ABC exports a CNF that asks whether any input makes the two graphs disagree. These are explicit comparison formulas, rather than copies of ABC’s internal SAT queries. We also include deliberate legal configuration mismatches. The 296 planned comparisons yield 292 distinct formulas: 272 UNSAT and 20 SAT, split into 233 training and 59 validation formulas. The specialist’s training and validation score ratios are 14.28/11.21 and 13.27/10.38, respectively. Deployment. The original workload contains the 292 retained circuit pairs. The new workload contains a 32-bit

E.1.2 ORFS circuit equivalence. System and export. OpenROAD-flow-scripts (ORFS) provides scripts and example designs for hardware implementation. We use its designs for a second circuit-equivalence study. We apply eight synthesis-mapping flows to nine circuit tops from AES, JPEG, and Ethernet designs. Yosys produces reference and revised logic graphs; AIGER’s aigmiter and aigtocnf tools construct and export the comparison formulas. Registers are treated as corresponding inputs and outputs, so these are combinational comparisons rather than proofs over arbitrarily many clock cycles. All 72 distinct formulas are UNSAT, with 57 used for training and 15 for validation. The training and validation score ratios are 50.64/28.64 and 72.12/40.42. Deployment. We use the 72 original circuit pairs and a new set comprising GCD, FIFO, RISC-V arithmetic-unit, and UART designs under the same eight mapping flows (Table 3). Native is the Yosys-bundled ABC cec procedure. Our path rebuilds the comparison and CNF, then solves and optionally checks the answer. Shared synthesis is excluded; the timed task is equivalence checking, not placement, routing, or the full ORFS flow. Limits match the OpenTitan experiment. All 72 original and 32 new cases finish in both specialist modes. Native-time ratios are 2.16/0.61 on the original set and 2.81/1.21 on the new set. Checking removes the originalworkload gain, but the new-workload gain remains.

The Case for Automated Hyperspecialization: Evidence from SAT

Table 3. ORFS workload provenance. Each source top receives eight mappings. Input

Source

Formulas Original New

Nine AES, JPEG, and Ethernet tops Same mapped circuit pairs GCD, FIFO, RISC-V ALU, and UART

Size 72 CNFs 72 pairs 32 pairs

E.1.3 C verification with CBMC. System and export. CBMC checks C code by turning bounded program executions and safety properties into logical constraints. We use AWS C Common’s verification harnesses, which are small drivers that supply inputs and state the properties to check. They cover utilities such as buffers, arrays, hash tables, and priority queues. We build all 173 harnesses with CBMC 6.4.0 and preserve their upstream options. A wrapper copies each live CNF at CBMC’s external SAT interface before Kissat answers it. A separate export supplies variable names; it does not replace the live query. Six harnesses make no SAT call. The remaining 167 yield distinct UNSAT formulas, split into 133 training and 34 validation formulas. The training and validation score ratios are 0.41/0.79 and 1.53/1.93. Deployment. The original workload contains the 167 formulaproducing AWS proofs, and the new workload contains all 15 coreJSON proofs (Table 4). We time fresh CBMC runs from prepared program representations, including symbolic execution and Boolean encoding. The specialist replaces the external SAT callback; checked UNSAT answers must pass DPR-trim and CakeLPR. Native AWS runs use CBMC 6.4.0 with MiniSat. Native coreJSON runs follow upstream CI: CBMC 6.3.1, its required contract instrumentation and patch, and per-proof MiniSat, CaDiCaL, or Kissat choices. Original runs have 3600 seconds; new runs have 300 seconds. On the 159 original targets shared by both specialist modes, the ratio is 1.18/0.32. Eight checked targets exhaust CakeLPR’s separate 4-GiB heap. The unchecked specialist finishes all 167, but its ratio on that full set is 0.69, so the shared-subset point does not establish an overall gain. On new targets, native finishes 15, the unchecked specialist 14, and the checked specialist 13. Ratios on the 13 shared targets are 0.38/0.11; the other specialist runs time out. The validation-formula improvement therefore does not carry through to faster native verification.

E.1.4 Rust verification with Kani. System and export. Kani checks Rust properties by compiling verification harnesses into CBMC’s program representation. One harness can issue several SAT queries. We use Kani 0.67.0 on all 112 harnesses in the s2n-quic CI matrix, covering its codec, core protocol, and platform crates. A copying wrapper around external Kissat records each live per-property formula. This produces 302 distinct formulas, split into 241 training and 61 validation formulas. Of these, 237 are SAT and 65 are UNSAT; a SAT query does not by itself mean the overall harness found a bug. The external interface can encode queries differently from native incremental solving. Training and validation score ratios are 1.43/1.74 and 10.54/6.48. Deployment. We rerun the 112 s2n-quic harnesses and test 14 new harnesses from Hifitime, a Rust time library (Table 5). Native is Kani’s default CaDiCaL workflow. Our path uses the external specialist with optional model or GRAT checks. Each timed invocation checks one harness after its own untimed code-generation warmup. Hifitime uses its required contract and stubbing options and a library buildconfiguration adjustment; harness bodies are unchanged. Original runs have 3600 seconds and 32 GiB; new runs have 300 seconds and 16 GiB. Ratios are 0.86/0.62 on 109 shared original targets; two specialist targets time out and another fails during checking. Seven Hifitime harnesses make no SAT call and are excluded from solver comparisons. Among the seven that do, native finishes all seven, while the specialist finishes four without checking and three with checking. Three inputs are unsupported, and the checked epoch-equality harness times out. Ratios on the three shared targets are 0.76/0.59. Native remains faster despite the isolated-formula gains. Matching source harnesses can also produce different later SAT queries after changing the solver. Table 5. Kani workload provenance. Seven new harnesses invoke SAT. Input

Source

Formulas

s2n-quic codec, core, and platform crates Complete s2n-quic CI matrix Hifitime timescale, general, and epoch harnesses

Original New

Size 302 CNFs 112 harnesses 14 harnesses

Table 4. CBMC workload provenance. Six AWS harnesses produce no SAT query. Input

Source

Formulas Original New

AWS C Common verification harnesses Same formula-producing AWS proofs All upstream coreJSON proof targets

Size 167 CNFs 167 proofs 15 proofs

E.1.5 Dependency resolution with Conda. System and export. Conda chooses compatible package versions for a software environment. Its classic backend repeatedly calls pycosat while constructing and improving a package plan. We run Conda 26.5.3 with pycosat on the EO-datascience and xESMF environment specifications, using a frozen condaforge index. We copy the current clauses at every classic

Harrison Green, Claire Le Goues, and Fraser Brown

SAT callback, then let the unchanged pycosat path answer. The 154 calls yield 151 distinct formulas: 105 SAT and 46 UNSAT, split into 120 training and 31 validation formulas. These include intermediate optimization queries. Training and validation score ratios are 10.39/3.91 and 11.76/4.83. Deployment. The original workload reuses those two environments. The new workload uses three XROMS CI specifications for Python 3.11–3.13 and Xskillscore’s minimumtests specification (Table 6). We time offline environment planning with conda env create –dry-run, excluding installation and copying the frozen index. Native baselines are classic/pycosat and libmamba 2.5.0, a different dependency resolver built on libsolv. The specialist replaces every SAT callback inside classic Conda; it is not inserted inside libmamba. Checked answers use model checks or VeriPB/CakePB. Original plans have 3600 seconds and new plans 600 seconds. Against pycosat, ratios are 1.67/0.84 on both original plans and 1.20/0.89 on one shared new plan. Against libmamba, ratios are 0.054/0.027 on both original plans and 0.053/0.040 on two shared new plans. On the four new plans, pycosat finishes three, libmamba four, and the specialist three without checking and two with checking. One checked plan takes 599.38 seconds, close to the limit, with only one repetition. The specialist helps classic Conda without checking on the shared cases, but libmamba remains much faster. Different SAT assignments and the omission of pycosat’s propagationlimit shortcut can change later queries and final package choices. New environments also share a small initialization formula with the original corpus. Table 6. Conda workload provenance. Counts refer to planning, not installation. Input

Source

Formulas Original New

EO-datascience and xESMF environments Same environment specifications Three XROMS specifications and Xskillscore minimum-tests

Size 151 CNFs 2 plans 4 plans

The Case for Automated Hyperspecialization: Evidence from SAT

F

Extended: Cross-family sensitivity and performance

In Figure 20 we plot cross-family performance of every top training-selected specialist on every other family (capped at 100 instances). Most specialists implement some sort of early detection to decide whether they are capable of running on a given formula. A family is not attempted only if the specialist returns UNKNOWN within one second on every sampled instance. Otherwise, it is attempted; the matrix distinguishes attempted families with no verified solves from those with at least one verified solve. Incomplete evaluations are excluded from the coverage counts.

Harrison Green, Claire Le Goues, and Fraser Brown

Cross-family solver performance Relative performance

Matrix cells

PAR-2 / best for target 1×

10²×

10⁴×

≥10⁶×

Darker squares indicate lower PAR-2

No attempt

Solved; not best

Attempted; no solve

Best (all ties)

Transfer coverage ≥1 solve

0 solves

Rejected result

1

25

50

Source solver (rank)

75

100

125

150

175

189 1

25

50

75

100

Target family (rank)

125

150

175

189

0

100

188

Families attempted Excludes own family

Figure 20. Cross-family performance of each top solver on up to 100 validation instances from every other family. Each cell shows the performance of a given specialist solver (y-axis) on a target family. The diagonal shows solver performance on its own target family. The right panel shows the solver’s transfer coverage: which other families it attempted and had at least one solve.

The Case for Automated Hyperspecialization: Evidence from SAT

G

Choice of agent

We conducted a small pilot evaluation (in May 2026) to decide which base model and harness to use for our experiments. We found that both GPT-5.5 (in Codex) and Claude Opus 4.8 (in Claude Code) were able to successfully and consistently synthesize working specialized solvers, while smaller proprietary models and frontier open-weight models took longer and failed more frequently. After GPT-5.6 Sol was released (which has the same pricing as GPT-5.5), we switched evaluations to use that model. Due to the costs associated with running so many agents at scale, we limit our evaluation in this paper to a single representative frontier model. The point of the paper is to demonstrate an existence proof of the utility of hyperspecialization with frontier models available today. We expect that future models will continue to improve at both the ability to write performant solver code and in cost-efficiency, both axes would continue to make hyperspecialization more economically viable. We leave exploration of comprehensive multi-agent benchmarking to future work.

H

Extended: Cardinality constraint encodings

Here we present complete details of our cardinality constraint encoding case study. H.1

Benchmark problems

We evaluated hyperspecialization on three crafted benchmarks covering both SAT and UNSAT cases and a variety of formula size ranges. H.1.1 GPHP. Generalized, capacitated pigeonhole principle problems ask if it is possible to assign pigeons to holes. Unlike the standard framing, holes in the generalized case can accommodate multiple pigeons up to a fixed capacity. A Boolean variable 𝑥 𝑝,ℎ represents assigning pigeon 𝑝 to an allowed hole ℎ. The formula defines the following constraints: Í • A pigeon must occupy exactly one hole: ℎ 𝑥 𝑝,ℎ = 1 Í • Each hole has at-most-𝑐ℎ capacity: 𝑝 𝑥 𝑝,ℎ ≤ 𝑐ℎ We sample instances with between 17 and 25 pigeons, 6–9 holes, and capacities of 2–4. All instances are UNSAT by construction: there is always one more pigeon than available capacity. H.1.2 LDPC decoding. Low-density parity-check [35] decoding problems ask whether an error pattern can explain an observed parity-check syndrome. A Boolean variable 𝑒𝑖 indicates whether bit 𝑖 was flipped during transmission. Given a binary parity-check matrix 𝐻 and syndrome 𝑠, the formula defines the following constraints: • The error pattern must satisfy the observed parity checks: 𝐻𝑒 = 𝑠 (mod 2).

Encoding

Year

Naive (direct) Totalizer [8] Sequential counter [97] Sorting network [10, 32] Cardinality network [6, 7] Modulo totalizer [82] 𝑘-modulo totalizer [77]

— 2003 2005 2006 2009 2013 2015

Table 7. Cardinality constraint encoding formats tested.

• The Í error pattern must contain exactly 𝑘 flipped bits: 𝑖 𝑒𝑖 = 𝑘. We sample instances with 105 error bits, 52 parity checks, and 𝑘 = 12. Each parity check involves 13 bits. The cardinality encoding applies to the single exact-weight constraint; the CNF encoding of the parity checks is fixed across encoding methods. All instances are SAT by construction: we plant an error pattern of weight 12 and compute its syndrome. H.1.3 Vertex cover. Vertex cover problems ask whether a graph admits a set of at most 𝑘 vertices containing at least one endpoint of every edge. A Boolean variable 𝑥 𝑣 represents selecting vertex 𝑣 for the cover. The formula defines the following constraints: • Every edge (𝑢, 𝑣) must have at least one selected endpoint: 𝑥𝑢 ∨ 𝑥 𝑣 . Í • The cover must contain at most 𝑘 vertices: 𝑣 𝑥 𝑣 ≤ 𝑘. We sample instances with 2,100 vertices, 3,570 edges, and 𝑘 = 1, 330. The cardinality encoding applies to the single global cover-size constraint; each edge contributes one binary clause. We construct the graphs by reducing degreethree Tseitin parity formulas to 3-CNF and then to vertex cover. Even- and odd-charge source instances yield 50 SAT and 50 UNSAT instances, respectively. H.2

Benchmark construction

For each benchmark, we sampled 100 formulas, calibrated to be hard for baseline solvers, and then encoded these 100 formulas separately with a variety of cardinality encodings implemented in PySAT [49, 50], listed in Table 7. We split each group of 100 formulas into 80 training instances and 20 validation instances (using the same seed) and independently synthesize specialists for each benchmark/encoding pair (using only GRAT-based specialists for parity with baselines). H.3

Results

In Figure 8, we present average validation PAR-2 values for each benchmark/encoding pair, for both baseline solvers and generated specialists. Here, we examine how these results

Harrison Green, Claire Le Goues, and Fraser Brown

relate to formula size. Figure 21 shows baseline performance, while Figure 22 shows specialist performance. The x-axis of each subplot is the mean number of variables (top row) or clauses (bottom row). Encoding choices substantially affect formula sizes. For the baseline solvers (Figure 21), we see some correlation with performance. Solvers generally perform worse as the number of clauses or variables grows for both GPHP and vertex cover. Indeed this is not surprising. It has been observed before that encoding size plays a role in solver performance, although it is not the sole factor [1]. Interestingly, we observe no such clear trend for the specialists (Figure 22). In some cases, the trend is almost reversed. The sequential encoding for GPHP performs the best among all of the baseline solvers but is one of the worst for the specialist. Conversely, the largest and simplest encoding (naive encoding) is one of the worst for the baselines but yields the best specialist. For LDPC decoding, the 𝑘-modulo totalizer is decisively worse than all other formats for the specialists despite being the newest of the encodings and producing the smallest formulas in both variables and clauses. Yet for vertex cover, it flips and yields the best specialist. One potential hypothesis is that there are two competing factors that influence specialist performance: 1. Formulas should be described succinctly in an easily parseable way (fancier encodings create more complex structures and auxiliary variables which may be harder or slower to parse consistently). 2. At a certain point, raw formula size does become important. For example, vertex cover problems are roughly two orders of magnitude bigger (in bytes) than both the other benchmarks. Here, encodings that save parsing time (even if more complex) may be worthwhile. Subsequent exploration is out of scope for this paper, but interesting future work could be to develop an understanding of optimal encodings for hyperspecialization (which may be separate from encodings for general-purpose solvers).

I

Extended: Ensuring and evaluating task alignment

Here we describe in detail the objectives of restricted internet access (§I.1), how we iteratively improved the agent prompt and harness before the full evaluation (§I.2) and how we analyzed runs after the evaluation to ensure compliance (§I.3). I.1

Agent internet access

During initial pilot experiments, we experimented with allowing agents to access the internet, with the intuition that such a resource would be useful to learn about specialized techniques and (potentially) research information about the specific high-level problem if applicable. While these agents did perform well, we found it near impossible to prevent the

Formula size and baseline performance Validation PAR-2 (s)

SBVA '23

GPHP Naive

Kissat '24

LDPC decoding KM-Tot

Seq.

Tot.

M-Tot

Sort.

KM-Tot

Card.

Tot.

Card.

M-Tot

AE-Kissat '25

Satsuma '26

Vertex cover

Seq.

KM-Tot Sort.

Tot. M-Tot

Seq. Card.

1.2k 600 400 100

1k

2k

Mean variables Seq.

M-Tot

KM-Tot

Tot.

Card.

6k

10k

1m

Mean variables Naive

KM-Tot

Seq.

Sort.

M-Tot

10k

6k

Mean variables

Sort. Card.

KM-Tot Tot.

Card. M-Tot

Tot. Seq.

1.2k 600 400 1k

Mean clauses

15k

100k

Mean clauses

2m Mean clauses

Figure 21. Baseline validation PAR-2 versus mean wholeCNF variable count (top row) and clause count (bottom row), across cardinality encodings. Columns show benchmark groups. Both axes are logarithmic, with a shared time range. Formula size and specialist performance Validation PAR-2 (s) GPHP

LDPC decoding

Tot.

KM-Tot

Seq.

1.2k

Card.

30

KM-Tot

1

Sort.

Naive

0.03 100

KM-Tot

0.03 M-Tot 1k

Card. Seq.

Naive

10k

Mean clauses

10k

0.003 0.001

1m

Mean variables M-Tot

3

Tot.

0.01 Sort.

KM-Tot Card.

0.03

6k

KM-Tot

Seq.

1

M-Tot

Seq.

0.3

Mean variables

Card. Tot.

30

Sort.

2k

Mean variables

1.2k

Tot.

0.003

1k

M-Tot Tot.

0.01

0.001

M-Tot

Vertex cover

3

M-Tot Card.

Tot.

Sort.

Card. KM-Tot

0.03

Seq.

6k

0.3

15k

Mean clauses

100k

Seq.

2m

Mean clauses

Figure 22. Validation PAR-2 of training-selected specialists versus mean whole-CNF variable count (top row) and clause count (bottom row). Filled squares indicate that all 20 validation instances were solved; open squares indicate at least one failure. Both axes are logarithmic; time ranges differ across benchmark groups.

agent from inadvertently leaking information about the withheld validation set (since the Global Benchmark Database is a public resource), and thus contaminating the evaluation. Therefore, we chose to run all agents with restricted internet access. True network isolation would be bulletproof, but would also prevent agents from installing runtime packages or other dependencies (which we did want to allow). Therefore, we attempted to enforce the restricted internet access by both disabling the web_search tools in Codex and by explicitly

The Case for Automated Hyperspecialization: Evidence from SAT

prompting the agent not to use the internet. We validated that these measures were effective after the fact (§I.3) and reran any runs which violated these rules (in practice, only 7). In retrospect, however, none of the agents attempted to install any runtime packages or other dependencies, thus this restriction was unnecessary. A more robust solution would be to properly enforce full network isolation at the sandbox level.

Integrity audit: round 1

I.2

Iterative harness and environment development

Our initial agent configuration was intentionally minimal. We started with a simple prompt and Docker environment. We iteratively ran pilot experiments and used Docent [76] to analyze agent behaviors. Docent is a tool for using language models to scan, summarize, and cluster findings and proved useful for quickly identifying failure modes and avenues for cheating. We found it very useful for quickly iterating on framework design and prompt engineering. As an example, our initial prompt resulted in agents more frequently attempting to build general-purpose solvers, so we added language to enforce the task of hyperspecialization. We also found that agents would frequently attempt to use certain common command-line tools (jq,ripgrep, time, xz, gdb, etc...) which were not available in our restricted environment, so we updated the environment to include all of these tools by default. Initially, our prompt language describing the timeout (that we would run solvers for 10 minutes) was not clear enough, causing many agents to mistakenly implement their own timeout logic inside their generated solvers, needlessly terminating them just before 10 minutes if no solution had been found. We were able to use Docent to detect and fix this issue before the full evaluation. Docent also helped us discover several environment configuration failures, such as memory-induced crashes in part of the evaluator that pre-checked SAT models (replaced with gratchk), and file permission issues when mounting the evaluation dataset in Docker. I.3

Cheating analysis pipeline

After the full evaluation, we performed a comprehensive analysis of generated artifacts (transcripts and solver code) to ensure compliance with the intended task. In particular, we wanted to ensure that no agent accessed the internet (thus could not leak information about the withheld validation set) or bundled external solvers as part of the submitted code. While manually reviewing all transcripts and solver code is infeasible at scale, we deployed a comprehensive two-part analysis. A: LLM-based trace analysis. We uploaded transcripts to Docent and provided GPT-5.6 Luna (high) a detailed prompt, tasking it with identifying any evidence of network

Cohort

567 runs

Screens

No lead 339 Any lead 228

Trace (Luna)

Clear 563 Candidate 4

Code (Luna)

Clear 563 Candidate 4

Cumulative

Clear 560 Confirmed 7

Final

Clear 560 Confirmed 7 Initial screen combinations

None (339)

Code (3)

Regex (221)

Regex + trace (3)

All three (1)

Figure 23. First round of the integrity audit. access. We supplemented this analysis with a deterministic Python script that searched over the agent-authored commands for network-related keywords, including curl/wget, remote Git operations, HTTP libraries, and package manager commands. Matching excerpts were provided as evidence leads to GPT-5.6 Luna (high) during the transcript analysis on Docent. B: LLM-based code analysis. We provided GPT-5.6 Luna (high) with the full solver source, including build and run scripts for each run and tasked it with identifying if there was evidence of bundling an external solver. Cumulative decision. For runs flagged by either the transcript or solver analysis, we gathered the supporting evidence and prompted GPT-5.6 Sol (high) to determine whether a legitimate violation occurred. All runs labeled as cheating were manually reviewed and we also spot-checked several non-flagged runs. I.3.1 Results. We performed two rounds of analysis. In the first round (Figure 23), we analyzed the 567 initial evaluation runs (189 families × 3 proof formats). Of these, four candidates were flagged in the trace analysis and another four (with one overlap) in the code analysis. All of these were marked as confirmed by the aggregator (and human review). After rerunning these seven invalidated runs, we performed a second audit on the new seven runs. None of these runs were flagged in either the trace or code analysis. Only the non-cheating runs were used for data results in this paper.

Harrison Green, Claire Le Goues, and Fraser Brown

Integrity audit: round 2 Cohort

7 runs

Screens

No lead 4 Any lead 3

Trace (Luna)

Clear 7

Code (Luna)

Clear 7

Cumulative

Clear 7

Final

Clear 7

Initial screen combinations None (4)

Regex (3)

Figure 24. Second round of the integrity audit.

The Case for Automated Hyperspecialization: Evidence from SAT

J

Complete specialist gallery

2d-strip-packing

This gallery contains a short description for each of the 189 families we evaluate on and the three train-selected specialists for each. Æ Note that the content generated in this section was generated by tasking GPT-5.6 Luna (high) with summarizing both families and solvers. We found no inaccuracies when spot-checking but it is possible that some descriptions are slightly inaccurate (although unlikely to be egregiously so). Reading the numbers. No verif. is baseline PAR-2 divided by specialist PAR-2; + verif. adds all recorded verification time to both scores before taking the ratio. Values above 1× favor the specialist (green); values below 1× favor the baseline (rust); ratios rounding to 1.00× are gray. Bold marks the largest unrounded ratio among the three specialists in that family, separately for each timing column (including ties). Solved counts verified validation results; failures retain the PAR-2 penalty, so speedups can also reflect differences in coverage.

01-integer-programming

001

6 GBD instances · 2 validation · baseline 2/2

Find a vector of zero-one values that satisfies a system of integer linear equations. The benchmark encodes the equations as CNF using Boolean circuits, with auxiliary variables representing intermediate computation. Proof format

GRAT

No verif.

+ verif.

Solved

137 ×

76.8 ×

2/2

Circuit reconstruction extracts XOR rows and recovers the binary linear system from a prescribed prefix and bounded-input sequential layout. Meet-in-the-middle handles smaller cases, while wider bounded searches use projection and local heuristics before CDCL fallback. Checked complete assignments yield SAT; unsupported or exhausted searches return UNKNOWN, with no UNSAT certificate path. DPR

0.664 ×

0.694 ×

2/2

Weighted union-find recovers integer rows from a prescribed equivalence-linked circuit layout; exact elimination and bounded nullspace search then seek Boolean models. Bit-parallel and lattice-guided searches supplement it, with CDCL after heuristic failure. Fallback learning logs witness-free DPR additions; direct SAT emits no proof output, and unsupported or internal failures return UNKNOWN, not UNSAT. VeriPB

106 ×

10.9 ×

2/2

Ordered Boolean-function reconstruction and parity-chain packing reduce the recognized CNF to an integer system of common bounded-width rows. Modular RREF derives pivots, with direct enumeration for small nullity and meet-in-the-middle for larger supported nullities. Forward evaluation and clause checking validate SAT candidates; unsupported shapes or failed searches return UNKNOWN, with no UNSAT certificate path.

002

46 GBD instances · 10 validation · baseline 10/10

2D strip-packing asks whether rectangles can be placed in a strip so no pair overlaps on both axes while encoded capacity and geometry conditions hold. CNF represents axis-overlap relations, rectangle-to-column incidence, and auxiliary gates enforcing the required structural constraints. Proof format

No verif.

+ verif.

Solved

GRAT

0.210 ×

0.424 ×

9/10

Activity-based watched-literal CDCL branches with a bonus for overlap variables, then uses 1-UIP learning, restarts, and clause reduction for general fallback search. It handles only the recognized encoding schema: SAT assignments are checked clause by clause, while UNSAT traces are written as textual DRAT for external elaboration and checking; other inputs return UNKNOWN. DPR

0.208 ×

0.424 ×

9/10

Overlap-first branching biases watched-literal CDCL toward recovered pair-overlap variables before auxiliary variables, with first-UIP learning and standard heuristic fallback. Recognition is limited to the bounded normalized schema; SAT assignments are checked, while UNSAT searches are replayed when needed before emitting textual DPR learned-clause and deletion lines; unsupported inputs or proof failures return UNKNOWN. VeriPB

0.616 ×

0.805 ×

10/10

Structural-prefix branching prioritizes overlap and incidence variables, leaving gate variables to propagation and falling back to an unset suffix variable when needed. 1-UIP learning supports the search; SAT models are checked, while UNSAT emits VeriPB RUP steps and a conclusion, with trimming only in small cases; unsupported inputs or model/proof failures return UNKNOWN.

agile

003

2,597 GBD instances · 520 validation · baseline 499/520

The benchmark asks whether a bit-blasted Boolean or bit-vector circuit has an assignment satisfying its CNF encoding. The clauses express wire aliases, XOR relations, gate definitions, and residual circuit assertions. Proof format

No verif.

+ verif.

Solved

GRAT

0.120 ×

0.125 ×

166/520

Failed-literal probing, binary-implication congruence, and exact XOR recovery drive shallow split refutations on smaller instances, with a narrow CDCL fallback. Larger instances receive bounded model trials and may return UNKNOWN; SAT outputs checked assignments, while supported UNSAT paths emit DRAT clauses ending in the empty clause for external checking. DPR

0.208 ×

0.213 ×

341/520

A randomized constructive pass recovers aliases, XORs, and forward gates to build models, then falls back to bounded elimination, probing, and CDCL. Checked SAT models may return before proof content is added; UNSAT writes witness-free DPR additions for external checking, while failed checks yield UNKNOWN. VeriPB

0.100 ×

0.105 ×

98/520

A 64-lane bit-parallel circuit simulation samples circuit inputs, enumerating some patterns and assigning deterministic pseudorandom words to others; surviving assignments are expanded and checked. Failure to find a witness is not UNSAT, while eligible residuals use signed-alias quotienting, gate congruence, and CDCL with RUP PB proof output, and other cases return UNKNOWN.

Harrison Green, Claire Le Goues, and Fraser Brown

algebra

004

alloy-vpn-models

006

8 GBD instances · 2 validation · baseline 0/2

15 GBD instances · 3 validation · baseline 3/3

The benchmark asks whether two nontrivial binary coefficient vectors have a product equal to the identity in a table-indexed algebra. CNF encodings use pairwise AND variables and XOR constraints requiring odd parity in the identity class and even parity in every other product class.

These benchmarks ask whether a Boolean assignment satisfies a circuit-shaped CNF encoding a relational model, including auxiliary gate variables and one asserted condition. The instances use structured binary implications and longer gate clauses rather than arbitrary CNF. Proof format

No verif.

+ verif.

Solved

Proof format

No verif.

+ verif.

Solved

GRAT

0.220 ×

0.222 ×

1.00 ×

1.00 ×

0/2

2/3

GRAT

After recognizing the prescribed CNF layout, it enumerates weight-three supports for one coefficient vector and solves the resulting product constraints for the other with packed GF(2) Gaussian elimination. A validated witness yields a complete checked SAT assignment; unsupported inputs or failure to find this restricted witness return UNKNOWN, with no UNSAT certificate path. DPR

1.00 ×

0/2

1.00 ×

After recognition, it searches coefficient bits with row and column bitsets, greedy parity-improving flips, restarts, and occasional boundary-basis choices. If that incomplete search fails, coefficient-decision CDCL can produce a checked model or record DPR additions on a successful UNSAT path; recognition, model completion, or certification failure returns UNKNOWN. VeriPB

2.00 ×

2.00 ×

DPR

005

0.445 ×

0.435 ×

3/3

It reconstructs signed-AND gates with parity union-find, then searches leaf representatives using DAG propagation and circuit-aware CDCL. A fallback can emit unit and learned clause additions as a DRAT trace for external elaboration and checking; unsupported encodings or failed searches return UNKNOWN, while SAT assignments are expanded and CNF-checked. VeriPB

1/2

Strict structural decoding evaluates a fixed 21-word two-generator construction over represented inverse pairs and compares its product parities with the decoded right-hand side. It reconstructs and checks all auxiliaries before emitting a complete assignment; failed recognition or construction returns UNKNOWN, with no UNSAT certificate path.

algorithm-equivalence-checking

Signed-AND circuit recognition contracts parity-equivalent variables and propagates assignments through a gate DAG, then tries restricted searches before a circuit-aware CDCL fallback. SAT assignments are expanded and checked against the original CNF; a level-0 fallback conflict can emit a DRAT trace for external elaboration, while unsupported shapes or failed searches return UNKNOWN.

0.802 ×

0.766 ×

3/3

On selected accepted shapes, support-pair gate normalization reorders clauses before watched-literal CDCL; other accepted shapes use the generic watched-literal path. SAT assignments are printed as complete models, but UNSAT has no active certificate path: lower-scope results lack certification and the largest hard-coded scope returns UNKNOWN.

antibandwidth

007

36 GBD instances · 8 validation · baseline 4/8

187 GBD instances · 38 validation · baseline 27/38

The benchmark asks whether a Boolean circuit has an input on which two represented computations produce different outputs. Tseitin-style CNF introduces variables for gate values and asserts an OR of XOR output differences, so UNSAT corresponds to equivalence for the encoded pair.

Antibandwidth asks whether graph vertices can occupy distinct positions so adjacent vertices are far apart by at least w. CNF uses vertex-position variables, coverage constraints, and clauses forbidding edge endpoints from being too close. Proof format

No verif.

+ verif.

Solved

Proof format

No verif.

+ verif.

Solved

GRAT

0.894 ×

0.931 ×

20/38

GRAT

0.619 ×

0.673 ×

1/8

Miter-shape recognition validates the restricted circuit structure, then bounded variable elimination simplifies it before watched-literal CDCL with first-UIP learning and restarts. Recognition does not establish sorting semantics and unsupported shapes return UNKNOWN; validated SAT assignments are printed, while UNSAT produces a DRAT-style clause stream. DPR

0.689 ×

0.755 ×

2/8

A narrow gate recognizer accepts only the prescribed gate grammar and XOR/OR miter tail, then bounded variable elimination feeds watched-literal CDCL. Elimination resolvents and learned clauses are emitted as witness-free additions with deletions; contradiction yields UNSAT, but satisfying results or unsupported inputs return UNKNOWN. VeriPB

0.614 ×

0.679 ×

1/8

Exact miter recognition admits only the supported topological XOR/OR shape, then bounded elimination and self-subsuming resolution preprocess the CNF before watched-literal CDCL with first-UIP learning. UNSAT paths emit VeriPB resolution and RUP steps; SAT paths print checked assignments, while unsupported shapes or internal failures return UNKNOWN.

Release-time permutations and focused tabu swaps seek a labeling that separates every edge, guiding exchanges by conflicts and distance deficits. It validates and emits a completed SAT assignment, but rejects other structures as UNKNOWN, has no UNSAT certificate path, and may continue without resolving a failed search. DPR

1.54 ×

1.52 ×

27/38

Capacity-constrained graph coloring, layered ordering, and local swaps first seek an antibandwidth order; otherwise watched-literal CDCL completes the recognized formula. A validated completion supplies the SAT certificate; UNSAT uses witness-free direct or learned additions plus an empty clause, while unsupported structures return UNKNOWN. VeriPB

0.919 ×

0.937 ×

20/38

Maximum-linear-arrangement seeding and stochastic swap search build a vertex permutation, using min-conflicts and penalty objectives to improve edge separation. It emits a validated SAT assignment; the checked recognizer accepts only this structure, bounded search returns UNKNOWN without an ordering, and no UNSAT certificate path is implemented.

The Case for Automated Hyperspecialization: Evidence from SAT

argumentation

008

auto-correlation

010

217 GBD instances · 40 validation · baseline 9/40

51 GBD instances · 11 validation · baseline 2/11

The benchmark asks whether a directed argumentation instance admits a conflict-free, defended, or stable set of arguments, or a complete labeling. CNF encodings use status variables and clauses expressing attacks, defense, coverage, and sometimes exactly-one choices, though supported layouts are only restricted patterns.

The benchmark asks whether a binary sequence has bounded aperiodic autocorrelation or whether two cyclic subsets have constant combined intersection counts at every nonzero shift. CNF encodings use primary variables and auxiliary gates or counters to express these constraints.

Proof format

No verif.

+ verif.

Solved

Proof format

No verif.

+ verif.

Solved

1.60 ×

1.60 ×

5/11

GRAT

0.970 ×

0.978 ×

7/40

GRAT

Exact recognition of the expected two-label clause layout drives projection onto in variables, then bounded graph search and randomized maximal conflict-free-set repair precede watched-literal CDCL. SAT assignments are checked against the original CNF; only this recognized projection can emit an UNSAT certificate, while other branches seek SAT and failed proof reruns return UNKNOWN. DPR

0.866 ×

0.872 ×

3/40

Structural recognition of strict two-block and three-block layouts drives projection onto in variables, followed by bounded graph search and randomized maximal conflict-free-set repair before watched-literal CDCL. SAT assignments are checked against the original CNF; only the two-block rerun emits UNSAT, while other layouts or failed searches return UNKNOWN, without internally checking that stream. VeriPB

0.936 ×

0.943 ×

6/40

Repeated-block status recovery exposes an attack graph for grounded fixed-point propagation, randomized kernel repair, and compressed CDCL over IN variables. SAT witnesses are checked on the original CNF; failure to find a stable kernel can return UNKNOWN on weaker encodings, while only conservative recognized cases emit an UNSAT proof.

at-least-two-sol

009

Structural recognition drives multiplier-orbit enumeration and cyclic intersection matching for cyclic subset instances, while tabu bit flips search recognized aperiodic sequence layouts. Smaller cyclic cases use randomized annealing; larger or unsupported cases, low-goal aperiodic cases, and exhausted searches return UNKNOWN, while residual WalkSAT completion can produce SAT assignments without an UNSAT certificate path. DPR

0.871 ×

0.871 ×

0/11

It searches recognized cyclic subset layouts by enumerating bounded multiplicative-orbit unions and matching cyclic difference signatures, including complementary, swapped, and rotated variants. After fixing primary bits, watched-literal propagation plus an internal full clause check validate complete SAT assignments; failed recognition or construction returns UNKNOWN, with no UNSAT or DPR certificate path. VeriPB

1.91 ×

1.91 ×

6/11

It dispatches recognized aperiodic layouts to correlation-maintaining simulated annealing with restarts, and cyclic layouts to bounded multiplicative-orbit enumeration with hashed meet-in-the-middle autocorrelation matching. Propagation, uniform fills, and residual WalkSAT completion precede full clause verification; unsupported, exhausted, and low-bound aperiodic cases return UNKNOWN, with no general SAT fallback or UNSAT/VeriPB certificate path.

18 GBD instances · 4 validation · baseline 2/4

The benchmark asks whether a Boolean CNF formula has at least two satisfying assignments, rather than merely one. A common encoding places two copies of the formula in one CNF, links corresponding variables with equality indicators, and requires at least one pair to differ. Proof format

No verif.

+ verif.

Solved

GRAT

0.599 ×

0.688 ×

0/4

A bounded 2-CNF implication/SCC search and a permutation-CSP search first seek two models after exact layout recognition. Fallback watched-literal CDCL uses a blocking clause for the second model; unsupported layouts or failure to find that model return UNKNOWN rather than UNSAT, while only fallback UNSAT paths emit textual DRAT for external elaboration. DPR

0.785 ×

0.883 ×

1/4

An active 2-CNF-plus-wide-clause search and a permutation-CSP search first seek two models after exact layout recognition. Fallback watched-literal CDCL uses a blocking clause for another model; unsupported layouts or failed second-model validation return UNKNOWN, while UNSAT paths emit witness-free textual DPR clause additions ending in the empty clause, to be elaborated and checked externally. VeriPB

0.599 ×

0.688 ×

0/4

An AIG structural-hashing shortcut and bounded Davis-Putnam elimination target selected UNSAT cases, retaining information for RUP replay and model extension. Fallback component-wise watched-literal CDCL finds one model, blocks it, and seeks a second; checked witnesses are accepted, while failed recognition or second-model search returns UNKNOWN and UNSAT paths close with VeriPB RUP certificates.

automata-synchronization

011

12 GBD instances · 3 validation · baseline 3/3

Does a bounded-length word send every state of a complete two-letter deterministic automaton to one state? Its time-expanded CNF tracks reachable states and letter choices, propagates transitions, and requires at most one final reachable state. Proof format

No verif.

+ verif.

Solved

GRAT

21.1 ×

2.21 ×

3/3

For recognized inputs, structural Cerny recognition first yields a closed-form word; otherwise bit-packed image-set beams and reverse-BFS pair merging seek a bounded reset word. Candidate words become checked assignments; failed searches invoke bounded small-subset or antichain image-exclusion clauses for external DRAT elaboration, while unsupported or out-of-bound cases return UNKNOWN. DPR

0.101 ×

0.134 ×

1/3

After recognizing the fixed two-letter image-set schema, it tries Cerny detection, image-set beams, greedy pair merging, then watched-literal CDCL. Direct words become checked assignments; CDCL emits learned and pair-distance clause additions, reporting UNSAT only after a level-zero contradiction. Missing transition clauses disable pair-distance preprocessing, and unsupported or uncertified cases return UNKNOWN. VeriPB

489 ×

26.7 ×

3/3

After recognizing the reachable-subset encoding, it uses Cerny threshold dispatch and exact-encoding interval reasoning, then falls back to duplicate-eliminating bit-parallel beams. If searches find no word, bounded pair/triple-distance reasoning emits VeriPB proofs; SAT assignments are clause-checked, but the solver does not validate proof files and unsupported or unproved cases return UNKNOWN.

Harrison Green, Claire Le Goues, and Fraser Brown

baseball-lineup

012

belpyramid-puzzle

014

40 GBD instances · 8 validation · baseline 7/8

57 GBD instances · 12 validation · baseline 6/12

The benchmark asks whether exactly K items can be selected so that every binary attribute receives at least its required coverage. Its CNF encodes the selection, item-attribute incidence, and cardinality or coverage counts with auxiliary variables and sequential counters.

These instances ask whether two Boolean networks with shared inputs differ on some output. CNF clauses encode AND gates and output comparisons; a required mismatch makes SAT witness inequivalence and UNSAT indicate equivalence; this is not every possible encoding.

Proof format

No verif.

+ verif.

Solved

Proof format

No verif.

+ verif.

Solved

GRAT

1.98 ×

1.94 ×

7/8

GRAT

0.638 ×

0.795 ×

1/12

After recognizing the row-mask-counter schema, weighted randomized greedy selection with targeted min-conflicts swaps drives the primary search. It falls back to a bounded DPLL tree; availability conflicts emit RUP-style units and an empty clause, while unresolved cases return UNKNOWN. SAT assignments reconstruct auxiliaries and are checked against all input clauses. DPR

1.98 ×

1.65 ×

7/8

Strict row-mask-counter recognition precedes deficit-weighted greedy selection with restarts and swap-delta improvement. Failure triggers exact row conflict search only for small K; larger unresolved cases return UNKNOWN. UNSAT uses sparse-column units or conflict refutations, may add witness-free clauses, and is not internally verified; SAT models are clause-checked. VeriPB

0.992 ×

0.993 ×

6/8

RUP-checked simulation lemmas drive solving on normalized AIG miters, with multiword signatures proposing equivalences before a monolithic watched-literal CDCL fallback. First-UIP learning and reason minimization support fallback search; SAT models are checked, the pyramid branch returns UNKNOWN, and UNSAT emits textual DRUP-style additions for external checking. DPR

0.698 ×

0.870 ×

2/12

Signature-guided equivalence sweeping drives incremental assumption-based CDCL on a recognized AND-miter layout. Multiword signatures propose equivalent or complementary variables, with a complementary-input prepass; bounded queries retain learned clauses and final conflicts emit proof clauses, while failed searches or unsupported layouts return UNKNOWN without a SAT witness. VeriPB

0.690 ×

0.825 ×

2/12

Scarcity-weighted greedy search with breakout swaps follows strict incidence-counter recovery. Failure falls back to a bounded primary-variable tree, then weighted PB separation with lifted-counter proof logging; bounded or unsupported cases return UNKNOWN. SAT candidates are clause-checked; UNSAT emits RUP or PB derivations, including deficient-support contradictions.

Comparator-cone probing leads the compact-miter search: assumed output mismatches are tested with restricted watched-literal CDCL, then the full miter receives unrestricted CDCL. Learned clauses remain global and UNSAT is logged with VeriPB RUP records, but rejected finite-domain layouts and all non-UNSAT outcomes return UNKNOWN without a SAT witness.

battleship

binary-pigeon-hole

013

015

45 GBD instances · 9 validation · baseline 8/9

5 GBD instances · 1 validation · baseline 0/1

The benchmark asks whether selected cyclic modular lines cover every point of an n-by-n toroidal board. It encodes line choices with Boolean variables, coverage clauses for every point, and within-block at-most-one clauses.

Binary pigeon-hole CNFs ask whether p pigeons can be assigned distinct holes using binary codes. A direct encoding gives each pigeon a bit block, excludes unused codes, and forbids two pigeons from sharing a valid code.

Proof format

Proof format

No verif.

+ verif.

Solved

GRAT

4,340 ×

1.40 ×

1/1

GRAT

No verif.

421 ×

+ verif.

Solved

35.4 ×

9/9

A direct square proof handles complete square instances, while bounded prime-power construction and weighted min-conflicts seek checked SAT assignments elsewhere. If these searches fail, affine-symmetry lifting or watched-literal CDCL can emit DRAT traces for applicable cases; unsupported or unresolved inputs return UNKNOWN, and UNSAT requires successful proof generation. DPR

0.414 ×

0.420 ×

6/9

Compressed-word min-conflicts tracks line coverage with exact deltas and restarts, trying one permitted omitted pair with both rows forced before one-row search unless its negative-case guard applies. After heuristic failure, bounded unit propagation emits a DPR-style stream with witness-free additions; limits or unsupported layouts return UNKNOWN. VeriPB

2.49·104 ×

4,470 ×

9/9

A direct cutting-planes contradiction handles complete square instances, while a prime-field two-affine-pencil construction supplies checked SAT assignments on qualifying complete instances. Otherwise grouped set-cover min-conflicts uses coverage counts, tabu moves, and bounded restarts before selected CDCL fallbacks; specialized proof routes certify limited UNSAT cases, while unresolved recognized inputs return UNKNOWN.

Exact structural recognition replaces general SAT search. For recognized formulas, it translates codes into unary membership and capacity constraints, repeatedly reduces pigeon-hole instances with fresh bridge variables, and emits an empty clause in a DRAT-style stream for external elaboration and checking. Recognition or proof-writing failure returns UNKNOWN. DPR

3,170 ×

0.620 ×

1/1

Exact recognition of the complete direct encoding drives a specialized path without general search. SAT assigns pigeon i code i and rechecks the clauses; UNSAT adds exact-code indicators and unary pigeon-hole constraints, then uses witness-free clause additions and propagation-redundancy symmetry steps to reach an empty clause; unsupported cases return UNKNOWN. VeriPB

3.51·104 ×

264 ×

1/1

Exact canonical-encoding recognition, including an order-independent retry, and a binary-code assignment shortcut replace general SAT search. SAT assigns pigeon i code i and checks clauses; UNSAT emits a VeriPB certificate with exact-code indicators, per-hole pseudo-Boolean capacity derivations, and per-pigeon binary resolution trees; unsupported inputs or proof-generation failures return UNKNOWN.

The Case for Automated Hyperspecialization: Evidence from SAT

binary-tree-parity

016

bitvector

018

2 GBD instances · 1 validation · baseline 0/1

594 GBD instances · 119 validation · baseline 49/119

These instances ask whether a Boolean assignment satisfies three-variable parity equations whose supports form two full binary trees, possibly in different variable orderings. In CNF, each equation is represented by clauses forbidding the four assignments of one parity, with some instances allowing a shortened block.

The benchmark asks whether a Boolean assignment can satisfy a CNF encoding of quantifier-free bit-vector constraints, with variables representing bit values and circuit intermediates. Clauses enforce translated operations and asserted constraints. Proof format

No verif.

+ verif.

Solved

Proof format

GRAT

0.739 ×

0.855 ×

20/119

GRAT

No verif.

+ verif.

Solved

7.42·105 ×

3.23·105 ×

1/1

Packed GF(2) Gaussian elimination solves recognized parity groups, with bounded affine-space enumeration and direct checking of candidates against the CNF. It falls back to order-independent matching for shortened gates; SAT emits a checked assignment, while unsupported or unsuccessful cases return UNKNOWN because no UNSAT certificate path is implemented. DPR

3.21·105 ×

6.25·104 ×

1/1

Exact recognition of the gate order and orientation precedes packed GF(2) Gaussian elimination and bounded Gray-code enumeration of affine solutions, with SAT candidates checked against clauses. It retries weakened gates by fixing the omitted-clause assignment and attempts Davis-Putnam refutation only for small no-model cases; unsupported or uncertified outcomes return UNKNOWN. VeriPB

3.17·105 ×

6.39·104 ×

1/1

Affine reduction through the recognized first tree produces a packed GF(2) system for the second. It also enumerates bounded shortened-gate branches and checks candidates against the CNF. Failed-literal propagation yields UNSAT only on a RUP contradiction; failed or unsupported cases return UNKNOWN, while SAT emits a checked assignment.

It recognizes local gate blocks and, when coverage is sufficient, tries bit-parallel functional-block search; otherwise it uses watched-literal CDCL with activity branching, 1-UIP learning, restarts, and clause reduction. SAT models are rechecked, while UNSAT is represented by textual DRAT additions for external elaboration; failures can return UNKNOWN. DPR

0.737 ×

0.846 ×

VeriPB

0.756 ×

0.852 ×

017

61 GBD instances · 13 validation · baseline 10/13

The benchmark asks whether Boolean variables satisfy a structured set of constraints; some instances organize them as one-of-many choices at positions with compatibility or circuit conditions. CNF uses one-hot clauses and auxiliary Tseitin variables, but these recognized shapes do not define every possible encoding. Proof format

No verif.

+ verif.

Solved

GRAT

0.461 ×

0.691 ×

6/13

Conservative XITS recognition enables canonical partition branching in small-class cases; other inputs use activity-ordered first-UIP CDCL, with failed-literal probing and bounded elimination only on the generic path. Checked SAT models are printed, while reported UNSAT is accompanied by textual DRAT ending in an empty clause; unresolved search or reconstruction failure returns UNKNOWN. DPR

0.362 ×

0.532 ×

4/13

Bounded Davis-Putnam elimination protects the one-hot and counter prefix; Xits cases also bias state choices and probe same-state pairs before CDCL fallback. SAT assignments are reconstructed and checked; internal UNSAT is logged with preprocessing and learned clauses plus an empty clause for external DPR elaboration, while unrecognized inputs or failed checks return UNKNOWN. VeriPB

0.485 ×

0.436 ×

6/13

Detected Xits shapes activate positive phases on choice blocks, Xits-gated minimization, and learned-clause reduction; other inputs use activity-based first-UIP CDCL with phase saving and Luby restarts. Checked SAT assignments are printed, while UNSAT requires written VeriPB RUP records ending in an empty record; parse, check, or proof-write failure returns UNKNOWN.

23/119

Root failed-literal probing and bounded proof-logged variable elimination precede generic watched-literal first-UIP CDCL; no bit-vector semantics are used. SAT assignments are independently checked, while UNSAT emits VeriPB RUP steps, deletions, and a final empty clause, with parsing, checking, or proof-writing failures yielding UNKNOWN.

bounded-model-checking bioinformatics

21/119

Local gate-pattern detection supplies only a heap tie-break, not bit-vector or XOR reasoning; the solver otherwise runs generic watched-literal first-UIP CDCL with phase saving and restarts. Its addition-only proof records learned clauses and the empty clause at root conflict, while SAT models are checked and printed separately, not in the proof.

019

31 GBD instances · 7 validation · baseline 7/7

These benchmarks ask whether a Boolean assignment satisfies a Tseitin-encoded circuit check. Variables represent circuit or state signals, and CNF clauses impose local gate and checking constraints. Proof format

No verif.

+ verif.

Solved

GRAT

0.0612 ×

0.0605 ×

7/7

It recognizes dense Tseitin-like AND/OR and XOR/XNOR structure, then applies bounded elimination to gate outputs; unsupported shapes return UNKNOWN rather than invoking general search. CDCL with first-UIP learning is the fallback, emitting DRAT additions for external elaboration on UNSAT and checking reconstructed SAT models against the original clauses. DPR

0.0549 ×

0.0390 ×

7/7

It starts with reverse-order bounded variable elimination, retaining non-tautological resolvents and recording eliminated values for reconstruction; watched-literal CDCL with first-UIP learning is the fallback when no contradiction appears. UNSAT uses addition-only RUP-style clause additions, while reconstructed SAT assignments are checked against original clauses; failed reconstruction or validation returns UNKNOWN, and proof logging is configurable. VeriPB

0.0471 ×

0.0168 ×

7/7

It begins with parity-DSU substitution and conservative gate hashing, adding bounded elimination only for a recognized signature; unrecognized or malformed inputs return UNKNOWN. CDCL with first-UIP learning is the fallback, with UNSAT logged as augmented VeriPB RUP additions and reconstructed SAT assignments checked against the original CNF.

Harrison Green, Claire Le Goues, and Fraser Brown

brent-equations

020

cellular-automata

022

20 GBD instances · 4 validation · baseline 3/4

52 GBD instances · 11 validation · baseline 8/11

These benchmarks ask whether the 3 by 3 matrix-multiplication tensor over GF(2) has a rank-r decomposition into binary factor vectors. A Tseitin-style CNF represents products and XOR accumulation, yielding 729 parity equations that describe the tensor.

These benchmarks ask whether a cyclic binary row has a predecessor trajectory under a local cellular-automaton rule for T steps ending in prescribed values. CNF represents cell states by time layer, links local neighborhoods with transition clauses, and fixes terminal values with unit clauses.

Proof format

No verif.

+ verif.

Solved

GRAT

1.06 ×

1.06 ×

3/4

Proof format

No verif.

+ verif.

Solved

A 26-term Strassen-style construction drives recognized unfixed high-rank cases, while fixed cases use DSU component decomposition, bounded multi-choice search, and bipartite matching after gate-network recognition. Incremental residual local search handles eligible high-rank failures; clause-checked SAT assignments are emitted, but no UNSAT certificate path exists and unsuccessful cases return UNKNOWN.

GRAT

1.29 ×

1.44 ×

9/11

3/4

DPR

DPR

1.06 ×

1.06 ×

XOR-root monomial flattening and bipartite matching first construct a primary assignment for recognized high-rank encodings. When direct construction fails, an open-ended weighted WalkSAT search may precede watched-literal CDCL; SAT assignments are checked and emitted, while unsupported, low-rank, or UNSAT cases return UNKNOWN and no UNSAT certificate path is implemented. VeriPB

1.06 ×

1.06 ×

3/4

Direct reconstruction of the fixed generator schedule drives sparse rank-one completion; substituted dense cases use bit-parallel tabu search with incremental deltas and repair. Watched-literal CDCL follows failed direct searches; clause-checked SAT assignments are emitted, but low-rank, contradictory, or failed searches return UNKNOWN and no UNSAT certificate path exists.

cardinality-constraints

0.520 ×

VeriPB

2/11

0.692 ×

0.698 ×

6/11

Recognized Rule-110 layouts first try a cyclic-predecessor check, then use backward row-biased CDCL; population signatures may activate bounded elimination, while unrecognized inputs return UNKNOWN. Tracked UNSAT replay emits VeriPB RUP dependencies and an empty conclusion, while direct predecessor derivations are limited to targets without a cyclic one-step predecessor; SAT assignments are checked.

circuit-equialence-checking

The benchmark asks whether at most K lattice points can hit every specified geometric object, such as a square or triangle. CNF uses Boolean variables for point selections, clauses requiring each object to be hit, and auxiliary clauses encoding the at-most-K bound. Proof format

No verif.

+ verif.

Solved

GRAT

0.506 ×

0.779 ×

3/4

Incidence recognition drives a canonical hitting-set witness for aligned squares, with min-conflicts as a SAT fallback; propagation completes the model before checking it. For arbitrary squares and aligned triangles, mapped templates or DRAT-producing 1-UIP CDCL handle UNSAT; aligned squares lack a CDCL fallback, and unsupported or failed paths return UNKNOWN. 0.301 ×

0.633 ×

2/4

Bounded geometric regeneration and weighted min-conflicts construct a capped hitting set; propagation then completes and checks the original CNF for SAT. On failure, bounded resolution elimination feeds GenericCDCL, whose resolution and learned-clause additions support an UNSAT certificate; unsupported or altered inputs, or failure of that path, return UNKNOWN. VeriPB

0.455 ×

Structural clause-width dispatch selects bounded and deterministic phase passes for non-ECA layouts, then falls back to watched-literal CDCL with first-UIP learning; recognized ECA layouts return UNKNOWN unless ECA_CDCL enables ordinary CDCL, with no specialized reverse-search path. SAT assignments are checked, while UNSAT learning is emitted as witness-free DPR additions, including the empty clause.

021

17 GBD instances · 4 validation · baseline 4/4

DPR

Backward traversal of recognized layered gates biases search from the fixed terminal row, while a shallow clause-width SPG heuristic tries sequential phase passes before the generic watched-literal CDCL fallback. SAT assignments are checked; UNSAT results use textual DRAT additions, with recognized-path clauses replayed chronologically, while an exhausted SPG budget returns UNKNOWN.

0.302 ×

0.633 ×

2/4

Fixed-cardinality simulated annealing searches recognized exact geometric layouts while maintaining unhit objects; unit propagation and small recursive DPLL then complete the assignment and check every original clause. No UNSAT certificate path is active: recognition, construction, or completion failure returns UNKNOWN.

023

19 GBD instances · 4 validation · baseline 0/4

The benchmark asks whether two Boolean circuits compute the same function. A CNF miter combines their outputs with an XOR and forces it true, so SAT gives a distinguishing input while UNSAT establishes equivalence for the gate encoding. Proof format

No verif.

+ verif.

Solved

GRAT

1.00 ×

1.00 ×

0/4

Strict LUT/miter recognition drives bit-parallel simulation of small root supports, checking any distinguishing assignment against the original clauses. Watched-literal DPLL handles larger supported roots and logs learned or cube-blocking clauses as textual DRAT additions, ending UNSAT with an empty clause; unsupported layouts or excessive support return UNKNOWN. DPR

1.00 ×

1.00 ×

0/4

Semantic sweeping groups equal or complemented signals using bit-parallel truth-table simulation, then tests candidate implications by propagation and RUP. Common-cut lemmas and bounded CDCL add witness-free clause additions, ending proved UNSAT with an empty clause; bounded failure returns UNKNOWN, and SAT models are checked. VeriPB

1.00 ×

1.00 ×

0/4

Packed signatures find equal or complemented candidate correspondences, certified first by bounded parity-DSU exact cuts with bit-parallel evaluation. Relevant-cone CDCL falls back when cuts fail, reusing proved pairs and logging RUP or polynomial-resolution steps; internal SAT only rejects candidates, no SAT assignment is emitted, and unsupported or unproved roots return UNKNOWN.

The Case for Automated Hyperspecialization: Evidence from SAT

circuit-equivalence-checking

024

circuit-multiplier

026

20 GBD instances · 4 validation · baseline 1/4

20 GBD instances · 4 validation · baseline 2/4

The benchmark asks whether two Boolean circuit signals agree on every shared input assignment. Gate relations are encoded in CNF and a difference (XOR) output is asserted, so SAT supplies a counterexample and UNSAT establishes equivalence.

The benchmark asks whether two nontrivial positive binary numbers, represented little-endian, can multiply to the fixed target encoded by the instance. CNF clauses constrain operand bits and auxiliary multiplier, adder, and carry wires so an assignment realizes that product.

Proof format

No verif.

+ verif.

Solved

GRAT

3.01 ×

3.07 ×

3/4

Bit-parallel SAT sweeping leads the search: signatures on a recognized acyclic LUT network propose equal or complemented signals, and CDCL assumptions test them against counterexamples. Bounded functional search covers exceptional inputs; fallback CDCL logs textual DRAT for external checks on UNSAT, while unsupported inputs return UNKNOWN and SAT yields a checked model. DPR

1.52 ×

2/4

1.56 ×

Assumption-based relation queries lead the search: incremental watched-literal CDCL tests candidate signal equalities from a bounded signature sweep, with counterexamples refining the signatures before final fanin queries. Failed candidates are discarded; unsupported or uncertifiable encodings return UNKNOWN, SAT gives a checked model, and only the final UNSAT path emits a dependency-tracked RUP proof core. VeriPB

0.763 ×

0/4

0.782 ×

Bit-parallel simulation leads: replayed lanes can produce clause-checked SAT models; otherwise bounded gate elimination feeds embedded CDCL, whose UNSAT path logs a VeriPB proof with RUP constraints and an empty clause, while fallback SAT is not checked against the original formula and nonmatching encodings return UNKNOWN.

Proof format

No verif.

+ verif.

Solved

GRAT

0.789 ×

0.789 ×

1/4

Watched-literal CDCL is the active engine, with first-UIP learning, restarts, rephasing, and clause reduction after a syntactic two-block check. SAT models are checked against the original CNF; closed non-SAT searches return UNKNOWN because no retained DRAT proof is produced. DPR

0.620 ×

0.620 ×

VeriPB

0.757 ×

0.757 ×

025

50 GBD instances · 10 validation · baseline 0/10

Proof format

No verif.

+ verif.

Solved

GRAT

9.95 ×

9.95 ×

9/10

Structural recognition and population-counter probing recover candidate phases and the exact cover size, with a per-bit fallback. Weighted greedy cover search uses focused exchanges and changing target weights; counter propagation completes SAT models, whose original clauses are checked, while unsupported or unsuccessful cases return UNKNOWN without an UNSAT certificate path. 9.97 ×

9.92 ×

9/10

Phase normalization identifies candidate inputs and cover constraints, while bounded cardinality probing recovers the target size in two exact layouts. A randomized weighted add/remove search seeks the fixed-cardinality cover, then bounded propagation/DPLL completes variables; checked SAT assignments are emitted, but failed or unsupported cases return UNKNOWN because no UNSAT certificate path exists. VeriPB

9.94 ×

9.88 ×

027

15 GBD instances · 3 validation · baseline 3/3

The CNF asks whether some edge assignment contains k distinct selected vertices forming a clique and assigns every vertex a color so adjacent vertices differ. Edge, selection, and color variables encode these conditions with exactly-one, distinctness, and incompatibility clauses.

Structured circuit-minimization encodings ask whether exactly a prescribed number of Boolean candidates can cover every requirement; selector clauses express coverage, and Tseitin population-count clauses enforce the exact cardinality.

DPR

1/4

Bounded variable elimination follows strict recognition, removes only internal variables while protecting operand words, and retains reverse-reconstruction records; unmatched inputs return UNKNOWN. Residual watched-literal CDCL uses first-UIP learning and restarts; reconstructed SAT models are checked against the original CNF, while internal UNSAT returns UNKNOWN without a VeriPB refutation.

clique-coloring circuit-minimization

0/4

Structural recognition admits only the documented operand blocks and circuit layout; malformed or unmatched inputs return UNKNOWN without generic SAT fallback. ProbSAT-style local search tracks unsatisfied clauses and break counts, then falls back to watched-literal CDCL for checked SAT models and witness-free DPR additions plus a root-contradiction marker on UNSAT.

9/10

Recognition of two exact signatures recovers selector phases, cover rows, and target cardinality from the population-count core, with repair for a shortened-gate variant. Weighted fixed-cardinality search focuses on uncovered rows and scores swaps with tabu diversification; clause-checked candidates yield SAT models, while failed or unsupported cases return UNKNOWN without an UNSAT path.

Proof format

No verif.

+ verif.

Solved

GRAT

31.1 ×

2.30 ×

3/3

Exact schema recognition and closed-form model construction are the active techniques. Its SAT branch checks a complete assignment; when k>c, it emits a DRAT proof with fresh variables and a Hall-style pigeonhole refutation, while unrecognized or unsupported cases return UNKNOWN instead of using generic search. DPR

10.7 ×

0.330 ×

3/3

Exact schema recognition gates both branches. When k>c, fresh-variable, witness-free clause additions and a subset-Hall refutation provide the UNSAT certificate; in the constructive regime, it constructs and checks a closed-form clique-coloring assignment. Unrecognized or unsupported cases return UNKNOWN rather than invoking generic search. VeriPB

26.4 ×

0.513 ×

3/3

Streaming recognition and direct model checking lead the solver. It constructs a checked SAT assignment when k<=n and c>=k; otherwise it emits bounded pseudo-Boolean proofs for n<k or k>c using position-pigeonhole reasoning or vertex symmetry, RUP, and cutting planes, while recognition or guard failures return UNKNOWN without general search.

Harrison Green, Claire Le Goues, and Fraser Brown

clique-formulas

028

clustered-random

030

4 GBD instances · 1 validation · baseline 1/1

43 GBD instances · 9 validation · baseline 8/9

The benchmark asks whether k ordered positions can choose vertices from different parts of a complete multipartite graph. Its CNF has one positive choice clause per position and binary clauses forbidding two choices in one position or the same part.

The benchmark asks whether a Boolean assignment satisfies every clause in a uniform 3- or 4-CNF formula whose clauses are arranged in overlapping clusters on limited variable supports. It represents a restricted clustered random SAT structure, not every possible encoding.

Proof format

GRAT

No verif.

+ verif.

Solved

447 ×

47.9 ×

1/1

Cross-row forbidden-neighbor signatures reconstruct rows and parts, followed by an exact conflict audit for the balanced encoding. It checks a complete SAT assignment when positions fit the parts; otherwise, for at most 16 parts, it writes a DRAT certificate for external elaboration and checking using OR definitions and subset-Hall clauses, while unsupported cases return UNKNOWN. DPR

820 ×

17.9 ×

1/1

DSU recognition recovers selector positions and parts, requiring the balanced clone encoding and full incompatibility set rather than general search. It checks a SAT assignment when positions fit recovered parts; otherwise it emits a DPR refutation with clone-collapse additions and a pigeonhole sequence, while unsupported inputs or failed checks return UNKNOWN. VeriPB

735 ×

9.00 ×

1/1

Twin-variable representative reduction and whole-part symmetry breaking compress the recognized balanced, auxiliary-free encoding instead of general clique search. It checks a SAT assignment when positions fit the parts; otherwise it writes a VeriPB proof ending in a RUP contradiction without verifying that proof; unsupported inputs or failures in model checks or proof output return UNKNOWN.

clique-width

029

Proof format

No verif.

+ verif.

Solved

GRAT

0.379 ×

0.498 ×

6/9

Focused local search with occurrence counters and bounded Hamming repair comes first, followed by watched-literal CDCL when no verified assignment is found. SAT returns a checked complete assignment; only the fallback logs learned clauses and an empty clause as a DRAT trace for external elaboration, while unsupported input or reconstruction failure returns UNKNOWN. DPR

0.481 ×

7/9

0.616 ×

Factor-local exhaustive repair and stochastic local search with belief-propagation decimation for width-4 cases precede factor-resolution clauses, bounded Davis-Putnam elimination, and watched-literal CDCL fallback. SAT exits return checked assignments, whereas only the fallback emits witness-free RUP-style clause additions toward an UNSAT derivation; unsupported inputs return UNKNOWN. VeriPB

0.470 ×

7/9

0.336 ×

Focused probSAT-style walking leads the search, with consensus-derived clauses on width-4 instances; failure triggers first-UIP CDCL with same-support consensus, bounded elimination, and a selective failed-literal pass. SAT returns a checked assignment, while UNSAT is certified through sliced VeriPB RUP steps; unsupported inputs or internal errors return UNKNOWN.

27 GBD instances · 6 validation · baseline 1/6

The benchmark asks whether a recovered graph can be built through n stages using at most k temporary labels, with each stage recording component relations, label choices, and graph constraints. These construction choices are represented by CNF variables and clauses. Proof format

No verif.

+ verif.

Solved

GRAT

1.00 ×

1.00 ×

1/6

A partition-state merge search tries forest-postorder, linear-neighborhood, and beam-bounded general constructions, pruning merges by outside-neighborhood signatures and implication conflicts; reused-label enumeration is an active fallback. The narrow recognized layout is required; structural assignments receive SCC-based 2-SAT completion and full clause checking, while search or checking failure returns UNKNOWN and no UNSAT certificate is emitted. DPR

1.00 ×

1.00 ×

1/6

A capped linear-order DFS seeks a graph ordering; success adds relation literals, while failure or cap exhaustion sends a binary normal form with representatives and Sinz-style counters to watched-literal CDCL. Only a narrow normalized layout is recognized; 32-bit masks limit support, and failed or unsupported cases return UNKNOWN without an UNSAT or DPR certificate. VeriPB

1.00 ×

1.00 ×

1/6

Graph recovery and a bounded component-hierarchy trial provide the active heuristic, after which failure falls back to unrestricted watched-literal CDCL rather than implying UNSAT. The exact generator-specific layout is required; SAT assignments are checked internally, while UNSAT emits a PB RUP proof, and malformed or unsupported inputs return UNKNOWN without external checking.

coloring

031

594 GBD instances · 119 validation · baseline 58/119

The benchmark asks whether each graph vertex can receive a color while respecting every encoded incompatibility between endpoint choices. These constraints are translated into CNF using compact Boolean codes or one-hot choices in the recognized instances, rather than one universal coloring encoding. Proof format

No verif.

+ verif.

Solved

GRAT

0.776 ×

0.784 ×

38/119

Strict recognizers send compact cases through bit-mask forward checking, with small-domain branching and impact-based state ordering; one-hot cases use CDCL directly. SAT assignments are checked against the CNF; failed bounded searches use CDCL or elimination, emitting DRAT clause additions for UNSAT, while unresolved cases return UNKNOWN. DPR

0.845 ×

0.854 ×

44/119

Incidence-signature recognition reduces supported compact encodings to finite-domain masks and binary support tables, followed by arc consistency and tabu/min-conflicts search. Validated assignments provide SAT witnesses; eligible failures fall back to watched-literal CDCL, emitting DRAT additions including an empty clause for UNSAT, while larger failed compact cases return UNKNOWN by default. VeriPB

0.879 ×

0.888 ×

47/119

Recognized encodings become finite-domain CSPs: min-conflicts repair is followed by propagation and bounded DSATUR-style search. SAT assignments are checked against the original CNF; exact rope signatures may use triangle-cycle dynamic programming, while CDCL or rope DP emit VeriPB proof steps, including RUP steps, for UNSAT; inconclusive one-hot or dense-choice searches return UNKNOWN.

The Case for Automated Hyperspecialization: Evidence from SAT

coloring-clique

032

core-based-generator

034

11 GBD instances · 3 validation · baseline 3/3

20 GBD instances · 4 validation · baseline 3/4

The benchmark asks whether an existentially chosen graph can contain a k-vertex clique while admitting a proper coloring with c colors. The CNF uses edge, clique-position, and color variables for graph choices, vertex selection, distinctness, and coloring conflicts.

The benchmark asks whether a Boolean CNF is satisfiable when a residual core is surrounded by repeated three-row gadgets over signed variables. Each gadget encodes all but one sign pattern with chain auxiliaries and links the omitted pattern to the core.

Proof format

Proof format

No verif.

+ verif.

Solved

GRAT

35.9 ×

61.6 ×

4/4

GRAT

+ verif.

Solved

1.35·10 −4 × 7.88·10−4 ×

No verif.

2/3

Exact structural recognition accepts only the expected numbering and complete clause set; malformed or unrecognized formulas return UNKNOWN. Recognized instances receive a clause-checked SAT assignment when k <= c; when k > c, the solver introduces position-color variables and emits a textual DRAT pigeonhole refutation, with generation failures returning UNKNOWN. DPR

1.35·10−4 ×

7.88·10 −4 ×

2/3

Exact clause-set recognition accepts only the recovered k = c+1 encoding; equivalent encodings return UNKNOWN. For matches, it skips SAT search, adds selected-color and derived-color clauses, reduces the contradiction to pigeonhole, and emits a DPR certificate ending in the empty clause; proof failures return UNKNOWN. VeriPB

1.35·10 −4 ×

7.87·10 −4 ×

2/3

Symmetry-based paired vertex swaps canonicalize selected vertices after accepting the exact contiguous layout and complete templates; alternate layouts or unsupported formulas return UNKNOWN. For c >= k, it checks a SAT assignment; for c < k, RUP exclusions and cutting planes derive a VeriPB pigeonhole contradiction, with generation failures returning UNKNOWN.

coloring-mycielski-graph

033

19 GBD instances · 4 validation · baseline 0/4

The benchmark asks whether a recursively constructed Mycielski graph has a proper coloring with one fewer color than its level. CNF encodings use vertex-color variables, exactly-one constraints, and edge clauses forbidding equal colors, with supported compact and permuted variants. Proof format

GRAT

No verif.

1.33 ×

+ verif.

0.786 ×

Solved

1/4

One-hot recognition drives an extension-variable recoloring proof that repeatedly reduces a Mycielski coloring instance to fewer colors. Shadow copies handle hinted levels, bit encodings first define color indicators, and an opt-in CDCL fallback can verify SAT assignments; without it, unsupported inputs return UNKNOWN, while refutations are emitted as DRAT text for external elaboration. DPR

1.01·104 ×

120 ×

4/4

It recognizes supported compact, permuted, and canonical layouts, reconstructing graph, color, and hint structure rather than running general SAT search. Complete hints drive RUP shadow-edge derivations and a PHP refutation; incomplete hints use fresh-variable one-hot recursion only at bounded ranks, while rejected inputs return UNKNOWN and no SAT-witness path exists. VeriPB

2.94·104 ×

271 ×

4/4

It recognizes canonical and recovered bit, one-hot, and XOR layouts, then uses color symmetry to fix the newest root and recursively derive a coloring instance with one fewer color. Proof-only indicators and recoloring steps support the induction to a final contradiction; failed recognition or certificate generation returns UNKNOWN, with no SAT-witness path.

Missing-sign decoding fixes gadget variables and reduces recognized inputs to output units plus a residual core for watched-literal CDCL with first-UIP learning. SAT assignments are checked; UNSAT emits a textual DRAT trace with gadget clauses, learned residual clauses, and a final contradiction, while rejected structures or failures return UNKNOWN. DPR

1.05 ×

1.01 ×

3/4

Strict structural extraction and preprocessing remove recognized padding, pure variables, and low-product occurrence pairs while recording resolvents, leaving a reduced core for watched-literal CDCL. SAT models are reconstructed and checked; UNSAT writes a textual DPR trace that can include witness-free clause additions, while unsupported structures or proof failures return UNKNOWN. VeriPB

1.05 ×

3.22 ×

3/4

Local search first attacks the bounded, structurally extracted core; if it fails, its best phase seeds watched-literal CDCL. SAT models are extended through the gadgets and checked; UNSAT emits VeriPB cutting-planes and RUP derivations, while unsupported structures or internal failures return UNKNOWN.

cover

035

18 GBD instances · 4 validation · baseline 1/4

The benchmark asks whether an n-point Steiner triple system has a cap of at least k points with no complete triple. CNF uses point variables, triple clauses forbidding all three selected, and auxiliary counters for the size threshold. Proof format

No verif.

+ verif.

Solved

GRAT

1.00 ×

1.00 ×

1/4

Recognized layouts first use randomized cap search with maximal-cap growth, local exchanges, tabu or annealing moves, and a special projective quotient search. Selected targets use watched-literal CDCL; checked SAT models are printed, while UNSAT emits learned clauses and the empty clause as DRAT. Unrecognized inputs or noncertifying failures return UNKNOWN. DPR

1.00 ×

1.00 ×

1/4

Recognized Steiner/BDD layouts use quasigroup cap search: completion conflicts trigger endpoint removal, refill, and periodic ruin-and-recreate. Selected cases use a propagation tree or watched-literal CDCL, emitting clause additions, some witness-free, that end in the empty clause; unsupported layouts and proof failures return UNKNOWN, while remaining regimes lack an UNSAT path. VeriPB

1.00 ×

1.00 ×

1/4

Embedded affine and product cap constructions lead, with bit-mask branch-and-bound and selected tabu-style repair on recognized layouts. Selected cases then use proof-logging CDCL or a tree proof; direct SAT models are checked and printed, but CDCL SAT results are not printed as models, and unsupported or noncertifying cases return UNKNOWN.

Harrison Green, Claire Le Goues, and Fraser Brown

crafted-cec

036

cryptography

038

20 GBD instances · 4 validation · baseline 4/4

860 GBD instances · 172 validation · baseline 84/172

The benchmark asks whether a primary-input assignment can make an asserted output of a signed AND circuit true. Each topologically ordered gate is represented by one wide clause plus binary clauses, and the final output is asserted by a unit clause.

These benchmarks ask whether a chosen Boolean input block can make a one-block cryptographic hash or compression circuit meet specified output constraints. The circuit becomes CNF with variables for input bits and intermediate values, plus clauses fixing constants, selected inputs, and outputs.

Proof format

No verif.

+ verif.

Solved

GRAT

23.7 ×

23.5 ×

4/4

Packed genetic search simulates the recognized circuit in bit lanes, mutating assignments to seek a 12-input model and rechecking clauses. The SAT path has no UNSAT route; a narrow two-input branch uses equivalence checks and RUP case proofs with CDCL fallback, emitting a proof stream on closure, while unsupported layouts or unfinished proofs return UNKNOWN. DPR

21.1 ×

18.7 ×

4/4

Packed bit-parallel evaluation combines an incumbent with one-bit neighbors; coordinate ascent and sparse mutations seek the asserted root, with large inputs falling back to random assignments. After validation it emits a SAT model but no DPR certificate; heuristic failure has no UNSAT or timeout result, while binary-root and unrecognized layouts return UNKNOWN. VeriPB

9.89 ×

9.44 ×

4/4

Weighted breakout search scores packed assignments and one-bit neighbors, raises weights on false root targets, and uses random moves to seek a satisfying 12-way output without handling a negative root. It reconstructs gates and checks clauses before emitting the model; no UNSAT certificate path is active, and unsupported shapes or other roots return UNKNOWN.

cril-misc

037

Proof format

No verif.

+ verif.

Solved

GRAT

0.961 ×

0.961 ×

74/172

Recognized hash layouts enumerate a bounded free message prefix, evaluate the hash, and use CDCL to complete auxiliary variables; other formulas use watched-literal CDCL with 1-UIP learning and restarts. Models are checked and returned; direct search lacks an UNSAT certificate and reports UNKNOWN on exhaustion or completion failure, while fallback logging supplies DRAT for elaboration. DPR

1.01 ×

1.01 ×

80/172

Exact structural recognition drives free-message enumeration and native hash evaluation, followed by CDCL completion of auxiliary variables; unrecognized inputs use watched-literal CDCL with 1-UIP learning and restarts. Verified SAT assignments are returned, while only the fallback can log learned DPR/RUP-style additions and a final empty clause; direct-search failure or unavailable logging yields UNKNOWN. VeriPB

0.896 ×

0.896 ×

66/172

Functional recovery of ordered gate blocks and 64-way bit-sliced input search drives the specialized path, filtering derived values against fixed clauses. SAT assignments are checked against original clauses; only compact CDCL emits a VeriPB RUP proof for UNSAT, while unsupported structures, fast-path exhaustion, and non-SAT sampled cases return UNKNOWN.

62 GBD instances · 13 validation · baseline 4/13

The benchmark asks whether Boolean signal and gate-output variables can satisfy a CNF encoding of a circuit or netlist. Clauses encode local gate relations, which may include parity-like constraints, gate definitions, or equivalences. Proof format

No verif.

+ verif.

Solved

GRAT

0.848 ×

1.04 ×

1/13

Structural recognition limits this solver to a large, clause-dense circuit-CNF envelope; parity-like blocks remain ordinary clauses, not GF(2)-eliminated. Watched-literal first-UIP CDCL handles recognized inputs; unsupported or unresolved cases return UNKNOWN, SAT assignments are checked, and UNSAT learned clauses are emitted as raw DRAT additions for external elaboration. DPR

0.799 ×

0.977 ×

0/13

It recognizes ordered ternary parity blocks and NAND definitions, then uses bipartite matching to orient parity equations before branching with structure-informed preferences in watched-literal CDCL. Unsupported or unmatched cases return UNKNOWN; SAT models are checked, while UNSAT emits witness-free DPR clause additions, including an empty clause, but the source does not establish proof correctness. VeriPB

0.799 ×

0.977 ×

0/13

Order-sensitive XOR/NAND template recognition selects probe variables, followed by hidden-unit extraction and failed-literal probing before watched-literal first-UIP CDCL. UNSAT traces use RUP steps and a VeriPB conclusion, SAT models are fully assigned and rechecked, and input or post-check failures return UNKNOWN.

cryptography-ascon

039

26 GBD instances · 6 validation · baseline 5/6

These formulas ask whether two unknown message bytes can make a fixed Ascon-Hash v1.2 digest through a bit-blasted Boolean circuit. The benchmark uses a specific optimized, topologically ordered CNF encoding with auxiliary variables and fixed constraints, rather than arbitrary cryptographic formulas. Proof format

No verif.

+ verif.

Solved

GRAT

1,010 ×

1.55 ×

6/6

Bit-parallel enumeration evaluates all 16-bit source assignments through the recognized circuit and intersects the target constraints. A surviving assignment is completed and checked clause by clause; exhaustion emits textual DRAT additions for external elaboration, including generated shortcuts and a resolution tree, while unsupported shapes or failed checks return UNKNOWN. DPR

747 ×

1.50 ×

6/6

Bit-sliced truth-table evaluation exhaustively tests 16-bit source assignments, using small gates and frontier summaries to filter target constraints. A surviving input is completed and checked against every clause; otherwise the solver emits witness-free assignment-blocking clauses and a resolution sequence for a DPR certificate, with recognition or certificate failures returning UNKNOWN. VeriPB

1.17 ×

0.400 ×

4/6

Incidence decomposition and bit-parallel compiled-gate evaluation enumerate the 16-bit source assignments while separating four small side components. A surviving lane is reconstructed with recursive DPLL and fully clause-checked; with none, the solver emits assignment-blocking RUP leaves and a balanced VeriPB polynomial-elimination tree, while unsupported structures or proof failures return UNKNOWN.

The Case for Automated Hyperspecialization: Evidence from SAT

cryptography-cbmc

040

design-debugging

042

13 GBD instances · 3 validation · baseline 3/3

119 GBD instances · 24 validation · baseline 24/24

These benchmarks ask whether Boolean constraints from a bounded cryptographic circuit check, typically a CBMC AES encoding, have a satisfying assignment. Circuit inputs and intermediate values become Boolean variables, while gate equations, bookkeeping conditions, and an assertion or miter obligation become CNF clauses.

These benchmarks ask whether fixed stimuli and, when present, state or observed values are consistent with a combinational or unrolled sequential circuit. Tseitin-style CNF clauses constrain Boolean signals, while unit clauses fix selected values; SAT means a completion exists and UNSAT means inconsistency.

Proof format

No verif.

+ verif.

Solved

Proof format

No verif.

+ verif.

Solved

GRAT

2.72 ×

8.95 ×

3/3

GRAT

3.04 ×

1.93 ×

24/24

Root propagation, binary-equivalence substitution, bounded variable elimination, and failed-literal probing simplify recognized CNFs before activity-guided CDCL with restarts and clause reduction. It checks SAT assignments against the original clauses and logs preprocessing and search additions for externally checked UNSAT proofs; unsupported inputs and some cases after elimination return UNKNOWN. DPR

0.419 ×

1.36 ×

3/3

Signature filtering identifies Tseitin-like inputs, after which root simplification, implication-SCC substitution, bounded elimination, and failed-literal probing precede focused output-assumption search for a unique width-16, 56, or 184 candidate, with generic CDCL as fallback. Logged witness-free clause additions and empty-clause exits support UNSAT, while SAT assignments are reported; unsupported inputs or unresolved post-elimination searches return UNKNOWN. VeriPB

0.159 ×

0.654 ×

2/3

Lookup-map congruence, signed AND/NAND matching, and binary-implication equivalences drive clause rewriting, bounded elimination, and failed-literal probing before watched-literal CDCL with 1-UIP learning. UNSAT receives VeriPB RUP records ending in an empty constraint; SAT is reported only when no variables were eliminated, while malformed inputs, setup failures, or post-elimination SAT cases return UNKNOWN.

cryptography-simon

041

71 GBD instances · 15 validation · baseline 5/15

These instances ask whether a 32-bit key maps fixed plaintext to ciphertext through Simon-style Feistel rounds, with a related key obtained by bit permutation. The CNF bit-blasts keys, round states, and gate intermediates, and fixes both endpoint words. Proof format

No verif.

+ verif.

Solved

GRAT

1,110 ×

1,100 ×

15/15

Structured Hamming-weight-ordered enumeration exploits the related-key bit permutation after exact canonical recognition, with SIMD batches and an exhaustive key-search fallback. A found key is reconstructed and checked against every input clause before a SAT assignment is emitted; unsupported, exhausted, or failed-validation cases return UNKNOWN, with no UNSAT certificate path. DPR

624 ×

596 ×

15/15

AVX-512 exhaustive enumeration of the first key, with a last-round inverse filter, follows fixed-key probes and exact clause-multiset recognition. Each found key is expanded and checked against every clause before a SAT assignment is emitted; unsupported, failed-validation, or exhausted cases return UNKNOWN because no UNSAT certificate path is implemented. VeriPB

716 ×

679 ×

15/15

Strict ordered-template recognition is followed by word-level endpoint screening and exhaustive vectorized key enumeration, including a copied-half test before the final nonlinear round. Each surviving key is expanded and clauses are validated before a SAT assignment is emitted; rejected templates or keyless instances return UNKNOWN, with no UNSAT certificate path.

Uniform-phase completion precedes search, assigning DIMACS variables in fixed order and accepting SAT only after checking every original clause. Failed trials invoke first-UIP CDCL with activity-based branching and phase saving, logging learned clauses as a DRAT refutation; SAT paths return checked models, and malformed or inconclusive cases return UNKNOWN. DPR

4.43·10 −3 ×

0.0379 ×

22/24

Gate-pattern congruence closure recovers supported AND, XOR, and ITE relations and merges them with a DSU quotient after propagation and phase trials. Failed-literal probing can emit clauses toward a DPR contradiction, but unsupported residual encodings or inconclusive search return UNKNOWN; no general DPLL or CDCL fallback exists. VeriPB

8.85·10 −3 ×

0.0704 ×

23/24

Occurrence-counter propagation scans CSR literal occurrences to decrement residual counts, with CDCL fallback using 1-UIP learning, phase saving, and restarts. UNSAT produces a VeriPB RUP proof, or a conditional compact pol chain for qualifying root conflicts; SAT assignments are clause-checked, while malformed or unresolved cases return UNKNOWN.

diagnosis

043

130 GBD instances · 26 validation · baseline 19/26

The benchmark asks whether a finite-domain diagnosis model has a consistent assignment, with each object selecting one of six states across repeated scenarios. Observations and transition or relation constraints are compiled into CNF using one-hot choices and auxiliary variables. Proof format

No verif.

+ verif.

Solved

GRAT

0.586 ×

1.40 ×

12/26

It recognizes the six-state exact-one structure, performs binary-implication SCC substitution and bounded variable elimination, then solves accepted formulas with watched-literal CDCL and first-UIP learning. Unsupported inputs return UNKNOWN; SAT outputs are checked assignments, while UNSAT uses textual DRAT for external elaboration; preprocessing conflicts can lack a final empty clause. DPR

0.520 ×

1.22 ×

10/26

It dispatches on recognized disjoint six-state exact-one groups and uses watched-literal CDCL with first-UIP learning, implication-graph minimization, activity branching, and restarts. SAT returns a checked assignment; UNSAT uses logged additions and deletions ending in an empty clause, with recovery after the serialization guard, while unsupported or certificate failures return UNKNOWN. VeriPB

0.661 ×

1.39 ×

14/26

It uses binary-clause parity union-find to quotient variables and validate a SAT model from the reduced instance; otherwise inputs use watched-literal CDCL with first-UIP learning. Unsupported or unresolved cases return UNKNOWN; UNSAT emits VeriPB RUP additions with hints, deletions, and an empty-clause conclusion, but shortcut paths have no UNSAT certificate.

Harrison Green, Claire Le Goues, and Fraser Brown

dimacs-sorter

044

discrete-logarithm

046

32 GBD instances · 7 validation · baseline 7/7

20 GBD instances · 4 validation · baseline 1/4

These benchmarks ask whether one option can be selected from each finite-domain block, with auxiliary Boolean choices, so that all remaining constraints are satisfied. In CNF, each block uses an at-least-one clause and pairwise at-most-one clauses, followed by residual circuit or row constraints.

The benchmark asks whether a bounded exponent x satisfies a^x mod n = r. CNF encodings represent exponent bits and staged accumulator, conditional-product, and modular-reduction words, with clauses fixing the final residue; implementations recognize particular layouts.

Proof format

No verif.

+ verif.

Solved

GRAT

GRAT

0.0591 ×

0.0591 ×

3/7

Ordered one-hot and gate-DAG recognition is the distinguishing technique: recognized circuits are evaluated over domain choices and Boolean controls, while shuffled inputs receive only one-hot recognition. After an initial population pass, watched-literal CDCL supplies the fallback; failed search or checking returns UNKNOWN, with no UNSAT certificate path. DPR

1.31 ×

1.31 ×

7/7

Bit-parallel local search is the distinguishing technique: after recovering signed one-hot blocks and functional gates, it evaluates many candidate states at once and applies tabu reassignment with randomized restarts. Recognition is bounded; search is incomplete. SAT models are printed, while failures return UNKNOWN; no UNSAT or DPR certificate is emitted. VeriPB

2.50 ×

2.49 ×

7/7

Row abstraction and propagation distinguish this solver: recognized sorter rows become source-count constraints, narrowed by fixpoint intervals and bounded affine checks before branching on the row with fewest patterns. Leaves reconstruct gates and validate every clause; unsupported layouts or exhausted search return UNKNOWN, with no VeriPB UNSAT certificate path.

dining-philosophers

045

37 GBD instances · 8 validation · baseline 8/8

These benchmarks ask whether a bounded dining-philosophers transition system can reach a deadlocked ring at some time frame. CNF encodes Boolean state and action signals, transition circuitry, and deadlock indicators, with a clause requiring one indicator to be true. Proof format

No verif.

+ verif.

Solved

GRAT

0.573 ×

0.595 ×

8/8

After recognizing its gate/XOR structure, the solver orders branches on deadlock outputs by AND-cone size; a size heuristic changes only that order. It then uses watched-literal CDCL, checks SAT assignments, and logs UNSAT clauses as textual DRAT additions for external checking; unsupported inputs return UNKNOWN. DPR

0.255 ×

0.177 ×

8/8

It recovers paths from Tseitin structure, prioritizes temporal goals, and for the exact UNSAT shape adds action-coverage clauses before bounded Davis-Putnam elimination, followed by CDCL. Non-special recognized inputs use CDCL; SAT models are checked, while UNSAT logging uses witness-free DPR additions with preludes outside live RUP checks, and unrecognized inputs or search/check failures return UNKNOWN. VeriPB

0.115 ×

0.0576 ×

7/8

Bounded Davis-Putnam elimination first shrinks recognized formulas while recording data for reverse model reconstruction, then watched-literal CDCL handles the reduced instance. SAT assignments are reconstructed and checked against all original clauses; UNSAT preprocessing and learned clauses are logged as VeriPB RUP constraints with an empty constraint; unsupported inputs, reconstruction failures, or validation failures return UNKNOWN.

Proof format

No verif.

+ verif.

Solved

2.85·104 ×

7,850 ×

4/4

Prefix CDCL probes recover intermediate residues; GCD-based reconstruction recovers the modulus before baby-step/giant-step search finds an exponent. That exponent and the circuit state are assigned as roots, and occurrence-indexed propagation completes and checks clauses; unsupported or ambiguous recovery, search or completion failures return UNKNOWN, with no UNSAT certificate path implemented. DPR

4.09·104 ×

1,070 ×

4/4

Baby-step/giant-step search directly recovers a bounded exponent after structural recognition, then pins exponent, accumulator, and product words to reduce circuit completion to propagation. A watched-literal pass is followed by an active CDCL fallback and clause checking; unsupported layouts, noninvertible giant steps, and search or completion failures return UNKNOWN, with no UNSAT or DPR certificate path. VeriPB

3.55·104 ×

885 ×

4/4

Extended baby-step/giant-step search peels gcd factors and uses GCD-based modulus recovery to solve the bounded logarithm. Stage residues and quotient hints pin exponent and state bits; alternating forward and reverse propagation completes auxiliaries and checks clauses, while unsupported layouts or search/completion failures return UNKNOWN, with no general SAT fallback or UNSAT/VeriPB certificate path implemented.

The Case for Automated Hyperspecialization: Evidence from SAT

edge-matching

047

ensemble-computation

049

97 GBD instances · 20 validation · baseline 8/20

13 GBD instances · 3 validation · baseline 3/3

The benchmark asks whether a square grid can be filled using each tile once while colors on shared edges agree between neighbors. CNF uses one-hot placement and internal-edge groups, with clauses enforcing legal local orientations; recognized layouts are specific encodings, not the whole edge-matching family.

These benchmarks ask whether a bounded collection of set-union operations can generate required target subsets from basic elements. They encode operation choices and resulting memberships as structured Boolean circuits, often expressed as CNF with auxiliary gate variables.

Proof format

No verif.

+ verif.

Solved

Proof format

No verif.

+ verif.

Solved

GRAT

2.17 ×

2.16 ×

14/20

GRAT

0.152 ×

0.271 ×

1/3

Structural recognition targets compact and unshuffled historical layouts, using grid embedding, arc consistency, and depth-first edge synchronization, with watched propagation and false-first completion on the historical path. SAT models undergo full clause checking; only a root-level propagation conflict emits a one-step RUP/DRAT empty-clause addition, while unrecognized layouts and other failures return UNKNOWN. DPR

1.10 ×

1.10 ×

8/20

Compact tile-major recognition and local orientation enumeration drive a min-conflict constructor, followed by connected-frontier backtracking with edge propagation, MRV choices, capacity checks, and symmetry breaking. The fallback forcing and watched-literal search prints SAT only after complete clause validation; unsupported, unsatisfiable, or exhausted cases return UNKNOWN because no UNSAT certificate path is implemented. VeriPB

1.64 ×

1.61 ×

12/20

Generalized exact-cover search uses colored secondary edge columns, bounded local edge tuples, and minimum-active-row branching after recognizing the compact placement-and-color structure. An alternate path requires a recovered contiguous square gate layout, forces identity placement, and uses watched-literal propagation; complete models are clause-checked, while failures return UNKNOWN because no UNSAT certificate path is implemented.

edit-distance

048

20 GBD instances · 4 validation · baseline 0/4

The benchmark asks whether a signed complete graph can be partitioned into clusters so that disagreements with target edge signs stay within an encoded edit budget. Pair-equality variables represent co-membership, transitivity clauses enforce consistent clusters, and additional bounded-width clauses constrain the allowed edit count. Proof format

No verif.

+ verif.

Solved

GRAT

2.00 ×

2.00 ×

2/4

The solver first uses partition local search with singleton, all-in-one, and graph-derived starts, then vertex relocations and cluster merges. It fixes pair variables, uses watched-literal CDCL to complete and validate a SAT model, and returns UNKNOWN on missed partitions, unsupported encodings, or failed completion; no UNSAT certificate path is implemented. DPR

2.00 ×

2.00 ×

2/4

Positive-gain agglomeration, single-vertex descent, and correlation-clustering restarts generate fixed pair assignments, then watched-literal CDCL completes and checks a SAT model. On failure, induced-P3 packing and bounded minimum-fill Davis-Putnam elimination feed residual CDCL and emit proof records for external checking; recognition mismatches return UNKNOWN, while UNSAT requires empty-clause or conflict evidence, not heuristic failure. VeriPB

2.00 ×

2.00 ×

2/4

Multi-start cluster-editing search uses graph-based and random partitions, greedy vertex moves, merges, exchanges, and perturbations to find a budget-feasible partition. It fixes pair variables and runs watched-literal CDCL only on the budget suffix, checks every original clause, and returns UNKNOWN on unsupported encodings or failed search/completion; no UNSAT certificate path is implemented.

Template recognition identifies bounded direct or complement layouts and encodes candidate subsets as integer masks. Randomized disjoint-union and min-conflicts searches propose programs; assumptions on result bits are propagated and clause-checked, while unsupported inputs or failed search return UNKNOWN and no active UNSAT or DRAT emission path exists. DPR

0.101 ×

0.181 ×

0/3

Bounded structural recognition filters for the six-row, k-column indicator layout and uses the recovered block to seed branching phases and order. Watched-literal CDCL with first-UIP learning searches the CNF; SAT models are completed and checked, while unsupported inputs or unproved UNSAT outcomes return UNKNOWN because proof logging is disabled. VeriPB

0.152 ×

0.271 ×

1/3

Acyclic XOR/AND recognition groups rows into masks and extracts targets. A construction handles one-one cases; randomized common-subexpression elimination and OR closure repair handle others, with CDCL fallback on an abstract subset closure. Checked SAT witnesses are emitted; failures or unsupported cases return UNKNOWN, with no UNSAT or VeriPB certificate path.

equivalence-chain

050

19 GBD instances · 4 validation · baseline 0/4

This benchmark asks whether a Boolean assignment satisfies a structured CNF. Groups of four ternary clauses encode parity relations over three variables, while binary and other ternary clauses impose residual constraints, sometimes using auxiliary variables. Proof format

No verif.

+ verif.

Solved

GRAT

1.00 ×

1.00 ×

0/4

Affine parity extraction and bounded-weight meet-in-the-middle decoding drive this solver: it recognizes four-clause ternary parity groups, eliminates them over GF(2), and searches for a low-weight parity-check solution. It propagates candidates through residual constraints and checks SAT assignments; without general fallback or UNSAT certification, structural, decoding, or completion failure returns UNKNOWN. DPR

1.00 ×

1.00 ×

0/4

Bounded-weight syndrome decoding is its core: it reduces recognized XOR structure to a system and uses meet-in-the-middle search for support of weight at most seven. It reconstructs variables and applies DPLL to residuals; checked SAT assignments succeed, while unsupported structure or failed decoding/completion returns UNKNOWN without an UNSAT path. VeriPB

1.00 ×

1.00 ×

0/4

Affine propagation and randomized conflict-core learning drive its search: Gaussian elimination expresses variables as affine forms, then residual-clause decisions yield dependency-based affine nogoods. It checks reconstructed SAT assignments; unsupported structures or failed search return UNKNOWN, with no UNSAT refutation implemented.

Harrison Green, Claire Le Goues, and Fraser Brown

equivalence-chain-principle

051

erdos-discrepancy

053

4 GBD instances · 1 validation · baseline 0/1

21 GBD instances · 5 validation · baseline 1/5

These CNFs ask whether a chain of Boolean tables can satisfy fixed endpoint assignments while auxiliary matching choices transport truth values between adjacent tables. In the benchmark encoding, CNF clauses express exact-one matching choices and guarded implications between corresponding entries.

The benchmark asks whether a finite sequence of positive or negative ones can keep every prefix sum along each homogeneous progression d, 2d, and so on within absolute value three. Its CNF uses sign variables plus auxiliary state variables to encode these bounded-prefix conditions.

Proof format

No verif.

+ verif.

Solved

Proof format

No verif.

+ verif.

Solved

GRAT

1.00 ×

1.00 ×

0/1

GRAT

1.01 ×

1.01 ×

1/5

Bounded structural recognition activates a fiber-independence invariant, propagating selector-controlled equivalences across the chain and using a backward dependency pass from a contradictory endpoint fiber. For recognized instances, it emits a textual DRAT clause-addition stream ending in the empty clause for external elaboration; unsupported, malformed, or proof-generation failures return UNKNOWN, with no general SAT fallback. DPR

1.00 ×

0/1

1.00 ×

Bounded recognition handles k=3 patterns by recursively transporting guarded existential-row clauses and branching on one matching coordinate; it does not require every pairwise-negative exact-one clause. Recognized k=4 patterns invoke watched-literal CDCL with 1-UIP learning and log UNSAT clause additions; checked SAT assignments are supported, while unsupported cases or solver failures return UNKNOWN. VeriPB

1.00 ×

0/1

1.00 ×

Structural recognition dispatches accepted chains to either a pseudo-Boolean equivalence proof, using matching-tree cardinality inequalities that telescope across layers, or a guarded implication proof that propagates nonempty fibers through an n-ary path tree. It emits a VeriPB proof rather than performing SAT search; inputs outside the reconstructed layout, proof-generation failures, and output failures return UNKNOWN.

equivalence-checking

052

28 GBD instances · 6 validation · baseline 0/6

These benchmarks ask whether two acyclic Boolean networks produce the same output for every primary-input assignment. A Tseitin CNF introduces variables for internal signals and asserts a final mismatch output, so a satisfying assignment supplies a distinguishing input and unsatisfiability establishes equivalence. Proof format

No verif.

+ verif.

Solved

GRAT

1.00 ×

1.00 ×

0/6

Bit-parallel simulation proposes equal or complementary signals, which the solver verifies while decomposing negated mismatch goals through recognized gate polarities. An incremental watched-literal CDCL fallback handles unresolved goals and logs learned clauses plus an empty clause for external UNSAT checking; unsupported or incomplete searches return UNKNOWN, and no SAT-model output path is implemented. DPR

1.00 ×

1.00 ×

0/6

Recursive OR decomposition and LUT shortcuts try to force final-miter literals false before cone-local watched-literal CDCL handles the remainder. Its three-input four-clause check is generic rather than mux-specific; on UNSAT, local clauses are mapped back and logged with an empty clause for external checking, while SAT completion is checked and failures return UNKNOWN. VeriPB

1.00 ×

1.00 ×

Hash-consed complemented-edge DAG reconstruction turns recognized Tseitin gates into simplified Boolean structure, while simulation signatures only nominate candidate equivalences and constants. Watched-literal RUP checks and bounded branching elaborate these claims; targeted CDCL emits RUP proof steps for closed UNSAT cases, but unresolved residues return UNKNOWN.

0/6

Primary-boundary recognition guides activity and phase initialization for the restricted three-state encoding, after which watched-literal CDCL performs the active search with learning, restarts, and clause reduction. SAT assignments receive full clause verification; rejected inputs, search or verification failure, and an UNSAT result return UNKNOWN, with no UNSAT certificate path. DPR

905 ×

80.2 ×

5/5

Mod-3 phase seeding, with ladder reconstruction for shuffled instances, targets recognized encodings before watched-literal CDCL completes or repairs assignments through learning and restarts. It verifies clauses for SAT; unsupported inputs, recognition failure, search failure, and an UNSAT result all return UNKNOWN, with no DPR proof path for UNSAT. VeriPB

0.815 ×

0.815 ×

0/5

A conditional family-specific phase warm start seeds sequence variables from a generated instance, then watched-literal CDCL searches recognized inputs with learning and restarts; only phases, not learned state, are carried over. It rebuilds states and rechecks clauses before SAT; recognition, reconstruction, or search failure and root contradiction return UNKNOWN, with no UNSAT certificate path.

fdmus

054

1,000 GBD instances · 200 validation · baseline 200/200

Given a Boolean circuit, the benchmark asks whether some input makes its output 1 and makes it 0 after either of two coordinate flips. CNF uses three circuit copies, sharing inputs except at those flips, with gate and output constraints. Proof format

No verif.

+ verif.

Solved

GRAT

3.22 ×

2.08 ×

200/200

Transitive-fanout equivalences collapse unaffected gates onto first-copy representatives and restrict decisions to changed inputs and affected cones. Watched-literal CDCL handles the reduced instance and SAT returns a complete assignment; exact-layout rejection or solver/proof failures return UNKNOWN, while UNSAT retains a textual DRAT/RUP-style certificate. DPR

3.03 ×

0.794 ×

200/200

Congruence propagation equates gates across copies; quotienting and bounded elimination simplify affected cones. Residual watched-literal CDCL searches the remainder and emits textual DPR additions for preprocessing and learning, with an empty clause on UNSAT. SAT reconstruction checks the full original assignment; unsupported layouts or solver, proof, or reconstruction failures return UNKNOWN. VeriPB

3.51 ×

0.352 ×

200/200

Structural hashing in lockstep with disjoint-set equalities derives gate congruences across the three copies, followed by root-level unit propagation. A two-watched-literal CDCL fallback handles remaining cases and logs equivalences and learned clauses as VeriPB RUP steps; SAT returns a checked full assignment, while unsupported layouts or solver/proof-writing failures return UNKNOWN.

The Case for Automated Hyperspecialization: Evidence from SAT

fermat

055

fixed-shape-random

057

8 GBD instances · 2 validation · baseline 2/2

30 GBD instances · 6 validation · baseline 2/6

This family asks whether two bounded, equal-width integers p and q satisfy p squared minus q squared equals N, usually with ordering and range conditions. A CNF encodes the squaring and subtraction circuits with auxiliary variables.

This benchmark asks whether shared Boolean base variables can satisfy all components of a bounded, fixed-shape CNF encoding. Each component uses center or root auxiliaries and selectors to express compatible local supports, and the question becomes whether one assignment satisfies every original clause.

Proof format

GRAT

No verif.

+ verif.

Solved

6.07·104 ×

1.48·104 ×

2/2

Template recognition recovers the odd target N rather than searching arbitrary CNF. It factors N, selects a balanced divisor pair for p and q, propagates through the circuit, and checks every clause before emitting a SAT assignment; unsupported or unsuccessful cases return UNKNOWN, with no UNSAT certificate path. DPR

4.25·104 ×

1,430 ×

2/2

Strict recognition of 8-65-bit primary vectors and a ripple-subtractor layout extracts the odd target N. Pollard-rho factorization and divisor splits construct p and q, after which ordered propagation and full clause checks validate a SAT assignment; unsupported or failed cases return UNKNOWN, with no UNSAT or DPR path. VeriPB

4.79·104 ×

1,300 ×

2/2

Parity and carry-chain recognition extracts an odd target from the narrow CNF layout rather than solving arbitrary formulas. It factors the target, selects a closest nontrivial factor pair, builds a full model by forward propagation, and checks every clause; failures return UNKNOWN, with no UNSAT or VeriPB proof path.

finite-state-machines

056

12 GBD instances · 3 validation · baseline 3/3

These benchmarks ask whether a bounded, unrolled finite-state transition system has an assignment satisfying its transition constraints and a stated property condition. State, signal, and auxiliary gate equations are converted to CNF, typically through Tseitin clauses. Proof format

No verif.

+ verif.

Solved

GRAT

0.822 ×

1.62 ×

3/3

Bounded variable elimination resolves pivots only when pair, width, and growth limits permit, retaining resolvents for reconstruction and logging. It then uses watched-literal CDCL with 1-UIP learning; SAT reverses elimination and checks a model, unsupported layouts return UNKNOWN, and UNSAT is written as DRAT additions for external elaboration and checking. DPR

0.478 ×

0.971 ×

3/3

It recognizes only TIP/PICO layouts; TIP prioritizes frames containing bad-state goals. On PICO, bounded Davis-Putnam elimination uses gate definitions and capped pairwise resolution before residual CDCL; SAT reverses elimination and checks a model, unsupported or setup-failing cases return UNKNOWN, while UNSAT ends with witness-free clause additions and an empty clause. VeriPB

1.24 ×

0.318 ×

3/3

Binary-implication SCC preprocessing merges equivalent literals, selecting bounded elimination regimes before a watched-literal CDCL fallback. SAT restores eliminated variables and checks a model; UNSAT logs preprocessing and learned clauses as VeriPB RUP constraints, ending with an empty RUP and UNSAT command, while malformed or failed cases return UNKNOWN.

Proof format

No verif.

+ verif.

Solved

GRAT

4.64 ×

4.64 ×

5/6

On recognized inputs, noisy minimum-break local search works on the sparse CNF and checks candidates before reporting SAT. A failed walk does not imply UNSAT; watched-literal first-UIP CDCL is the fallback, logging learned clauses and a final empty clause as textual DRAT for external GRAT elaboration, while unsupported inputs return UNKNOWN. DPR

4.55 ×

5/6

4.55 ×

Recognized inputs use focused probSAT on compact CNF, then support-level min-conflicts over shared bases. Base models get exhaustive auxiliary completion; if these searches fail, watched-literal CDCL logs learned clauses and the empty clause as witness-free DPR additions, and heuristic failure is not UNSAT. Unsupported shapes return UNKNOWN; no elaboration or kernel validation is supplied. VeriPB

4.62 ×

4.62 ×

5/6

Plan projection first reduces each recognized gadget to minimal signed-primary supports, then weighted min-conflicts flips all false literals of selected plans while updating affected gadgets incrementally. Exhaustive auxiliary reconstruction and full CNF checking certify SAT; no UNSAT or VeriPB certificate path exists, so unsupported inputs return UNKNOWN and unsatisfiable inputs may continue until externally stopped.

floodit-puzzle

058

40 GBD instances · 8 validation · baseline 8/8

These instances ask whether a fixed-source Flood-It board can be completely flooded within at most m moves by choosing colors that absorb adjacent monochromatic regions. The CNF uses one-hot color choices and time-indexed flooded-state variables. Proof format

GRAT

No verif.

+ verif.

Solved

745 ×

61.2 ×

8/8

Disjoint-set contraction and bitset state tracking support multi-heuristic greedy construction, with seeded noisy trials providing a fallback for finding a plan within the move bound. It recognizes only ten-color square grids; found plans are expanded into move and state variables and clause-checked before SAT, while construction failure returns UNKNOWN and no UNSAT certificate path exists. DPR

328 ×

8/8

5.71 ×

Monochromatic-region contraction supports greedy region-graph color selection with randomized trials as a fallback for finding a bounded move sequence. It reconstructs and clause-checks the full assignment before SAT; this narrow recognized encoding has no UNSAT certificate path, so recognition, search, or checking failure returns UNKNOWN. VeriPB

582 ×

6.51 ×

8/8

Randomized greedy rollouts on a contracted region graph lead into bounded beam search over flooded-component states. It recognizes only a ten-color layout, expands a found plan into all action and auxiliary variables, and rechecks every CNF clause before SAT; missed plans or failures return UNKNOWN, with no UNSAT certificate path.

Harrison Green, Claire Le Goues, and Fraser Brown

fpga-routing

059

genurq

061

97 GBD instances · 20 validation · baseline 16/20

12 GBD instances · 3 validation · baseline 3/3

These CNFs ask whether each route or vertex can select one of k channels so conflicting pairs receive different channels. They encode choices with group clauses and matched binary conflicts, with some instances using equality-expanded or binary representations.

Instances use small clause blocks with graph-like variable sharing; one form encodes parity constraints, while another combines positive binary clauses with an exact sequential encoding of an at-most-k constraint. The task is to determine whether all clauses can be satisfied for these restricted encodings.

Proof format

No verif.

+ verif.

Solved

GRAT

1.29 ×

1.04 ×

17/20

The active strategy reconstructs conflict graphs, uses bit-parallel clique search, and tries chordal or tabu coloring before finite-domain search and CDCL fallback. SAT assignments are checked against the CNF; UNSAT proof output supplies learned clauses or finite-domain nogoods for external DRAT elaboration, while unsupported or unfinished cases return UNKNOWN. DPR

4.07 ×

4.05 ×

19/20

The active strategy finds bitset (k+1)-cliques, then tries bounded coloring, binary equality chains, and tabu search before CDCL fallback. SAT assignments are checked against input clauses; UNSAT uses PR/PHP clique proofs or proof-logged CDCL with normalized/original fallback and witness-free additions. Unsupported inputs or unresolved searches return UNKNOWN; verification is external. VeriPB

4.09 ×

4.08 ×

19/20

The active strategy reconstructs route-conflict graphs and searches for bounded oversized cliques, using cutting-planes pigeonhole proofs when a clique is found; dense bit encodings can use one-hot extensions. Otherwise watched-literal CDCL supplies RUP logging, while SAT models are checked against original clauses and unsupported or failed cases return UNKNOWN.

generic-csp

060

83 GBD instances · 17 validation · baseline 8/17

Proof format

No verif.

+ verif.

Solved

GRAT

0.702 ×

0.793 ×

3/17

Bounded MAC first applies tuple-mask arc consistency, then the active CDCL search adds domain and functional clauses and uses conflict learning to seek a checked SAT assignment. UNSAT emits DRAT proof text for external elaboration and checking; unrecognized formulas or proof-output failure return UNKNOWN, and the local-search path is inactive. 0.892 ×

0.928 ×

6/17

Unique-extension propagation builds bootstrap seeds, then four-way DFS seeks checked SAT assignments and records rejected cubes. UNSAT is limited to at most 11 seeds and a complete bounded search with an externally checked clause/deletion log; otherwise unbounded local search is used, while unrecognized inputs return UNKNOWN. VeriPB

0.820 ×

No verif.

+ verif.

Solved

GRAT

64.9 ×

41.2 ×

3/3

Recognizes bounded-width parity blocks as XOR equations, solves their incidence graph by spanning forests, and enumerates one possible exceptional block. It verifies resulting assignments against the CNF before SAT, but has no UNSAT certificate path; unsupported or unresolved inputs return UNKNOWN instead of using general SAT search. DPR

56.4 ×

18.1 ×

3/3

Combines XOR block recognition and incidence-graph parity solving with a fallback for positive binary clauses and a sequential at-most-k counter. The base branch emits a checked SAT assignment; only bounded line-graph reconstruction can produce an UNSAT DPR proof, reducing the matching contradiction to PHP with PR swaps, while failed searches and unsupported shapes return UNKNOWN. VeriPB

72.7 ×

17.6 ×

3/3

Recognizes bounded XOR/table blocks, enumerates exceptional assignments, and solves incidence components by spanning trees. Compact SAT models are checked, while supported contradictions use VeriPB red/rup proofs or, after bounded line-graph reconstruction, pol proofs using a sequential counter; the debug route has no SAT branch and failures return UNKNOWN.

glassy-gen

062

140 GBD instances · 28 validation · baseline 9/28

Each instance asks whether a ternary constraint problem has one of four values for every variable while satisfying all relations. Its CNF encoding uses four-literal domain clauses and three-literal clauses forbidding disallowed value triples; the recognized subset has unique extension of any two values.

DPR

Proof format

0.678 ×

6/17

Support filtering and bounded CSP search branch on domains, then fall back to watched-literal CDCL with learned clauses when that search fails; SAT assignments are checked against the original CNF. UNSAT uses VeriPB text with red witnesses and RUP steps for external checking; recognition or proof setup/finalization failure returns UNKNOWN.

The benchmark asks whether a Boolean assignment satisfies every clause in a sparse near-14-regular CNF formula. It uses mostly three-literal clauses, at most one two-literal clause, and variables occurring 13 or 14 times, defining a structural envelope rather than the whole family. Proof format

No verif.

+ verif.

Solved

GRAT

1,340 ×

1,330 ×

28/28

Bit-packed GF(2) elimination treats the three-literal clauses as favored odd-parity equations, then recursive implication-closure repair tests the candidate; failure falls back to replica-exchange Metropolis search with local CNF repair and reseeding. Only clause-checked SAT assignments are emitted; unsupported inputs return UNKNOWN and no active UNSAT certificate path exists. DPR

4.35 ×

4.35 ×

24/28

Ordered planted-prefix reconstruction leads the search, followed by GF(2) parity elimination that tests both a solution and its complement; failures fall back to watched-literal 1-UIP CDCL and then open-ended probSAT repair. Only verified SAT assignments are emitted, while unsupported or unverified cases have no UNSAT/DPR path and return UNKNOWN. VeriPB

60.0 ×

60.0 ×

28/28

Ordered selector reconstruction is tried first, with candidate enumeration and CNF verification, followed by packed GF(2) elimination of odd-parity equations; remaining recognized cases use XOR annealing and focused probSAT-style search with Hamming-ball repair. Only verified SAT models are emitted; unsupported or unsuccessful searches return UNKNOWN, with no UNSAT/VeriPB certificate path exists.

The Case for Automated Hyperspecialization: Evidence from SAT

graceful-production

063

graph-isomorphism

065

20 GBD instances · 4 validation · baseline 4/4

92 GBD instances · 19 validation · baseline 14/19

The benchmark asks whether a structured CNF encoding of finite-domain order and equality constraints is satisfiable. Boolean variables represent semantic choices and auxiliary comparison, arithmetic, cardinality, or gate relations, while clauses enforce consistency in a restricted ordered construction.

The benchmark asks whether two colored graphs have a vertex bijection that preserves the relevant adjacency relations. In the supported CNF form, long clauses require candidate mapping choices, while binary clauses forbid incompatible choices. Proof format

No verif.

+ verif.

Solved

Proof format

No verif.

+ verif.

Solved

GRAT

1.79 ×

1.71 ×

6.44 ×

2.08 ×

4/4

17/19

GRAT

Two-watched-literal propagation drives a positive, increasing-variable completion after root units, avoiding branching or backtracking. It recognizes only a narrow normalized DIMACS signature; an independent clause scan verifies the selected-literal SAT witness, while root conflicts alone emit a minimal proof artifact and completion conflicts or unsupported inputs return UNKNOWN. DPR

0.0179 ×

2/4

0.0418 ×

Positive increasing-index construction, driven by watched-literal propagation, is the active technique rather than general SAT search or a graph-level decoder. It independently checks every original clause and emits a clause-covering SAT witness; conflicts, unsupported inputs, or oversized outputs return UNKNOWN, and no UNSAT or DPR certificate path is implemented. VeriPB

0.0179 ×

2/4

0.0410 ×

Parameterized opening-signature recognition selects the supported ordered CNF subset, followed by unit propagation and increasing positive assignment without DPLL or fallback search. A verified essential-variable witness is emitted for SAT; structurally different, conflicting, or oversized inputs return UNKNOWN, with no UNSAT or VeriPB certificate path.

After recognizing the restricted candidate-mapping grammar, it tries a bounded grid-factor shortcut with graph-isomorphism search, then falls back to CDCL with binary-conflict propagation and clause learning. SAT assignments are checked against the original CNF; UNSAT is emitted as textual DRAT for external elaboration, while unsupported inputs or failed checks return UNKNOWN. DPR

0.476 ×

0.721 ×

VeriPB

0.476 ×

0.721 ×

064

20 GBD instances · 4 validation · baseline 4/4

The benchmark asks whether choices can cover every required source while each destination is used at most once. CNF encodings use coverage clauses and binary incompatibility clauses; supported instances have one more source than destination, producing a pigeonhole contradiction. Proof format

GRAT

No verif.

+ verif.

Solved

2.50·104 ×

3,200 ×

4/4

Bounded dual-rail recovery and line-graph inversion reconstruct a bipartite incidence graph, accepting only a checked connected component with one more source than destination. A swap-and-remove reduction then emits ordered DRAT additions, including extensions, to the empty clause; recognition, parsing, or proof-output failure returns UNKNOWN, with no SAT search. DPR

1.17·104 ×

2.99 ×

4/4

VeriPB

3.17·104 ×

1.56·104 ×

4/4

Parity union-find collapses equality and complement classes, then line-graph reconstruction recovers endpoint cliques and the one-unit source surplus. The solver emits RUP clauses and incremental pseudo-Boolean at-most-one derivations for a counting proof, but performs no SAT search or external proof check; unsupported or failed cases return UNKNOWN.

066

35 GBD instances · 7 validation · baseline 6/7

This benchmark asks whether a cycle of w distinct binary words of width b exists, with each pair of cyclically adjacent words differing in exactly one bit. The CNF uses a binary matrix and auxiliary variables for shifted tracks, XOR relations, and one-bit transitions. Proof format

No verif.

+ verif.

Solved

GRAT

0.683 ×

0.866 ×

5/7

A bounded 64-bit exact-cover search over cyclic translates is the SAT technique, constructing a shifted-track cycle and checking the resulting assignment against all clauses. If construction fails, watched-literal CDCL is used; unsupported cases may return UNKNOWN, and UNSAT certificates are emitted as DRAT streams ending in empty clause. DPR

Signed union-find normalization, propagation, and clique checks recognize a one-over-capacity pigeonhole core, while matching exposes a Hall circuit. For accepted cores, it adds fresh variables and emits deletion and witness-free clause additions completing a polynomial pigeonhole proof. It does not run SAT search, and unsupported or failed cases return UNKNOWN.

0/19

It solves the recognized incidence structure as an independent-transversal problem, using bitset arc consistency and minimum-domain DFS with conflict-based support checks. After checking complete assignments against the CNF, it emits SAT models; unrecognized shapes and failed searches return UNKNOWN, with no active UNSAT certificate path.

gray_codes grandtour-puzzle

0/19

It reconstructs recognized graphs, using paired color refinement and individualization search, while a bounded CFI branch solves induced GF(2) constraints and falls back on this search. SAT mappings are CNF-checked; narrow UNSAT proof search uses Hall/pigeonhole branching and RUP/PR only for full instances without phase flips or repair cells, while other nonisomorphic cases return UNKNOWN.

0.673 ×

0.849 ×

5/7

A bounded residue construction is the primary SAT technique: it chooses periodic transition representatives and track offsets, then checks the assignment against every clause. If construction fails, CDCL is the fallback; unsupported cases return UNKNOWN, and root-conflict UNSAT logging records only witness-free learned-clause additions, ending in empty clause without PR witnesses. VeriPB

1,540 ×

100 ×

7/7

Bounded canonical cyclic-factor search is the active SAT technique: it enumerates derivative choices and shift lifts, builds words, and checks adjacency, uniqueness, and all clauses. Failure invokes no CDCL; counting and parity arguments can emit VeriPB proofs in eligible nondivisible or odd-quotient cases, while other recognized cases return UNKNOWN.

Harrison Green, Claire Le Goues, and Fraser Brown

greentao

067

hamiltonian

069

12 GBD instances · 3 validation · baseline 0/3

1,100 GBD instances · 220 validation · baseline 206/220

Green-Tao-style instances ask whether prime arithmetic-progression edges can be two-colored so every positive edge contains color 1 and every negative edge contains color 0. In CNF, these two requirements are monotone positive and negative clauses.

The benchmark asks whether a structured monotone CNF has a selection of Boolean variables that selects exactly one variable in each positive constraint. Pairwise negative clauses forbid selecting variables that co-occur, reducing the decision question to a restricted exact-cover search.

Proof format

No verif.

+ verif.

Solved

GRAT

1.00 ×

1.00 ×

0/3

Stochastic incremental local search repairs violated clauses with occurrence-local satisfaction counts, weighted break/make scores, restarts, and noisy flips, followed by a longer fallback. Checked assignments are certificates and produce SAT output; the inactive exact proof path leaves failed searches as UNKNOWN, with no UNSAT certificate. DPR

1.00 ×

1.00 ×

0/3

Min-conflicts search leads with negative-edge covers, balanced starts, and focused flips; failure guides an exact fallback using blocker reduction, counter propagation, watched clauses, and gain branching. No cutoff applies; it emits witness-free clause additions and an empty-clause UNSAT certificate. SAT has no proof file; unsupported inputs return UNKNOWN. VeriPB

1.00 ×

1.00 ×

0/3

With shorter positive clauses, asymmetric greedy hitting-set initialization covers negative edges while avoiding complete positive edges; stalled cases use Novelty+ local search with make-break scores, restarts, and incumbent retention. Checked assignments are printed as SAT certificates; no UNSAT or VeriPB certificate path exists, so unsupported and unsatisfiable inputs return UNKNOWN.

grs-fp-comm

Proof format

No verif.

+ verif.

Solved

GRAT

0.248 ×

0.256 ×

141/220

Focused Metropolis flips and at-most-one-preserving packing searches provide SAT attempts, followed by exactly-one CDCL probes. SAT assignments are checked; if no model is found, an unbounded logged CDCL run emits learned clauses and an empty clause as DRAT text for external elaboration and checking, while unsupported, exhausted, or failed paths return UNKNOWN. DPR

0.473 ×

0.483 ×

185/220

Reversible exact-cover search branches on the least-available constraint with reversible bitset buckets; bounded probes precede an unbounded pass. Watched-literal CDCL follows a no-model search; SAT assignments are checked, while UNSAT runs emit learned clauses and an empty clause as witness-free DPR additions for external checking; unsupported layouts or proof failures return UNKNOWN. VeriPB

0.443 ×

0.423 ×

182/220

CDCL-first search operates only after strict port-layout recognition, then falls back after a non-root conflict to Algorithm X with singleton propagation, reversible state, and minimum-row branching. SAT models are checked against the CNF; UNSAT reruns emit pseudo-Boolean RUP steps and tree nogoods, while unsupported layouts, some large cases, or proof-output failures return UNKNOWN.

068

64 GBD instances · 13 validation · baseline 1/13

hamiltonian-cycle

070

24 GBD instances · 5 validation · baseline 4/5

The benchmark asks whether a bit-blasted Boolean circuit for floating-point commutativity has a satisfying assignment, representing a counterexample. Its CNF uses unit constraints and ordered local clauses for two- or three-input gates; recognized layouts are restricted and do not independently verify arithmetic semantics.

The benchmark asks whether a sparse graph contains a single cycle visiting every vertex once. In the recognized CNF layouts, paired directed-edge selectors enforce one chosen incoming and outgoing arc per vertex, with auxiliary finite-state counter constraints.

Proof format

No verif.

+ verif.

Solved

0.955 ×

1.12 ×

0/13

Proof format

GRAT

Gate-aware root propagation and bounded variable elimination, with defined outputs avoided as decisions, simplify accepted streams before watched-literal first-UIP CDCL on the residual formula. For UNSAT, it emits plain clause additions; unsupported inputs or a conflict-free completion are reported as UNKNOWN rather than producing a SAT witness. DPR

0.955 ×

1.12 ×

0/13

Repeated failed-literal probing, root shortening, binary-implication SCC processing, and bounded Davis-Putnam elimination reduce the recognized circuit formula before 1-UIP CDCL on the remainder. Its UNSAT path buffers derived clause additions, including probe, SCC, elimination, learned, and contradiction clauses; unsupported inputs or a satisfying search result are reported as UNKNOWN. VeriPB

0.955 ×

1.12 ×

0/13

Dynamic-cost elimination over gate outputs and binary-implication SCC substitution drive preprocessing; fixed-seed local search seeds phases for one-UIP CDCL with minimization and restarts. For UNSAT, it emits derived and trimmed learned clauses as unhinted RUP steps, then an empty RUP constraint and conclusion; unsupported layouts or SAT exhaustion return UNKNOWN.

GRAT

No verif.

+ verif.

Solved

251 ×

233 ×

5/5

Graph CDCL drives the search with degree-two constraints, DSU component checks, and incremental subtour cuts. Blossom 2-factor seeds and randomized DFS provide fallbacks; a found cycle is oriented, completed by watched-literal propagation, and clause-checked, while unsupported or exhausted searches return UNKNOWN rather than an UNSAT certificate. DPR

1.43 ×

1.43 ×

4/5

Structural recognition reconstructs the graph, then rollback degree-2 search uses subtour propagation and frontier memoization on detected narrow strips. A Tutte-style degree-2 factor and bounded exchanges, followed by full-CNF CDCL, provide fallbacks; verified assignments are emitted, while unsupported layouts or failed SAT-side searches return UNKNOWN because no UNSAT/DPR certificate path exists. VeriPB

1.42 ×

1.41 ×

4/5

Certifying CDCL with VeriPB RUP replay supplies UNSAT certification after a graph-guided probe using blossom matching, Posa rotations, and bounded rollback. Only the supported layouts are recognized; SAT candidates are completed and clause-checked with full assignments, while unsupported inputs or replay divergence yield UNKNOWN and the counter-divisibility helpers remain unused.

The Case for Automated Hyperspecialization: Evidence from SAT

hanoi

071

hardware-model-checking

073

6 GBD instances · 2 validation · baseline 2/2

76 GBD instances · 15 validation · baseline 8/15

The benchmark asks whether a Boolean assignment describes a legal shortest Towers-of-Hanoi transfer between three pegs. A CNF uses move, disk, support, and state variables, with clauses enforcing one move per step, legal support changes, and the initial and goal configurations.

The benchmark asks whether Boolean signals in a bounded hardware check can receive values satisfying every constraint. Gate relations, initial conditions, and checked outputs become CNF clauses, often in a structured Tseitin netlist, though encodings may be less regular. Proof format

No verif.

+ verif.

Solved

Proof format

GRAT

0.567 ×

0.674 ×

1/15

GRAT

No verif.

+ verif.

Solved

159 ×

64.7 ×

2/2

Structural recognition recovers conditional-frame tracks, transition order, and endpoint towers for bounded Hanoi signatures. It recursively constructs the canonical plan, maps moves to action variables, propagates units before checking every clause; it has no SAT fallback, and structural, propagation, or verification failure returns UNKNOWN, with no UNSAT certificate path. DPR

4.57·10 −4 ×

4.63·10 −4 ×

0/2

Structural recognition accepts only a fixed six-disk, 63-transition structure, reconstructs its persistent tracks, and synthesizes the canonical plan. It reduces residual clauses to 2-SAT by propagation and SCC fallback, checks complete assignments against all clauses, and emits SAT only; failures return UNKNOWN, with no UNSAT certificate path. VeriPB

4.57·10 −4 ×

4.63·10 −4 ×

0/2

Recognition recovers the six-disk path and persistent move tracks, then forces each moved disk to the canonical optimal sequence. Occurrence-based unit completion handles larger frontends, while the base residual falls back to learned-clause CDCL. SAT assignments are clause-checked; unsupported or non-SAT cases return UNKNOWN, and no UNSAT certificate is emitted.

hardware-bmc

0.794 ×

0.964 ×

5/15

Conservative gate recognition and canonicalization recover Boolean operations, then test structural candidates under negation before adding RUP-style proof clauses. Within certifying size limits, watched-literal CDCL handles fallback cases; larger formulas with structural tasks use quotient discovery, while checked SAT assignments are accepted and proofless UNSAT or failed discovery return UNKNOWN. VeriPB

0.586 ×

0.730 ×

1/15

Complemented-edge structural hashing canonicalizes recognized topologically numbered Tseitin gates, folds simplifications, and checks the resulting quotient by propagation. After structural admission, internal watched-literal CDCL searches the quotient; without admission or a structural contradiction, the solver returns UNKNOWN, while SAT assignments are clause-checked and UNSAT ends with a VeriPB RUP certificate.

hardware-verification

This benchmark asks whether bounded hardware execution can reach a bad state while obeying initial and transition constraints. The time-unrolled circuit, state, and gate constraints become CNF clauses, so SAT gives a witness execution and UNSAT rules out the bound. Proof format

No verif.

+ verif.

Solved

GRAT

0.868 ×

1.13 ×

5/10

Structural gate-pattern recognition admits a narrow hardware-CNF subfamily, then bounded variable elimination simplifies it before watched-literal CDCL search. SAT assignments are reconstructed and checked against the original CNF; UNSAT runs emit DRAT-style clause additions, deletions, and an empty clause for external checking, while rejected inputs, proof-output failure, or failed checks return UNKNOWN. 0.829 ×

1.09 ×

5/10

Wide-gate recognition and incremental assumptions over bad frames drive search, with propagation and bounded elimination feeding watched-literal CDCL. Rejected frames become complement units; SAT models are checked after reverse extension, while UNSAT may log witness-free clause additions, deletions, and an empty clause for checking; unrecognized inputs, proof-output failure, or failed reconstruction return UNKNOWN. VeriPB

DPR

072

49 GBD instances · 10 validation · baseline 6/10

DPR

Exact AND/XOR recognition, equivalence probing, and bounded variable elimination simplify smaller gate-heavy instances before search. Watched-literal CDCL handles larger or unreduced formulas; SAT models are reconstructed and checked, while UNSAT clauses are logged for external elaboration and malformed or failed internal paths return UNKNOWN.

0.797 ×

0.786 ×

5/10

Signed equivalence closure and gate-congruence recognition enable proof-logged substitutions and bounded variable elimination before watched-literal CDCL search. SAT models are reconstructed and checked against the original CNF; supported UNSAT runs emit VeriPB RUP steps through contradiction, whereas UNSAT outside supported certificate profiles or on unrecognized inputs returns UNKNOWN.

074

779 GBD instances · 156 validation · baseline 129/156

The benchmark asks whether Boolean signals in a circuit or hardware-oriented encoding can satisfy all required constraints. CNF typically expresses gate relations together with output, property, or state conditions, though the runtime accepts arbitrary DIMACS formulas. Proof format

No verif.

+ verif.

Solved

GRAT

0.637 ×

0.965 ×

106/156

A bounded AIG sweep recognizes signed-AND gate patterns, simulates assignments in parallel, and may reconstruct a model or add equivalence clauses after RUP checks. Otherwise it falls back to CDCL. SAT models are checked; UNSAT logging records learned additions and a root empty clause as DRAT for external checking. Malformed input returns UNKNOWN. DPR

0.460 ×

0.685 ×

83/156

A size-band portfolio starts with baseline CDCL, then reparses selected inputs into a fresh solver with recursive reason minimization. It otherwise uses first-UIP learning, watched propagation, activity branching, and restarts. SAT assignments are not independently clause-checked; UNSAT output records learned and empty-clause additions without original-clause copies, and malformed input or replay/proof failures return UNKNOWN. VeriPB

0.526 ×

0.688 ×

97/156

A circuit sweep recovers gates, uses bit-parallel signatures and bounded implication queries, and falls back to CDCL when unproductive. The fallback uses first-UIP learning and Luby restarts; SAT assignments are clause-checked, while UNSAT output records learned clauses as VeriPB RUP constraints, an empty contradiction, and an explicit conclusion. Malformed input or certificate failures return UNKNOWN.

Harrison Green, Claire Le Goues, and Fraser Brown

hashtable-safety

075

heule-nol

077

21 GBD instances · 5 validation · baseline 1/5

11 GBD instances · 3 validation · baseline 2/3

These instances are CNF satisfiability queries with numbered Boolean variables and clauses. The available solver artifacts do not establish which hashtable-level property the queries encode.

The benchmark asks whether an L-shaped board can use three colors without a monochromatic equal-arm L, with some cells pinned. The CNF gives each cell three color variables, an at-least-one clause, and negative clauses forbidding monochromatic colors on each L.

Proof format

No verif.

+ verif.

Solved

GRAT

0.815 ×

1.10 ×

0/5

Packed binary implications are used for propagation in a narrow variable-count range, with ordinary watched literals as the fallback. It runs first-UIP CDCL with activity-based branching, phase saving, and restarts, checks SAT assignments, and emits learned clauses as DRAT for external elaboration; input, proof, or exception failures return UNKNOWN. DPR

0.815 ×

1.10 ×

0/5

A strict CNF profile gate admits only very large, mostly short-clause formulas with many units; accepted inputs use two-watch CDCL with 1-UIP learning. Checked SAT models are reported, while UNSAT output uses witness-free DPR clause additions, a final empty clause, and only the first parsed original unit; unsupported or output failures return UNKNOWN. VeriPB

0.815 ×

1.09 ×

0/5

A fixed signature gate accepts only a narrow normalized bit-blasted CNF pattern; accepted inputs use watched-literal CDCL with 1-UIP learning and activity-based branching. UNSAT conflicts are logged as VeriPB RUP constraints with hints and a final empty RUP, while SAT returns a checked assignment; unsupported or malformed inputs return UNKNOWN.

heule-folkman

076

Proof format

No verif.

+ verif.

Solved

GRAT

37.4 ×

37.4 ×

3/3

Weighted breakout descent uses incidence-local recoloring scores, tabu moves, and restarts to escape local minima. The narrow canonical-board recognizer validates a found coloring and emits it as the SAT witness; unsupported inputs return UNKNOWN, while unresolved recognized cases can remain in the restart loop because no UNSAT certificate path is implemented. DPR

13.4 ×

13.4 ×

3/3

Incremental breakout local search combines weighted incident scores, smoothing, and deterministic restart salts. The exact 22 by 22, n=11 recognizer searches one-hot colorings although its raw clauses lack at-most-one constraints; it emits checked SAT assignments, proves only a directly forced monochromatic L by an empty clause, and unresolved cases can remain in the restart loop. VeriPB

194 ×

191 ×

3/3

Incremental weighted min-conflicts local search evaluates six recolorings around a violated corner, using noise, breakout weights, and deterministic restarts. After exact recognition for n <= 64, unsupported inputs return UNKNOWN; VeriPB pol or RUP certificates cover only two restricted UNSAT triggers, while checked SAT assignments are emitted and other unresolved cases may continue indefinitely.

11 GBD instances · 3 validation · baseline 3/3

The benchmark asks whether the edges of a K4-free graph can receive two colors so every triangle is non-monochromatic. A CNF representation uses one variable per edge and an all-positive/all-negative clause pair per triangle, requiring its three values not to be all equal. Proof format

GRAT

No verif.

+ verif.

Solved

429 ×

419 ×

3/3

Focused local search attacks monochromatic triangles with weighted break-score flips and updates, then falls back to watched-literal CDCL when bounded search fails. SAT assignments are checked before emission; CDCL UNSAT logs textual DRAT and emits the empty clause after a root conflict for external elaboration, while unsupported inputs return UNKNOWN. DPR

1.83 ×

1.83 ×

3/3

NAE local search uses focused bad-triangle walks and breakout/probSAT choices, then tries vertex cut-switching, bounded repair, elimination, and CDCL. SAT candidates are checked on the original; reduced UNSAT is rerun untouched, logging additions that may be witness-free and emitting the empty clause only after a root conflict for external checking; unsupported inputs return UNKNOWN. VeriPB

104 ×

98.2 ×

3/3

NAE WalkSAT first maintains monochromatic triangles incrementally across noise restarts and an alternate sampler, then falls back to watched-literal CDCL seeded with its best phase and a color-symmetry fixing. A complete assignment is expanded and checked against the original clauses; unrecognized or exhausted cases return UNKNOWN because no UNSAT certificate path is implemented.

hgen

078

336 GBD instances · 68 validation · baseline 36/68

These benchmarks ask whether a Boolean assignment satisfies a DIMACS CNF, usually with bounded-width clauses and regular occurrence, density, and polarity patterns. Some encodings use disjoint four-literal choice clauses and signed binary exclusions whose hole components are one fewer than choice groups. Proof format

No verif.

+ verif.

Solved

GRAT

1.32 ×

1.30 ×

42/68

Configuration-checking and ProbSAT-style local search maintains true-literal counts and a false-clause set, then uses unbounded focused restarts when attempts find no model. An elimination route uses watched-literal CDCL with first-UIP learning and writes learned clauses as textual DRAT additions for external elaboration; the local-search route has no UNSAT result path. DPR

1.64 ×

1.53 ×

47/68

A structural sparse-pigeonhole detector first constructs a RUP/PR/RAT-style certificate with deletions and witness-free clause additions; the writer is not internally checked. Otherwise focused stochastic local search uses break counts, weighted choices, restarts, and an active WalkSAT-style mode. Regular recognized formulas have no UNSAT certificate route; failed assignments are reported UNKNOWN. VeriPB

1.61 ×

1.63 ×

47/68

A direct pigeonhole detector emits a VeriPB-style proof by deriving hole at-most-one constraints and summing them with four-choice clauses. Other CNFs use inverse-power local search with spectral restarts, then CDCL with RUP logging; only the exact pigeonhole structure has a direct proof path, while failed paths can return UNKNOWN.

The Case for Automated Hyperspecialization: Evidence from SAT

hidoku

079

independent-set

081

34 GBD instances · 7 validation · baseline 7/7

30 GBD instances · 6 validation · baseline 2/6

The benchmark asks whether an n by n grid can place 1 through n^2 exactly once, respect fixed clues, and put consecutive values in king-adjacent cells. CNF encodings use cell-value variables, exact-one constraints, clue units, and adjacency implications, sometimes with auxiliary variables.

The benchmark asks whether exactly K of N=2^d binary-word vertices can be selected without selecting both ends of a graph edge. CNF uses one selection variable per vertex, clauses forbidding both endpoints of each edge, and auxiliary one-hot running-count variables enforcing the exact total.

Proof format

No verif.

+ verif.

Solved

Proof format

No verif.

+ verif.

Solved

GRAT

0.0818 ×

0.202 ×

0/7

GRAT

1.17 ×

1.29 ×

2/6

It enumerates a bounded set of row-wise Hamiltonian snake assignments under board symmetries, fixing the primary cell-value variables before searching auxiliary variables. Watched-literal DPLL completes and independently checks the CNF before emitting a total SAT assignment; unrecognized or unsuccessful cases return UNKNOWN, and no UNSAT certificate path is implemented. DPR

163 ×

0.278 ×

7/7

It tries row-boustrophedon snakes on presumed primary variables, then uses propagation and recursive DPLL to complete and check the CNF before emitting a SAT assignment. A specialized recognizer constructs a clause/deletion trace for a diagonal crossing and reports UNSAT only on contradiction, but covers only narrowly recognized 6x6 and 7x7 encodings; other failures return UNKNOWN. VeriPB

8,980 ×

2,620 ×

7/7

It recognizes a nested-AMO and king-neighbor encoding and, for a geometric clue obstruction, emits a VeriPB certificate using PB and RUP derivations. Otherwise it tries symmetric row snakes, then clue-aware Hamiltonian-path search with distance and connectivity pruning; SAT models are checked before emission, while failed or unrecognized cases return UNKNOWN.

hypertree-decomposition

080

56 GBD instances · 12 validation · baseline 8/12

The benchmark asks whether a hypergraph has an elimination order whose resulting bags satisfy the hypertree-width bound, with each bag covered by at most k hyperedges and special conditions. A CNF records order, completion relations, bag selections, and bound constraints. Recognizers may support restricted encodings. Proof format

No verif.

+ verif.

Solved

GRAT

1.23 ×

1.23 ×

7/12

Recognizes a fixed CNF layout, constructs a primal elimination order with dual min-fill trials, then searches bag hyperedge covers, using bounded branching only for smaller widths. Rollback Horn propagation completes remaining variables, after which every clause is checked before SAT output; unsupported layouts, width rejection, or failed construction/search return UNKNOWN, with no UNSAT certificate path. DPR

1.23 ×

1.16 ×

7/12

Recognizes degree-two incidence and an exact dual-line-graph primal graph, then fixes a primal order by greedy dual minimum-fill completion. A watched-literal backtracking search solves remaining guard, special-condition, and counter variables, with limited seed retries; structural rejection or unsolved residual returns UNKNOWN, and no DPR derivation or UNSAT certificate path is implemented. VeriPB

0.517 ×

0.522 ×

0/12

Uses randomized local search over orders, with bitset branching for bounded bag covers and restarts. After an order is found, it fixes arcs and selectors, uses Horn propagation for remaining variables, and checks every clause before SAT output; unsupported layouts or failed search/completion return UNKNOWN, with no VeriPB or UNSAT certificate path.

Weighted-syndrome construction is tried first for Z-channel graphs, followed by specialized exact or bit-parallel complement-clique search and bounded local search on remaining recognized cases. Candidates are given counter values and fully rechecked before SAT output; failed or unsupported inputs return UNKNOWN, with no UNSAT or DRAT certificate path. DPR

1.17 ×

1.29 ×

2/6

Layered complement-graph clique search drives construction, using Hamming-weight layers and bit-parallel branch-and-bound for transposition graphs, while deterministic local search handles other recognized cases. After assigning the running-count auxiliaries, it rechecks every clause before emitting a SAT assignment; failed or unsupported cases return UNKNOWN, and no UNSAT or DPR certificate path is implemented. VeriPB

4.63 ×

3.62 ×

5/6

Component decomposition and complement-clique search solve suitable cases exactly, while residual-degree greedy and exchange search cover larger graphs. A found SAT assignment reconstructs the automaton and is checked against every clause; only narrow component-tree, transposition, and Z-channel branches emit VeriPB UNSAT proofs, while other failures or unsupported cases return UNKNOWN.

independent-set-reconfiguration

082

20 GBD instances · 4 validation · baseline 0/4

Given a graph and two same-size independent sets, the task asks whether one can reach the other through a specified-length sequence of distinct independent sets, replacing one vertex at each step. CNF encodings represent graph edges, configurations, transitions, and endpoint constraints. Proof format

No verif.

+ verif.

Solved

GRAT

2.00 ×

2.00 ×

2/4

Target-directed bitset exploration builds the token-jumping component, then randomized shortest-path-guided trials seek a simple path of the prescribed length. A completed path is propagated and checked against the CNF for SAT; bounded component, parity, and small-clique routines write DRAT-style UNSAT certificates, while unrecognized or unresolved cases return UNKNOWN. DPR

3.95 ×

3.90 ×

3/4

Guarded structural routines emit PHP/RUP-style derivations, including witness-free clause additions, for a few UNSAT cases; otherwise bidirectional BFS with bounded randomized walks and ear insertion seeks a simple path of the required length. Propagation and clause rescanning complete SAT models; unsupported layouts or unresolved searches return UNKNOWN, with no general UNSAT path. VeriPB

2.00 ×

1.99 ×

2/4

Specialized pseudo-Boolean proof generators run first, while bit-mask token-jump generation and component enumeration support distance-pruned exact-length DFS or seeded restarts. SAT paths are propagated and rescanned against the CNF; no generic CDCL proof fallback is enabled, so unrecognized or unresolved cases return UNKNOWN.

Harrison Green, Claire Le Goues, and Fraser Brown

influence-maximization

083

karatsuba-multiplication

085

20 GBD instances · 4 validation · baseline 0/4

3 GBD instances · 1 validation · baseline 1/1

The benchmark asks whether at most K seed nodes in a directed graph can activate at least T nodes within ten synchronous threshold rounds, with seeds remaining active. A CNF represents seed choices, activity states, threshold propagation, and cardinality limits.

These CNFs ask whether two odd, nontrivial binary factors satisfy a bit-blasted Karatsuba encoding of a fixed odd product. Factor bits and auxiliary circuit variables are linked by clauses, making the benchmark a Boolean satisfiability question.

Proof format

No verif.

+ verif.

Solved

Proof format

No verif.

+ verif.

Solved

GRAT

1.33 ×

1.33 ×

1/4

GRAT

0.115 ×

0.115 ×

0/1

Greedy seed construction and simulated-annealing swaps guide threshold diffusion on the recognized encoding. A watched-literal CDCL fallback handles unsuccessful searches without an internal stop condition; unresolved cases can return UNKNOWN. SAT checks a total assignment; UNSAT logs learned clauses and an empty clause as DRAT additions for external elaboration. DPR

1.33 ×

1.33 ×

1/4

Bit-parallel diffusion simulation drives greedy and annealed fixed-cardinality seed search on recognized layouts. If it fails, a watchdog-free CDCL fallback adds reduction, elimination, and learning clauses; unresolved or rejected models return UNKNOWN. SAT requires a total satisfying assignment; UNSAT requires potentially witness-free RUP additions to derive the empty clause. VeriPB

1.33 ×

1.33 ×

1/4

Composite odd-prefix probing guides modular-inverse recovery and watched-literal recursive DFS, prioritizing factor bits before auxiliary variables; clause learning, available prime tries, and factor-order symmetry prune the search. Only a checked satisfying assignment is emitted; no UNSAT certificate is implemented, so unsupported layouts, contradictions, and exhausted searches return UNKNOWN. DPR

0.115 ×

0/1

0.115 ×

Low-column residue decoding with modular-inverse propagation drives watched-literal recursive DFS, prioritizing factor bits before branching on remaining variables; prime and exact-width checks prune completed candidates. It emits a checked satisfying assignment but has no UNSAT certificate or generic fallback, so unsupported layouts, contradictions, and failed searches return UNKNOWN. VeriPB

0.115 ×

0/1

0.115 ×

Bitset threshold simulation scores greedy additions and annealed fixed-cardinality swaps on the recognized ten-layer layout. Activity-guided CDCL is conflict-bounded, so failed heuristic search or unfinished fallback returns UNKNOWN. SAT requires a checked assignment; UNSAT is reported only after level-zero closure with VeriPB RUP constraints and deletions.

Adaptive modular-residue recovery guides discrepancy-ordered factor search on recognized layouts, while bounded CDCL completes auxiliary variables and a general CDCL loop handles backdoor failure. Every reported assignment is checked against the original clauses; no UNSAT certificate path exists, so heuristic failure, contradictions, or unsuccessful completion return UNKNOWN.

interval-matching

knights-problem

084

23 GBD instances · 5 validation · baseline 0/5

086

21 GBD instances · 5 validation · baseline 2/5

The benchmark asks whether each group can choose one option without using a resource clique twice. Its CNF expresses choices with at-least-one clauses and conflicts with binary clauses, reducing the question to a matching that covers every group.

The benchmark asks whether an even square board has a closed knight tour, a cycle of legal knight moves that visits every square exactly once. CNF encodings select these moves, with auxiliary variables enforcing one-to-one incidence or position conditions.

Proof format

Proof format

No verif.

+ verif.

Solved

GRAT

3.34 ×

3.49 ×

4/5

GRAT

No verif.

+ verif.

Solved

2.97·104 ×

17.9 ×

5/5

Parity-DSU polarity recovery normalizes recognized clauses into rows and conflict cliques; augmenting-path matching constructs and checks a full assignment against every original clause. After a Hall deficiency, alternating reachability feeds a fresh-variable pigeonhole reduction emitting definition, row, conflict, and empty clauses for UNSAT; unsupported structure or proof failure returns UNKNOWN. DPR

1.26·105 ×

208 ×

5/5

Augmenting-path matching builds a complete assignment, checked against every original clause before SAT is reported. For a deficient matching, Hall validation and bitset-guided swap planning must succeed; the PR-style writer emits swap records, unit clauses, and an empty clause for UNSAT, while recognition, model-check, planning, or proof-output failures return UNKNOWN. VeriPB

1.50·105 ×

4,220 ×

5/5

Augmenting-path matching on recognized signed group/resource structure yields a complete assignment checked against every original clause before SAT is reported. For a deficient matching, alternating reachability supplies a Hall witness for a pseudo-Boolean derivation combining group clauses and clique bounds to derive an empty clause; recognition, model-check, or certificate failures return UNKNOWN.

Structural graph recovery is the distinguishing technique: recognized CNFs are matched to an even-board knight graph before a randomized Warnsdorff-style search selects a closed tour. Propagation and bounded completion fill the remaining variables and check all clauses, but failed recognition or search yields UNKNOWN; no general fallback or UNSAT certificate path is implemented. DPR

3.34 ×

3.48 ×

4/5

Structural encoding recognition recovers the knight graph and move groups, then randomized Warnsdorff search proposes a tour; propagation and bounded residual DPLL complete the assignment. For unary and binary formulas, SCC reasoning can certify UNSAT with contradictory unit clauses and an empty clause; unsupported encodings or failed search return UNKNOWN. VeriPB

1.67 ×

1.74 ×

3/5

Stable-color refinement and constrained backtracking recover knight-graph structure only for validated layouts, after which a Warnsdorff-style search selects a closed tour. Watched-literal propagation and bounded residual completion check the full clauses, but unsupported layouts or failed construction return UNKNOWN; no UNSAT certificate path is implemented.

The Case for Automated Hyperspecialization: Evidence from SAT

ktf

087

linvrinv

089

20 GBD instances · 4 validation · baseline 2/4

12 GBD instances · 3 validation · baseline 1/3

KTF instances ask whether a team can be selected from conflict groups so each skill is covered at least a required number of times within a weight budget. The CNF uses agent variables, skill counts, conflict cliques, and a budget constraint.

The benchmark asks whether n by n Boolean matrices over GF(2) can satisfy AB = I but not BA = I, testing one-sided inversion. The CNF uses matrix variables and AND/XOR auxiliaries, with a clause requiring BA to differ from I.

Proof format

No verif.

+ verif.

Solved

Proof format

No verif.

+ verif.

Solved

GRAT

0.511 ×

0.511 ×

0/4

GRAT

1.00 ×

1.00 ×

1/3

Bounded auxiliary-only elimination is followed by watched-literal CDCL, with learned clauses and positive-first branching. SAT reconstructs eliminated variables and checks the CNF; UNSAT at level zero is logged with clause additions and an empty clause for external DRAT elaboration, while unsupported input or internal failure returns UNKNOWN. DPR

1.02 ×

1.02 ×

2/4

It recognizes a narrow nested clique-and-skill encoding, compresses each clique to a level, and tries rarity-weighted greedy constructions. It completes auxiliaries and checks every clause for SAT; only demand two enables watched-literal CDCL with witness-free clause additions for UNSAT, while other failures or unrecognized layouts return UNKNOWN. VeriPB

1,440 ×

62.6 ×

4/4

An integer-dual search tests whether skill lower bounds exceed the decoded budget, emitting a pseudo-Boolean proof on success. Otherwise, randomized cost-aware construction and fail-first branching seek SAT models, complete auxiliaries by propagation, and check clauses; unsupported structures or exhausted search return UNKNOWN, with no general UNSAT fallback.

lam-discrete-geometry

088

A narrow structural signature drives watched-literal CDCL, preferring variables with the expected gate-occurrence pattern and otherwise using activity-based decisions. First-UIP learning, restarts, phase saving, and clause reduction guide search; UNSAT emits textual DRAT for external elaboration, while SAT returns a checked assignment and malformed, unsupported, or proof-file failures return UNKNOWN. DPR

1.00 ×

1.00 ×

1/3

XOR/AND flattening turns recognized networks up to four into parity clauses before two-watch CDCL, retaining the original CNF if flattening fails. First-UIP learning produces a textual DPR trace with parity, learned-clause, deletion, and empty-clause records on UNSAT; SAT returns a checked assignment, while malformed, unsupported, or larger inputs return UNKNOWN. VeriPB

1.00 ×

1.00 ×

1/3

Algebraic decoding reconstructs matrices and, at orders four and five, enumerates vectors for an injection-surjection VeriPB certificate. Other cases below six, or decoder failures, use watched-literal CDCL with First-UIP learning and RUP logging; malformed or unsupported inputs and orders six or larger return UNKNOWN, and SAT returns a checked assignment.

long-learned-clauses

090

20 GBD instances · 4 validation · baseline 3/4

16 GBD instances · 4 validation · baseline 0/4

These CNFs ask whether a partially fixed incidence structure can be completed by selecting incidence variables so required sets are hit while incompatible pairs and small forbidden patterns are avoided. They use restricted row-based encodings, sometimes with auxiliary intersection variables.

This benchmark asks whether a Boolean assignment satisfies every clause in a structured CNF formula. The clauses encode local relations and, in some instances, larger counting or graph-linked layouts.

Proof format

No verif.

+ verif.

Solved

GRAT

0.401 ×

0.619 ×

0/4

The supported pure 111-stride path uses a perfect-matching configuration lift, grouping incidence variables into row configurations and searching an exact-cover problem with bitset incompatibility masks and minimum-options branching. It has no generic fallback: unsupported structures return UNKNOWN, while UNSAT explanations are emitted as textual DRAT additions for external elaboration and checking. DPR

0.401 ×

0.619 ×

0/4

A domain-level search enumerates legal row domains for the compact stride-111 layout, propagating compatibility with bit masks and chronological branching. If the specialized path does not establish UNSAT, watched-literal CDCL with learned clauses and restarts is used; unsupported signatures return UNKNOWN, and UNSAT output is a witness-free clause-addition stream for external elaboration and checking. VeriPB

0.401 ×

0.619 ×

0/4

Semantic failed-literal probing and incidence-specific preprocessing prepare accepted instances for two-watched-literal CDCL, using first-UIP minimization, semantic phases, and Luby restarts. The bounded detector has no generic fallback: rejected inputs return UNKNOWN, while UNSAT attempts produce a VeriPB-style RUP proof for external checking and proof-creation failure also returns UNKNOWN.

Proof format

No verif.

+ verif.

Solved

GRAT

4.00 ×

3.96 ×

3/4

Small XOR truth-table blocks are grouped into incidence components; an odd-RHS, zero-incidence component triggers fresh XOR gates, resolution projections, and balanced-tree rotations for a refutation. The refutation is written as DRAT for external elaboration and checking; otherwise bounded randomized search checks models, and unsupported or failed cases return UNKNOWN. DPR

4.00 ×

3.95 ×

3/4

Small clauses become 2-4-variable parity relations; sparse XOR expressions and greedy overlap cancellation refute the odd-aggregate, two-occurrence case with witness-free DPR additions for external elaboration and checking. Otherwise, bipartite flow handles a relaxed exact-two four-variable pattern with one binary relation and builds a clause-checked model; unsupported cases return UNKNOWN. VeriPB

1.59·105 ×

1.20·104 ×

4/4

Sequential-counter and clique-covered graph layouts are recognized first, yielding a non-branching pseudo-Boolean proof of contradiction. Compact groups use parity tests or fewest-mask branching for a clause-checked model; structural proofs use only recognized clauses. A bounded TreeRup fallback can emit a RUP proof, and unrecognized or uncertified paths return UNKNOWN.

Harrison Green, Claire Le Goues, and Fraser Brown

matrix-multiplication

091

maxsat-optimum

093

53 GBD instances · 11 validation · baseline 7/11

60 GBD instances · 12 validation · baseline 9/12

These benchmarks ask whether a matrix-multiplication tensor over GF(2) can be expressed as a bounded sum of rank-one tensors. CNF variables encode factor coefficients and auxiliary products, while parity constraints require the resulting tensor entries to match.

The benchmark asks whether a Boolean assignment satisfies hard clauses while using no more than a specified number of soft-constraint relaxations. CNF encodings may express this bound with counters or structured choice, conflict, coverage, or selector gadgets, so recognized forms are restricted rather than universal.

Proof format

No verif.

+ verif.

Solved

GRAT

0.978 ×

0.987 ×

7/11

Structural recognizers recover tensor factors from AND/parity connectivity and try rank-six proof bridges or the embedded rank-23 construction, completing auxiliary variables by propagation before falling back to watched-literal CDCL. SAT models are clause-checked; UNSAT proof output is textual DRAT, with rank-six branches remapping bundled binary proof assets. DPR

0.373 ×

0.382 ×

0/11

An exact recognizer for the checked rank-23 Brent structure canonicalizes the GF(2) tensor, searches the embedded Laderman orbit, and propagates auxiliaries. Only nine compact signatures reach watched-literal CDCL; SAT assignments are checked, compact UNSAT logs witness-free DPR additions and an empty clause, while unsupported inputs or proof failures return UNKNOWN. VeriPB

0.937 ×

0.957 ×

7/11

Exact structural recognizers recover the 3x3 rank-23 Laderman structure or a rank-10 parity chain, using GF(2) elimination for residuals and a basis model for rank 10. Eligible cases use watched-literal CDCL with RUP logging; direct SAT returns checked assignments, while hard-listed or oversized cases return UNKNOWN without an UNSAT certificate.

maximum-constraint-partition

092

Proof format

No verif.

+ verif.

Solved

GRAT

0.348 ×

0.373 ×

0/12

Structured order-counter and selector recognition activates destroy/rebuild, crossover, and weighted neighborhood search, with a fresh SAT solve validating each extended candidate. Unrecognized inputs use watched-literal CDCL with textual DRAT-style logging; specialized branches emit SAT models but lack an UNSAT certificate and may return UNKNOWN or continue indefinitely. DPR

0.348 ×

0.373 ×

0/12

Exact-schema recognition activates least-model Horn closure and support-counted local search over choices, emitting checked SAT models. If that path fails or is inapplicable, watched-literal CDCL solves the original CNF and logs learned clauses as witness-free DPR additions. It tests only K, not a global optimum; exhaustion may return UNKNOWN, and immediate-conflict UNSAT lacks proof logging. VeriPB

0.407 ×

0.436 ×

2/12

Focused clause-weighted local search handles synthesis-shaped CNFs by flipping variables from false clauses and increasing weights on trapped clauses, returning only checked SAT models and no UNSAT proof path. Other recognized forms undergo subsumption, bounded variable elimination, and watched-literal CDCL with reconstructed assignments and VeriPB/RUP UNSAT logging; unsupported inputs or configured limits can yield UNKNOWN.

10 GBD instances · 2 validation · baseline 2/2

The benchmark asks whether 100 of 200 weighted items can be chosen so their total equals that of the remaining 100. Its CNF uses selector variables plus auxiliary cardinality and binary-sum circuitry, with equality outputs enforcing matching totals. Proof format

No verif.

+ verif.

Solved

GRAT

0.0797 ×

0.0798 ×

1/2

Bit-parallel evaluation of singleton scenarios recovers weights and total from the recognized 200-selector circuit, then dispatches by parity. Odd totals use least-significant-bit Davis-Putnam elimination and can emit DRAT additions for elaboration and checking; even totals use cardinality-grouped meet-in-the-middle search, but have no UNSAT path; unsupported layouts or failures return UNKNOWN. DPR

1,060 ×

74.5 ×

2/2

Bit-parallel singleton evaluation recovers weights and total from the recognized circuit, then dispatches by parity. Odd totals use a least-significant-bit slice with minimum-occurrence-product Davis-Putnam elimination and witness-free clause additions for UNSAT; even totals use cardinality-grouped meet-in-the-middle search for checked models, while unsupported layouts or failures return UNKNOWN. VeriPB

1,610 ×

70.9 ×

2/2

Scenario evaluation recovers weights and complementary sums from the recognized circuit, then dispatches by parity. Odd totals use affine GF(2) cone analysis to emit a pseudo-Boolean proof with RUP contradiction; even totals use greedy balancing and bounded exchanges for checked models but have no UNSAT path, while unsupported structures or failures return UNKNOWN.

md5-equivalence-checking

094

20 GBD instances · 4 validation · baseline 4/4

The benchmark asks whether a structured Boolean computation can be assigned values that satisfy its circuit constraints and a terminal comparison. CNF uses circuit and state variables with gate and control clauses, but recognized instances follow exact layouts rather than arbitrary encodings. Proof format

No verif.

+ verif.

Solved

GRAT

0.113 ×

0.218 ×

3/4

Occurrence-based unit propagation drives a DIMACS-order lucky-model construction, setting remaining encountered literals false and retrying with reversed order and phase choices within the bounded affine family. It checks every retained clause before emitting a complete satisfying assignment; malformed, unsupported, or failed attempts return UNKNOWN, and no UNSAT path is implemented. DPR

0.0382 ×

0.0739 ×

1/4

Fixed-point construction leads: occurrence-list propagation builds a base assignment, reuses its stable state across middle stages, and watched-literal DPLL searches the two terminal stages as fallback. It checks shifted clauses and terminal gates before emitting a complete assignment; failed or unsupported cases return UNKNOWN, with no UNSAT certificate path. VeriPB

0.0381 ×

0.0737 ×

1/4

Exact structural recognition recovers equality cones and variable relationships, then a fixed-point loop flips a target or free leaf while streaming gate evaluation. It handles only the default bounded recognized layout and checks every original clause before emitting a complete assignment; failures return UNKNOWN, with no UNSAT certificate path.

The Case for Automated Hyperspecialization: Evidence from SAT

mechanical-master-key

095

minimal-superpermutation

097

20 GBD instances · 4 validation · baseline 0/4

46 GBD instances · 10 validation · baseline 7/10

The benchmark asks whether keys can receive bounded-jump cut-depth words so every lock separates each unauthorized key from its authorized keys at some position, with CNF using one-hot key-depth variables, lock-depth variables, and separation witnesses.

The benchmark asks whether a fixed-length word over n symbols contains each required permutation as a contiguous substring. CNF encodes symbols at positions and links possible windows to occurrence, coverage, and cardinality constraints through auxiliary variables.

Proof format

No verif.

+ verif.

Solved

GRAT

5.27 ×

5.25 ×

4/4

Weighted local search assigns legal bounded-jump words to keys, scoring unauthorized pairs by authorized-depth counts and using whole-word moves, breakout weighting, and restarts. It accepts only the D=4, P=8 layout, checks the reconstructed SAT assignment, and returns UNKNOWN when recognition, search, or checking fails; no UNSAT certificate path is implemented. 3.59 ×

DPR

3.48 ×

3/4

Enumerated legal bounded-jump paths feed weighted local search, tracking authorization counts and violated pairs with whole-word moves and breakout weighting. It accepts only the exact CNF layout, checks reconstructed SAT assignments, and returns UNKNOWN for unsupported inputs or search failure; its proof fallback is disabled, so no UNSAT result is implemented. 1.71 ×

VeriPB

1.69 ×

2/4

Exact recognition gates sparse counterexample-repair, moving violating keys toward absent depths or unique authorized users with per-lock masks and counts. Other cases use legal words, whole-word RUP additions, and key-space CDCL with auxiliary elimination; SAT models are clause-checked, root conflicts yield UNSAT, while failures or unsupported inputs yield UNKNOWN.

minimal-disagreement-parity

096

Proof format

GRAT

No verif.

+ verif.

Solved

2.54·104 ×

57.1 ×

10/10

Hard-coded construction dispatch seeds candidate words, then whole-formula propagation and clause checking produce SAT assignments for supported layouts. If this fails on the recognized n=4 forbidden-canonical signature, occurrence-placement DFS emits RUP-style context clauses through a replayed recipe or semantic fallback; other or unverified cases return UNKNOWN. DPR

5.20·104 ×

54.8 ×

10/10

Fixed-witness seeding and indexed unit propagation test hard-coded n=4 and n=5 constructions, with bounded residual enumeration and a complete clause check before SAT output. Six exact n=4 variable-count and clause-structure cases use occurrence-window DPLL with copied or regenerated DPR occurrence-cover templates for UNSAT certificates; unsupported inputs or failed residual completion return UNKNOWN. 520 ×

VeriPB

80.8 ×

10/10

Coverage-first branching drives watched-literal completion after pinning shape-recognized formulas to canonical constructions, with chronological backtracking and full clause checks before SAT output. If completion fails, only n=4 cases enter proof-logging CDCL on uncovered permutation blocks, with VeriPB chains for UNSAT; n=5 failures and unrecognized or unresolved searches return UNKNOWN.

29 GBD instances · 6 validation · baseline 1/6

The benchmark asks whether shared Boolean variables produce parity outputs that disagree with observations within a bounded limit. Its CNF uses XOR or equality auxiliaries, a counter for the disagreement bound, and can include ordinary residual clauses. Proof format

GRAT

No verif.

+ verif.

Solved

2.65·105 ×

1.02·105 ×

6/6

Affine decoding through a weighted union-find and bit-parallel Gaussian elimination is the main strategy. It enumerates bounded error patterns, then completes auxiliary and weakened-chain variables with watched-literal search, verifies the full CNF, and prints SAT only then; unsupported recognition or failed decoding returns UNKNOWN, with no UNSAT certificate path. DPR

53.2 ×

53.1 ×

6/6

GF(2) elimination, min-fill factor decomposition, and separator/syndrome joins are the central search strategy. Junction-tree messages aid bounded reconstruction, and recovered assignments are back-substituted and checked against every input clause before SAT is printed; unsupported or failed paths return UNKNOWN with no UNSAT certificate path. VeriPB

2.96·104 ×

2.54·104 ×

6/6

Affine-mask propagation and Gaussian-elimination syndrome decoding drive the active search on recognized fixed-shape encodings. It enumerates bounded errors, checks decoded witnesses against the CNF, uses DPLL for a narrow weakened branch, and returns UNKNOWN for unsupported or failed cases; its VeriPB proof routine is not invoked for UNSAT.

minimum-disagreement-parity

098

30 GBD instances · 6 validation · baseline 0/6

The benchmark asks whether a binary source assignment induces affine parity outputs with at most K disagreement bits. It encodes the parity relations and Hamming-weight bound in CNF using auxiliary variables, while concrete encodings may vary. Proof format

No verif.

+ verif.

Solved

GRAT

5.39 ×

4.45 ×

5/6

Dual-syndrome meet-in-the-middle decoding uses GF(2) elimination to enumerate bounded supports and recover a checked assignment for the recognized canonical XOR-and-counter layout. If it fails, CDCL adds proof-only XOR consequences, emits DRAT for external checking, and reports UNSAT only after a root conflict; unsupported or inconclusive cases return UNKNOWN. DPR

3.00 ×

3.00 ×

4/6

Affine reconstruction recovers the map, then Gray-code meet-in-the-middle search indexes partial outputs and tests bounded Hamming neighborhoods for a source assignment. The candidate is propagated and checked against every clause before SAT output; restricted layouts, capacity limits, failed search, or no model yield UNKNOWN, with no UNSAT certificate path implemented. VeriPB

3.00 ×

3.00 ×

4/6

Lee-Brickell information-set decoding samples syndrome positions, solves binary systems, and tests low-error candidates with packed parity operations to recover a source assignment in the recognized affine-XOR/cardinality encoding. After validation, eligible small lower-bound cases use watched-literal CDCL with VeriPB RUP logging for UNSAT; bounded decoding failure or unsupported cases return UNKNOWN.

Harrison Green, Claire Le Goues, and Fraser Brown

misc-satex

099

modcircuits

101

19 GBD instances · 4 validation · baseline 3/4

20 GBD instances · 4 validation · baseline 1/4

The benchmark asks whether a Boolean assignment satisfies every clause of a CNF formula. Some instances encode one allowed choice per object with binary incompatibilities, while others encode Boolean gate relationships; these structures describe recognized subfamilies, not all CNF inputs.

These benchmarks ask whether a bounded acyclic Boolean circuit with two-input gates can realize a modular Boolean function, possibly with encoded input or output residues. CNF encodings represent gate sources, gate truth tables, row values, and outputs, with clauses enforcing consistency and the target behavior.

Proof format

No verif.

+ verif.

Solved

No verif.

+ verif.

Solved

8.83 ×

8.82 ×

4/4

Proof format

GRAT

GRAT

1.00 ×

1.00 ×

1/4

Graph-color recognition drives bounded TabuCol, while pure 3-CNF uses bounded ProbSAT with break counts; accepted models are checked against the original clauses. Fallback uses canonicalization, bounded elimination, and watched-literal CDCL, logging resolvents and learned clauses as DRAT for external elaboration; failed search or reconstruction yields UNKNOWN rather than UNSAT. DPR

11.6 ×

11.4 ×

4/4

Structural matching recovers shared color labels in recognized correspondence-coloring CNF, followed by bounded randomized TabuCol search and full-model checking. The narrow twoall shape instead uses bounded elimination and watched-literal CDCL with logged clauses for UNSAT; circuit and generic paths lack an UNSAT certificate path and may return UNKNOWN after failed search. VeriPB

0.781 ×

0.780 ×

3/4

Invariant-based recognition recovers list-coloring structure for TabuCol, while other admitted shapes use ProbSAT, min-break, or weighted-breakout search. Recognized AIG and completion cases may use bounded elimination with reverse extension, and heuristic failures fall to watched-literal CDCL with RUP logging; checked SAT assignments are accepted, but unmatched inputs return UNKNOWN.

miter

100

524 GBD instances · 105 validation · baseline 74/105

A Boolean miter CNF asks whether two Boolean circuits can produce different outputs on some input. Circuit wires and auxiliary variables are constrained by gate clauses, with clauses expressing the required output difference. Proof format

No verif.

+ verif.

Solved

GRAT

0.569 ×

0.632 ×

46/105

Signed-literal congruence is its main preprocessing: implication SCCs and recognized AND, XOR, or mux structures merge equivalent signals before bounded elimination. Failed-literal probing and CDCL search the residual; checked SAT models are reconstructed, while UNSAT is logged as textual DRAT for external elaboration and failures return UNKNOWN. DPR

0.379 ×

0.424 ×

15/105

Exact AND/XOR Tseitin recognition uses topological signed hashing and equivalence lemmas, with an all-AND unit-root path as a second recognizer. Residual cases use bounded elimination, profile-gated CDCL, or tree DPLL; checked SAT assignments are printed, while unsupported or exhausted cases return UNKNOWN and UNSAT uses witness-free RUP/DRUP-style additions. VeriPB

0.509 ×

0.567 ×

39/105

Parity union-find and bottom-up AND-cone hashing recover duplicate-wire equivalences, followed by bounded Davis-Putnam elimination, even though the input miter structure is not validated. Watched-literal CDCL handles the remainder with RUP-logged clauses; checked SAT models have no VeriPB SAT certificate, and proof, resource, or model failures return UNKNOWN.

Hard-coded recognizers install bit-parallel phase hints for selected MOD4, MOD5, and restricted MOD3 layouts, while otherwise retaining generic watched-literal CDCL search with exactly-one source-choice branching where applicable. SAT assignments are checked against the original clauses; UNSAT search logs learned clause additions and a terminal empty clause for external DRAT elaboration, and failed checks return UNKNOWN. DPR

0.750 ×

0.750 ×

0/4

Structural decoders derive hard-coded phase assignments for selected MOD4, MOD5, and restricted MOD3 layouts, then test constructions under temporary assumptions before CDCL fallback. Fallback uses implication-SCC substitution and bounded variable elimination, while first-UIP learning logs clause additions, including witness-free additions; checked SAT models return, but unsupported shapes or failed completion yield UNKNOWN. VeriPB

0.750 ×

0.750 ×

0/4

Embedded recipe matching decodes one-hot layouts and uses 64-bit truth-table simulation to assign gates, outputs, and row variables for selected parameter shapes. A polarity-shuffled ten-gate layout can trigger signed topology recovery with CDCL completion after construction fails; checked assignments are returned, while unsupported shapes or failed search yield UNKNOWN and no UNSAT certificate is implemented.

mosoi-289

102

36 GBD instances · 8 validation · baseline 4/8

Assign one of four colors to each grid cell so that no color occupies all four corners formed by two rows and two columns. A CNF uses four variables per cell, clauses restricting allowed colors, and four-literal negative clauses forbidding monochromatic rectangles. Proof format

No verif.

+ verif.

Solved

GRAT

1.00 ×

1.00 ×

4/8

Collision-scored min-conflicts search uses row-pair/color counts with legal recoloring and restarts. It checks candidate assignments against every parsed clause; bounded search failure returns UNKNOWN. Counting obstructions trigger a DRAT proof using a pigeonhole reduction from a usable 6x31 core; failed or unsupported cases return UNKNOWN. DPR

1.00 ×

1.00 ×

4/8

A specialized 6x30/30x6 matching construction is tried first; otherwise pseudorandom annealing minimizes monochromatic rectangles, with balanced column swaps preserving color counts in tight cases. Checked SAT assignments are emitted, while a generated 6x31 core can produce a DPR pigeonhole proof with fresh collision variables; unsupported sequential-counter/K4 cases return UNKNOWN. VeriPB

2.52·104 ×

788 ×

8/8

Incremental min-conflicts search on a recognized ordered or signed/permuted grid scores row-pair/color collisions and applies random perturbations. It verifies every input clause for each candidate before emitting a SAT assignment; resource-count contradictions produce a VeriPB proof, while the MIS fallback emits only a proof and unsupported inputs return UNKNOWN.

The Case for Automated Hyperspecialization: Evidence from SAT

multiplier-circuits

103

multiplier-verification

105

18 GBD instances · 4 validation · baseline 4/4

27 GBD instances · 6 validation · baseline 2/6

These benchmarks ask whether two bounded unsigned binary integers multiply to a fixed target. The target is encoded by unit clauses on product outputs, while Boolean gate clauses connect operand bits to a multiplier circuit and auxiliary variables.

Determine whether values for the free input variables can satisfy a structured acyclic Boolean mux network, a fixed-true terminal, and a final disjunction. The CNF defines each mux with six clauses and includes the network constraints plus the final clause.

Proof format

Proof format

No verif.

+ verif.

Solved

GRAT

1.05 ×

1.05 ×

2/6

GRAT

No verif.

+ verif.

Solved

547 ×

217 ×

4/4

Structural multiplier recognition recovers operand widths and the fixed product, then small-divisor and Pollard-Brent factorization finds width-fitting operands. The factors drive watched-literal propagation and a full clause scan for SAT; only a root-level unit-propagation conflict emits an empty-clause certificate, while recognition, factoring, or other unsatisfiable cases return UNKNOWN. DPR

713 ×

33.1 ×

4/4

Rigid-layout recognition and bit-parallel signature simulation recover a target from ordered encoding, then drive width-bounded divisor search with Pollard-Brent factoring. Candidate factors are checked against all original clauses before a SAT assignment is emitted; ambiguous signatures, unsupported layouts, or failed searches return UNKNOWN, with no UNSAT certificate path. VeriPB

853 ×

33.3 ×

4/4

Structural recognition reads product units and replaces SAT search with width-bounded factor search, using trial division, with Pollard-Brent available only up to 128 bits. It checks factors through the circuit and clauses before emitting a SAT model; unsupported or unfactored targets return UNKNOWN, with no UNSAT or VeriPB certificate path.

64-bit signature sweeping proposes RUP-validated unit and binary lemmas over the ordered ITE DAG. CDCL searches, tries final-literal assumptions, then falls back to whole-miter search; it emits DRAT-style clause additions and an empty clause for external elaboration, returns UNKNOWN on unsupported or incomplete cases, and has no SAT assignment output. DPR

0.702 ×

0.707 ×

VeriPB

1.05 ×

0.833 ×

104

20 GBD instances · 4 validation · baseline 0/4

The benchmark asks whether two circuit implementations, commonly 16-bit multipliers over two 16-bit operands, produce different outputs for some input. The CNF encodes the circuits and a final disagreement miter, so SAT gives a counterexample and UNSAT rules out disagreement for that encoding. Proof format

No verif.

+ verif.

Solved

GRAT

1.00 ×

1.00 ×

0/4

XOR/XNOR detection and bounded parity substitution strengthen recognized instances before watched-literal CDCL search with activity-based branching and clause learning. SAT is reported only with an assignment checked against the CNF; rejected patterns return UNKNOWN, while exhausted UNSAT search logs learned and strengthening clauses plus deletions and a final empty clause for external checking. DPR

1.00 ×

1.00 ×

0/4

Primary-input conflict projection uses an independent watched-literal RUP oracle to minimize learned clauses, driving asserting backjumps in watched-literal CDCL search. SAT models are reconstructed and checked; unsupported or failed cases return UNKNOWN, while UNSAT emits witness-free DPR additions, including learned operand-blocking clauses and a final empty clause for external elaboration and checking. VeriPB

1.00 ×

1.00 ×

0/4

Proof-first pseudo-Boolean generation recovers partial products and width-three compressor relations from a restricted Tseitin DAG, emits a refutation, and leaves proof checking external. If generation fails, bit-parallel evaluation searches for a counterexample and reports SAT only after clause checking; otherwise it returns UNKNOWN and has no UNSAT certificate path.

2/6

Mux-derived phase initialization guides watched-literal CDCL with first-UIP learning. It logs learned clauses with RUP and dependency hints and a final empty constraint for UNSAT; the SAT path checks and prints a model but emits no completed VeriPB certificate, while unsupported inputs or proof failures return UNKNOWN.

mutilated-chessboard multiplier-equivalence-checking

0/6

Gate-aware propagation stores each ITE as a six-clause object, rescanning affected gates while CDCL branches only on primary inputs. First-UIP learning with watched clauses supports UNSAT certification through witness-free learned clause additions and an empty clause; unsupported structure, duplicate final literals, or proof failures return UNKNOWN, while SAT prints an assignment.

106

22 GBD instances · 5 validation · baseline 2/5

The task asks whether dominoes can cover every remaining square after removing two same-color opposite corners from a checkerboard, without overlap. CNF encodings use placement variables, coverage clauses, and constraints preventing a square from being used twice. Proof format

No verif.

+ verif.

Solved

GRAT

1.53 ×

1.16 ×

3/5

Checkerboard matching recovery drives anti-diagonal frontier DP, using occupancy bit masks and a subset trie to reuse failed states. Preparation checks destination conflicts with binary propagation and emits bottom-up dependency lemmas, deletions, and an empty root clause for UNSAT; without a SAT witness path, satisfiable or unsupported cases return UNKNOWN. DPR

0.783 ×

0.812 ×

1/5

A strict structural recognizer accepts only validated direct-incidence or sparse-matrix pigeonhole layouts and returns UNKNOWN otherwise. Bounded Davis-Putnam elimination and watched-literal CDCL use first-UIP learning; the clause/deletion trace records resolvents, learned clauses, deletions, and an empty clause on a level-zero conflict, while non-UNSAT outcomes return UNKNOWN without model reconstruction. VeriPB

3.56·105 ×

1.33·104 ×

5/5

Matching recognition converts compact or sparse layouts into degree-2-to-4 counting constraints over mutexes, replacing general search. Unit propagation may justify missing mutexes in the sparse fallback; a VeriPB contradiction then combines coverage clauses with clique inequalities. Unsupported inputs or failed proof preconditions return UNKNOWN; no SAT witness path is implemented.

Harrison Green, Claire Le Goues, and Fraser Brown

oddball-weighing

107

ordering-principle

109

40 GBD instances · 8 validation · baseline 8/8

36 GBD instances · 8 validation · baseline 8/8

The benchmark asks whether each object can receive a nonzero ternary signature and orientation across non-adaptive weighings so every weighing has equal positive and negative counts. CNF encodings represent signature choices and directed matching edges pairing opposite outcomes.

The benchmark asks whether a directed relation on n objects can give every object a predecessor while obeying antisymmetry and transitivity. A CNF encoding uses variables for directed edges, predecessor clauses, binary clauses forbidding opposite edges, and ternary clauses expressing transitivity.

Proof format

Proof format

GRAT

No verif.

+ verif.

Solved

413 ×

36.9 ×

8/8

It recognizes a narrow matching-and-signature CNF pattern, then searches fixed canonical assignments with randomized-prefix meet-in-the-middle and free ordered subsets with codeword/orientation annealing. Successful constructions undergo streamed propagation and clause checks; an odd-row obstruction can produce textual DRAT for external elaboration, while other heuristic failures return UNKNOWN rather than UNSAT. DPR

359 ×

3.77 ×

8/8

On narrow canonical block/domain inputs, an odd-coordinate matching obstruction adds core clauses and an empty clause for UNSAT certification. Otherwise, exact dynamic programming gives way to annealing or distinct-tuple search; the SAT path lifts tuple and matching variables, propagates the CNF, and greedily completes clauses, while failures return UNKNOWN. VeriPB

0.692 ×

1.04 ×

7/8

On its recognized layouts, seeded annealing balances coordinate sums by changing orientations and, for free layouts, selected signature classes. Only fixed layouts with an odd support coordinate use the pseudo-Boolean matching proof path; other cases seek a model by streaming propagation after pairing opposite sides, with failures or oversized inputs returning UNKNOWN.

GRAT

No verif.

+ verif.

Solved

142 ×

11.7 ×

8/8

A bounded search for cycles among omitted transitivity clauses supplies checked SAT models when successful; a single missing predecessor or antisymmetry axiom also permits direct model construction. Otherwise, vertex elimination emits a textual DRAT clause sequence toward UNSAT; unsupported shapes or failed construction return UNKNOWN without general SAT search. DPR

138 ×

4.31 ×

VeriPB

120 ×

3.99 ×

108

10 GBD instances · 2 validation · baseline 2/2

The benchmark asks whether ternary Boolean XOR equations can be satisfied when each logical value is represented by the OR of a consecutive pair of variables. Each equation is converted to CNF by distributing over those pair expressions, producing the encoded instance. Proof format

No verif.

+ verif.

Solved

GRAT

0.281 ×

0.354 ×

2/2

Bit-packed GF(2) Gaussian elimination solves recognized systems and reconstructs a checked SAT assignment. Inconsistent cases use pair-symmetry reduction, bounded Davis-Putnam elimination, and watched-literal CDCL to build a textual DRAT refutation for external checking; only exact canonical OR-pair expansions are handled, and failures return UNKNOWN. DPR

0.434 ×

0.129 ×

2/2

Bit-packed GF(2) elimination first solves recognized equations and produces a clause-checked SAT assignment. For inconsistency, the solver peels a 2-core, applies pair-symmetry reduction and bounded Davis-Putnam elimination with witness-free clause additions, then uses CDCL to complete the refutation; only exact canonical OR-pair expansions are recognized, and failures return UNKNOWN. VeriPB

0.194 ×

0.275 ×

2/2

Dependency-carrying GF(2) elimination solves recognized systems and yields a clause-checked SAT assignment. For inconsistency, the solver peels and normalizes a 2-core, then uses watched-literal CDCL with RUP logging to emit a VeriPB proof; only the exact OR-pair encoding is supported, and unsupported inputs or failed proof construction return UNKNOWN.

8/8

Safe object elimination removes an object only when omitted transitivity clauses cannot obstruct it, then emits a VeriPB RUP proof of UNSAT. If elimination fails, it checks models only for a missing-transitivity cycle or one missing predecessor or asymmetry clause; unsupported or failed cases return UNKNOWN without general SAT search.

ordering-principle-xor or_randxor

8/8

Structural recognition recovers vertices, directed edges, and polarities, then validates the supported transitivity pattern. It either greedily eliminates vertices with witness-free RUP additions to derive UNSAT, or constructs and checks a model for narrowly supported missing-axiom or cycle cases; unsupported inputs and failed checks return UNKNOWN without general SAT search.

110

6 GBD instances · 2 validation · baseline 2/2

Given n vertices, can a transitive, antisymmetric relation give every vertex at least one of three prescribed predecessors? The supplied encodings represent each relation variable by an XOR pair and expand the resulting width-two and width-three clauses into CNF. Proof format

No verif.

+ verif.

Solved

GRAT

2,670 ×

15.8 ×

2/2

Candidate-state scanning applies the finite-order minimal-element contradiction, replacing a current candidate exactly when a relation holds. It bridges selected XOR definitions to logical clauses and emits a DRAT clause-addition refutation ending in the empty clause; unrecognized or output-failing cases return UNKNOWN rather than using general SAT search. DPR

8,640 ×

44.4 ×

2/2

Minimum-fill elimination orders the predecessor graph and gauge-fixes only XOR atoms needed by that plan. It emits witness-free DPR additions for bypass and reduced predecessor clauses, reporting UNSAT only after an empty clause is written; there is no SAT fallback, and unsupported or failed cases return UNKNOWN. VeriPB

2,710 ×

33.8 ×

2/2

Redundancy-based XOR symmetry breaking fixes each even physical bit before deterministic vertex elimination. It emits RUP predecessor clauses as candidate sets are unioned, then completes a VeriPB proof with a final NONE marker and UNSAT conclusion; failed recognition or proof generation returns UNKNOWN, and no SAT search is used.

The Case for Automated Hyperspecialization: Evidence from SAT

p-center

111

21 GBD instances · 5 validation · baseline 5/5

pebbling

113

55 GBD instances · 11 validation · baseline 11/11

The benchmark asks whether at most p candidate facilities can cover every demand point, with each demand represented by a positive facility clause. In CNF, auxiliary variables enforce the cardinality bound, while equivalent encodings may represent the same p-center question.

These instances ask whether a Boolean assignment can satisfy a DAG’s pebbling dependencies while forcing its sink false. CNF encodings use signed literal blocks, source clauses, Cartesian transition families, and sink units; some variants use a specific NAE gadget.

Proof format

No verif.

+ verif.

Solved

Proof format

No verif.

+ verif.

Solved

13.9 ×

0.292 ×

11/11

GRAT

3.03 ×

2.14 ×

5/5

GRAT

Failed-literal probing is the active strategy: root propagation tests variables and turns a conflicting polarity into an opposite unit until conflict or closure. It handles only a restricted layout with unit coverage clauses; no SAT path exists, so unclosed or SAT cases return UNKNOWN, while closed refutations record units and an empty clause. DPR

2.72 ×

0.416 ×

5/5

Facility-aware DPLL is the active strategy: watched propagation branches on a facility from the shortest uncovered cover clause, favoring frequent facilities, then uses auxiliary fallback decisions to complete the counter extension. It checks SAT assignments and emits postorder, witness-free RUP blockers for UNSAT; other encodings or failed checks return UNKNOWN. VeriPB

2.06 ×

0.663 ×

5/5

Specialized neighborhood branching is central: watched propagation selects a candidate from the tightest uncovered neighborhood, while counter clauses enforce capacity. Root propagation may produce an empty-RUP result; otherwise a CDCL fallback learns RUP clauses, while checked SAT assignments are emitted and unsupported layouts or failed checks return UNKNOWN.

parity-games

112

28 GBD instances · 6 validation · baseline 5/6

The benchmark uses structured CNF encodings of bounded sequences of parity-game strategies and their transitions. Satisfiability asks whether Boolean variables can realize a consistent encoded construction, using relation variables, transition clauses, and definitional witnesses. Proof format

No verif.

+ verif.

Solved

GRAT

1.05 ×

1.05 ×

5/6

Strict structural recognition gates two-watched-literal CDCL with 1-UIP learning, clause minimization, activity-based branching, and LBD retention. SAT models are checked; a non-SAT pass is rerun with logging, and only a logged UNSAT yields UNSAT, while nonmatching inputs or a failed rerun yield UNKNOWN. The proof path emits textual DRAT clauses. DPR

0.673 ×

0.673 ×

4/6

Binary-implication SCC compression, failed-literal probing, and bounded resolution elimination simplify recognized instances before CDCL; vivification is inactive. SAT assignments are reconstructed and checked against the original CNF, while UNSAT deductions are streamed as witness-free textual DPR additions; malformed or nonmatching inputs return UNKNOWN. VeriPB

0.673 ×

0.672 ×

4/6

Binary-implication SCC compression and bounded variable elimination precede CDCL with first-UIP learning. In SAT-first mode, it searches without proof retention and reruns only UNSAT results for version 3.0 pseudo-Boolean proof emission. SAT models are reconstructed and checked; no verifier is included, and parsing, checking, or proof-writing failures return UNKNOWN.

A clause-sharing recognizer factors transition families into Cartesian predecessor blocks and reconstructs the pebbling graph, with an UNSAT-only fallback for exact parity-gadget shapes. Staged resolution and RUP checks handle complete transitions and limited missing-cell repairs. CDCL can return a checked SAT assignment or DRAT trace; unsupported cases return UNKNOWN. DPR

109 ×

25.3 ×

11/11

Occurrence signatures and Cartesian-tail factorization recover a disjoint DAG, with a validated three-variable NEQ route using a fixed resolution template. Topological traversal emits resolution proofs, including repairs, while SAT construction handles only one recognized source, transition, or sink defect and checks its model. Unrecognized or failed constructions return UNKNOWN. VeriPB

30.0 ×

12.4 ×

11/11

NAE3 dispatch eliminates exact gadgets by local resolution; incidence recovery otherwise builds a Horn quotient from signed blocks. Forward chaining finds a conflict for proof or lifts a checked SAT model; only one-deletion OR cases are accepted. Unit-containing CDCL covers small recognized-width formulas; NAE SAT cases and unresolved inputs return UNKNOWN.

perfect-matching

114

17 GBD instances · 4 validation · baseline 0/4

Given a bipartite graph, can each of n+1 left vertices choose an incident right vertex without choosing any right vertex twice? The CNF requires one choice per left vertex and at-most-one use per right vertex, often exposing an unbalanced matching or pigeonhole core. Proof format

GRAT

No verif.

+ verif.

Solved

5.53·105 ×

3,080 ×

4/4

Structural recognition identifies an unbalanced graph-pigeonhole core by recovering consecutive row blocks and column exclusions from binary implication paths. It recursively eliminates rows and holes with fresh variables, using RAT definitions and RUP child rows to derive the empty clause; unsupported encodings, proof failures, and SAT cases return UNKNOWN. DPR

2.82·105 ×

1,840 ×

4/4

Hall-deficit recognition reconstructs the unbalanced matrix and checks binary implication reachability for column conflicts. It recursively removes a row and column, adding Tseitin definitions and witness-free conflict clauses, with RUP bridges when needed, until an empty clause; unsupported shapes, proof failures, and SAT inputs return UNKNOWN without SAT search or assignment output. VeriPB

3.26·105 ×

6.18·104 ×

4/4

Pigeonhole counting recognizes a row-major unbalanced core, recovering column at-most-one constraints from pairwise, sequential, grouped, or implication-based encodings. It writes a VeriPB proof using inequality chains or RUP-added conflicts followed by cutting-planes induction; full-CNF propagation is a fallback, while unrecognized or satisfiable inputs return UNKNOWN without general SAT search.

Harrison Green, Claire Le Goues, and Fraser Brown

petrinet-concurrency

115

phnf

117

54 GBD instances · 11 validation · baseline 6/11

27 GBD instances · 6 validation · baseline 2/6

These benchmarks ask whether a graph’s vertices can be assigned one of k colors so adjacent vertices differ. The accepted CNF uses vertex-color variables, per-vertex at-least-one clauses, and same-color edge exclusions, with a restricted symmetry-broken layout.

The task is to choose exactly one value from each disjoint group while avoiding forbidden pairs between groups. In CNF, group clauses require a choice, complete within-group binary exclusions enforce uniqueness, and cross-group binary clauses forbid incompatible value pairs.

Proof format

GRAT

No verif.

+ verif.

Solved

530 ×

12.8 ×

11/11

DSATUR coloring is the primary search, with colors renamed by first occurrence to respect the triangular symmetry breaking. If it fails, bounded bit-parallel clique search seeks a checked (k+1)-clique for a pigeonhole-style UNSAT proof; bounded TabuCol may find a SAT assignment and independently verify it; unresolved cases return UNKNOWN. It handles only this recognized layout. DPR

5.77 ×

0.640 ×

10/11

DSATUR coloring drives the SAT search on the strictly recognized layout, using dense bitsets; success emits a complete assignment. If it fails, greedy and exact clique search seek a (k+1)-clique for a pigeonhole-style UNSAT proof, but failure to find the clique or write the proof returns UNKNOWN; no explicit empty-clause derivation is written. VeriPB

5.77 ×

5.18 ×

10/11

Clique-anchored coloring and DSATUR drive SAT search on the recognized layout, with randomized list-coloring, MRV backtracking, and tabu/min-conflicts fallbacks. A bounded exact clique search can emit an UNSAT cutting-planes proof for a (k+1)-clique within the kernel guard; SAT models are clause-checked, but no external proof checker is invoked, and failure returns UNKNOWN.

philips

116

4 GBD instances · 1 validation · baseline 1/1

These CNFs encode a miter comparing two combinational multiplier circuits driven by the same binary inputs. Gate variables and local circuit relations are expressed as CNF constraints, so satisfiability asks whether the two outputs can differ. Proof format

No verif.

+ verif.

Solved

GRAT

91.1 ×

0.425 ×

1/1

Packed 64-bit truth-table enumeration over a recognized acyclic netlist is the active technique. A clause-checked candidate yields SAT; recognition or proof failures yield UNKNOWN, with no general SAT fallback. The no-candidate path uses RUP equivalences and a recursive Shannon-style refutation with deletions, blocking clauses, and a final empty clause for external checking. DPR

1.61 ×

1.06 ×

1/1

Gray-code enumeration of one recovered 10-bit operand drives the search. Each fixed case uses fresh CDCL with minimized learned clauses; a SAT case returns a checked model. Refuted cases write lifted learned clauses and resolve blocking clauses to the empty clause for external DPR elaboration; unsupported cases return UNKNOWN. VeriPB

438 ×

0.811 ×

1/1

Fixed-shape fingerprinting dispatches on twenty high-fanout inputs. The active path emits a complete binary tree with RUP blocking leaves and cutting-planes joins to a contradiction for external VeriPB checking. It does not internally validate the proof or provide a SAT fallback; recognition or generation failures return UNKNOWN.

Proof format

No verif.

+ verif.

Solved

GRAT

0.875 ×

1.39 ×

1/6

Rectangle compression is active: repeated row and column masks become Boolean features before CDCL searches the model. The bounded recognizer has no generic fallback, so rejected inputs and failed SAT checks return UNKNOWN. UNSAT logging records feature, rectangle, and learned clauses for external elaboration, but an initial-unit conflict may lack an empty clause. DPR

0.747 ×

1.37 ×

0/6

Parity factoring leads the active search: union-find links complementary binary relations, CDCL searches the latent CNF, and finite-domain search verifies assignments before path-consistency and finite-domain CDCL fallbacks. Unsupported structures return UNKNOWN; verified SAT assignments are emitted, while UNSAT uses a witness-free DPR log for external elaboration, with replay failure also returning UNKNOWN. VeriPB

0.747 ×

1.37 ×

0/6

Dominance substitution and failed-assumption probing lead the default search, followed by path-consistency implications and bit-mask propagation in CDCL. The recognizer accepts only bounded complete choice blocks and returns UNKNOWN otherwise; SAT assignments are reparsed and checked, while UNSAT records dominance and RUP steps with a final contradiction, reporting UNKNOWN if proof finalization fails.

pigeon-hole

118

154 GBD instances · 31 validation · baseline 23/31

The benchmark asks whether pigeons can be assigned to holes without exceeding the allowed hole capacity. CNF encodings use placement variables, clauses requiring each pigeon to choose a hole, and clauses preventing forbidden collisions; variants may add capacities or guards. Proof format

No verif.

+ verif.

Solved

GRAT

0.982 ×

1.57 ×

22/31

Exact grid recognition regenerates the expected clause set before accepting a layout and constructs a checked assignment for recognized SAT cases. Specialized UNSAT branches use subset-counting, occupancy-state, implication, finite-path, or relativized reasoning and emit DRAT streams; bounded CDCL handles limited cases, while unsupported or unfinished paths return UNKNOWN. DPR

8.82 ×

2.47 ×

30/31

Incidence reconstruction identifies pigeon rows, holes, and guarded collision structure, directing recognized instances to specialized paths rather than a general CNF solver. Only the standard branch constructs and checks SAT assignments; other recognized variants use clause additions, deletions, PR-style induction, RUP checks, or clique-restricted CDCL proofs, while failed recognition or unfinished paths return UNKNOWN. VeriPB

8.84 ×

14.0 ×

30/31

Structural recognition reconstructs matrix placement columns and collision edges, enabling a direct SAT assignment only on the matrix path. Recognized UNSAT variants use aggregate PB/RUP derivations, while a fixed MIS-wrapper signature alone enables CDCL with RUP-logged learned clauses; unsupported layouts or failed certificate generation return UNKNOWN.

The Case for Automated Hyperspecialization: Evidence from SAT

planning

119

popularity-similarity

121

667 GBD instances · 134 validation · baseline 81/134

42 GBD instances · 9 validation · baseline 3/9

These benchmarks ask whether a Boolean assignment satisfies every clause in a CNF encoding of a planning decision problem. Variables and clauses can represent actions, states, auxiliary choices, or occupancy constraints, while inputs use DIMACS rather than recovered planning semantics.

These instances are 3-CNF formulas, with three literals per clause. The available solver artifacts establish their clause structure but do not recover the generating popularity-similarity model.

Proof format

No verif.

+ verif.

Solved

GRAT

0.719 ×

0.850 ×

57/134

Short-clause occurrence weighting guides bounded DIMACS CDCL, with a dense-instance binary-implication path but no planning-specific reconstruction. Malformed inputs return UNKNOWN; oversized inputs receive only a two-pass model attempt and otherwise return UNKNOWN. Validated SAT assignments are emitted, while UNSAT retains learned clauses as ASCII additions followed by the empty clause. DPR

0.762 ×

0.889 ×

62/134

A narrow multi-robot path-planning reconstruction searches space-time paths with vertex and swap avoidance, then injects a candidate branch. A failed candidate falls back to CDCL rather than proving UNSAT; clause-checked SAT models are emitted, while UNSAT uses a learned-clause journal as DPR additions plus the empty clause, and malformed or unsupported inputs return UNKNOWN. VeriPB

0.759 ×

0.790 ×

62/134

Conservative layout-gated root probing can force learned root units before first-UIP CDCL; it is structural rather than planner-semantic. SAT assignments are reread and clause-checked, while UNSAT translates learned clauses into VeriPB RUP constraints with deletions and an empty constraint; malformed, oversized, or output failures return UNKNOWN.

polynomial-multiplication

120

60 GBD instances · 12 validation · baseline 5/12

The benchmark asks whether length-n polynomial multiplication over GF(2) can be expressed as a sum of T products of three binary factor vectors. Its CNF encodes factor coefficients, conjunctions, and XOR chains enforcing each convolution output. Proof format

No verif.

+ verif.

Solved

GRAT

7.57 ×

5.32 ×

11/12

Recognizing compact GF(2) layouts, it constructs a symmetric rank decomposition, evaluates product and XOR variables, and independently checks every clause before emitting a SAT assignment. For recognized UNSAT cases, including ordered GF(3) layouts, it content-checks and remaps bundled DRAT cores; unsupported layouts or failed checks return UNKNOWN. DPR

7.57 ×

6.24 ×

11/12

Recognizing regular layouts with sufficient rank, it constructs SAT assignments from GF(2) bilinear decompositions, with ordered propagation and full-CNF validation; the alternate layout has no SAT path. For UNSAT, it recovers limited permutations by incidence structure and maps checked embedded DRAT templates for bounded recognized cases; failures or unsupported cases return UNKNOWN. VeriPB

1.52 ×

1.53 ×

7/12

It recognizes the tensor header equations and constructs GF(2) decompositions from diagonal and pair products, using a special reduced-product basis and elimination for n=5. It evaluates all auxiliary variables and clauses before SAT output; it has no general search or UNSAT certification, so failed or unsupported cases return UNKNOWN.

Proof format

No verif.

+ verif.

Solved

GRAT

1.07 ×

1.12 ×

3/9

Bounded variable elimination limits resolvent growth, records eliminated variables, and passes the residual to watched-literal CDCL with degree-seeded EVSIDS. It accepts only the recognized fixed 3-CNF degree profile, returns UNKNOWN on other inputs, checks reconstructed SAT assignments, and may print UNSAT after failed reconstruction without a final empty-clause certificate. DPR

1.09 ×

0.988 ×

3/9

Focused local search uses break-score weighting on variables from unsatisfied clauses, then falls back to bounded Davis-Putnam elimination and a 1-UIP CDCL cube search. It handles the recognized 3-CNF shape, returns UNKNOWN otherwise, checks SAT assignments, and logs witness-free clause additions for external DPR elaboration when cube closure reports UNSAT. VeriPB

1.03 ×

0.898 ×

3/9

Popularity ordering and damped loopy-BP seed decisions and phases; watched-literal CDCL then uses first-UIP learning and restarts. It accepts only the recognized 3-CNF profile, returns UNKNOWN on unsupported inputs or conflict limits, and checks SAT assignments. Enabled proof logging emits learned clauses and level-zero conflicts as VeriPB RUP steps.

prime-factoring

122

185 GBD instances · 37 validation · baseline 27/37

These CNFs ask whether two bounded binary factors multiply to a fixed target integer. Variables encode factor bits, partial products, carries, and output wires, while gate clauses and fixed bits constrain the multiplication circuit. Proof format

No verif.

+ verif.

Solved

GRAT

0.845 ×

0.843 ×

24/37

Structural multiplier recognition recovers operand positions, polarities, and a bounded factor split, then checks the resulting assignment against every original clause. Specialized failures fall back to first-UIP CDCL; SAT yields a checked assignment, root-conflict UNSAT is logged as textual DRAT for external elaboration, and unresolved cases return UNKNOWN. DPR

0.596 ×

0.604 ×

17/37

Structural probing on recognized multipliers uses packed all-zero and singleton-pair tests to recover input polarity and bit order, then searches bounded factors. A restricted CDCL fallback checks SAT assignments. Only the special 8-by-8 EZFact no-factor case emits witness-free blocking clauses followed by a resolution tree; other no-factor results and CDCL UNSAT return UNKNOWN. VeriPB

1.03 ×

0.973 ×

27/37

Polarity-aware structural decoding reconstructs shuffled multiplier bits and tests arithmetic candidates, with a narrower multiplier recognizer as a second direct path. These direct paths provide checked SAT assignments but no UNSAT certificate. If they fail, CDCL logs level-zero UNSAT as VeriPB RUP traces with deletions; unsupported or resource-limited cases return UNKNOWN.

Harrison Green, Claire Le Goues, and Fraser Brown

prime-testing

123

profitable-robust-production

125

58 GBD instances · 12 validation · baseline 7/12

20 GBD instances · 4 validation · baseline 2/4

These benchmarks ask whether structured Boolean arithmetic constructions have satisfying assignments, including fixed-product factors, modular square roots, or repeated gate predicates. CNF encodings use Tseitin-style gate clauses, unit constraints, and equivalences over circuit and word variables.

The benchmark asks whether binary production decisions and auxiliary Boolean variables can satisfy constraints linking production and profit across scenarios. These constraints encode comparisons, arithmetic, and robust minima or related aggregates in CNF, often with order/comparator and Tseitin variables.

Proof format

No verif.

+ verif.

Solved

No verif.

+ verif.

Solved

1.27 ×

1.29 ×

8/12

Proof format

GRAT

GRAT

23.7 ×

17.6 ×

4/4

Arithmetic witness construction drives the main search: factorization or modular-root solving fixes word inputs, and unit propagation completes assignments followed by full clause checks; repeated blocks use watched-literal DPLL with offset translation. Supported failures may use bounded input-tree search, reporting UNSAT only with textual DRAT refutations; unsupported cases return UNKNOWN. DPR

1.56 ×

1.37 ×

9/12

Bit-parallel exhaustive search handles recognized repeated circuits, while arithmetic decoding uses factorization or modular roots followed by propagation and clause checks. Selected fallbacks use restricted CDCL, emitting clauses and an empty clause as DPR additions; unsupported cases return UNKNOWN, while proof-output failure can return SAT without an UNSAT certificate. VeriPB

1.27 ×

1.29 ×

8/12

Bit-parallel evaluation is the primary search for repeated circuits, while product and quadratic layouts recover arithmetic witnesses using factorization or modular-root methods, then complete assignments by propagation and validate every original clause. The solver has no general SAT search or UNSAT or VeriPB certification path; unrecognized or failed constructions return UNKNOWN.

product-configuration

124

23 GBD instances · 5 validation · baseline 5/5

For recognized inputs, support-row canonicalization fixes selected auxiliary rows and installs stride-based assignments, then watched-literal CDCL searches decomposed components with reused phases, learning, and restarts. These restrictions may exclude satisfying assignments; SAT models are checked against the original clauses, while contradictory or exhausted searches return UNKNOWN and no UNSAT certificate is produced. DPR

0.823 ×

0.825 ×

2/4

Binary-implication SCC substitution and bounded elimination remove equivalent or redundant comparator variables before watched-literal CDCL. On SAT, it reconstructs eliminated variables and checks every original clause before printing a model; no DPR proof is emitted, and rejected inputs, contradictions, or failed searches return UNKNOWN rather than UNSAT. VeriPB

0.804 ×

0.805 ×

2/4

False-first phase heuristics with occurrence-signature exceptions initialize watched-literal CDCL, while recognition of the expected ripple-gate prefix only selects the specialized input path. After a SAT result, the solver reparses and checks every original clause before outputting the assignment; it has no VeriPB derivation or UNSAT path, so search exhaustion and internal failure return UNKNOWN.

purdom-instances

126

19 GBD instances · 4 validation · baseline 4/4

These benchmarks ask whether a Boolean assignment can select product or configuration choices while satisfying requested features, exclusions, implications, and support conditions. The instances use flattened CNF with unit, binary, and wider clauses representing these constraints, without requiring recovery of named product domains.

The family asks whether two binary operand vectors multiply to a fixed integer. CNF encodings represent factor bits, partial products, and an AND/XOR carry network, with unit clauses fixing output bits; some encodings may shuffle variables or polarities.

Proof format

No verif.

+ verif.

Solved

Proof format

No verif.

+ verif.

Solved

GRAT

3.88 ×

3.31 ×

5/5

GRAT

0.878 ×

0.388 ×

4/4

Failed-literal probing builds root units in short-clause activity order, followed by watched-literal propagation and activity-guided CDCL with first-UIP learning and restarts. Rejected inputs or proof-output failures return UNKNOWN; SAT returns a checked model, while UNSAT records learned clauses and, on root conflict, the empty clause as textual DRAT additions. DPR

6.63 ×

0.502 ×

5/5

An exclusion-heavy recognizer selects a specialized path, combining root-unit and watched-literal propagation with activity-weighted, negative-first CDCL, first-UIP learning, and restarts. Unsupported structure or internal or proof-output failure returns UNKNOWN; SAT yields a checked model, while UNSAT appends learned clauses and, on root contradiction, an empty clause to textual DPR output. VeriPB

4.24 ×

1.09 ×

5/5

Shortest-open-clause branching follows failed-literal closure, guiding search; conflicts use first-UIP learning, nonchronological backtracking, and a conservative whole-decision nogood fallback. Unsupported structure or checking or writing failure returns UNKNOWN; SAT yields a checked assignment, while UNSAT writes learned clauses and an empty clause as a VeriPB RUP certificate.

Partial-product graph recovery identifies a shuffled, polarity-flipped multiplication circuit, then factors the fixed target and evaluates compatible operands for a clause-checked model. If no compatible pair is found, modular prefix lemmas guide a least-significant-bit-first RUP tree, which writes the UNSAT proof only if it completes; unsupported inputs return UNKNOWN. DPR

917 ×

0.278 ×

4/4

Gate recognition reconstructs a single-writer network, including a clean AND/XOR multiplier. It recovers an odd target, factors it, and checks a model; direct UNSAT cases emit a blocker-tree clause sequence, while augmented cases may use bounded factor search or unlogged CDCL for SAT, returning UNKNOWN on recognition or search failure. VeriPB

0.134 ×

0.189 ×

3/4

Carry-weight and bipartite-graph recovery reconstructs factor-bit order, after which factoring the target enables fixed-point AND/XOR evaluation and independent clause checking for composite cases. For prime targets, first-UIP CDCL supplies RUP-style proof logging; unsupported structures, factoring failures, or an incomplete proof return UNKNOWN.

The Case for Automated Hyperspecialization: Evidence from SAT

puzzle

127

quantum-kochen-specker

129

21 GBD instances · 5 validation · baseline 5/5

11 GBD instances · 3 validation · baseline 2/3

The benchmark asks whether fixed-orientation polyomino components can cover every cell of a rectangular board without overlap. A structured CNF encodes board-cell labels, one-label-per-cell constraints, piece-placement limits, adjacency, and forbidden translations.

The benchmark asks whether an n-vertex graph can meet a specified Kochen-Specker obstruction: Boolean edge variables choose adjacencies, triangle variables represent conjunctions of three edges, and CNF clauses enforce square-freeness and rule out the intended zero-one coloring, sometimes with symmetry breakers.

Proof format

No verif.

+ verif.

Solved

Proof format

No verif.

+ verif.

Solved

GRAT

0.0927 ×

0.0866 ×

4/5

GRAT

0.588 ×

0.695 ×

1/3

It searches recognized fixed, translation-only shapes with bit-parallel exact-cover DFS, constrained-cell branching, and failed-state memoization. Found tilings are expanded and checked before SAT output. When search fails, mapped-placement CDCL attempts UNSAT and emits DRAT additions; unsupported cases or failed proof generation return UNKNOWN without external verification. DPR

0.0546 ×

0.0790 ×

2/5

It searches recognized fixed, translation-only shapes by bit-parallel exact-cover backtracking with constrained-cell branching and failed-state memoization. Found tilings are expanded and checked before SAT output. When search fails on eligible small boards, auxiliary CDCL logs DRAT additions, including witness-free clauses; unsupported cases or an inconclusive fallback return UNKNOWN without external verification. VeriPB

0.164 ×

0.0752 ×

4/5

The solver first tries a hard-coded 17-vertex graph witness, propagates triangle auxiliaries, and verifies complete SAT assignments against the original clauses. Other recognized inputs use edge-biased CDCL with elimination and failed-literal probing; first-UIP clauses support external DRAT elaboration, while only an empty clause yields UNSAT and unsupported or unfinished cases return UNKNOWN. DPR

0.654 ×

0.759 ×

1/3

A hard-coded 17-vertex construction solves residual constraints, propagates triangle auxiliaries, and verifies its assignment against the original CNF. Other recognized cases use edge-prioritized CDCL with selected layer removal and first-UIP learning; witness-free DPR additions support UNSAT certification, while unsupported or unfinished cases return UNKNOWN and the shortcut lacks UNSAT certification. VeriPB

0.469 ×

0.595 ×

0/3

It extracts placement masks from bidirectional implications and solves exact cover by minimum-domain branching, grouped translations, and failed-state memoization. SAT assignments are checked; UNSAT reruns a DFS emitting VeriPB symmetry breakers and RUP nogoods ending in contradiction. No general-SAT fallback is active, so rejected inputs or proof failures return UNKNOWN.

Edge-prioritized CDCL drives the accepted n=17 minimum-degree regime, with watched literals, activity-based branching, phase saving, deletion, and restarts. It checks SAT assignments against the original CNF and reports UNSAT only after serializing its trace as RUP steps for external checking, ending in an empty clause; unsupported inputs or trace failures return UNKNOWN.

pythagorean-triples

quasigroup-completion

128

130

21 GBD instances · 5 validation · baseline 2/5

261 GBD instances · 53 validation · baseline 47/53

Given a collection of Pythagorean triples, assign each represented integer one of two colors so that no triple is monochromatic. In CNF, each triple is represented by complementary monotone clauses, while binary clauses capture constraints left after a vertex is fixed or removed.

The benchmark asks whether a partially filled quasigroup or Latin-square table can be completed so each cell receives one value and each value occurs once in every row and column. CNF uses cell-value choices with one-hot groups and pairwise conflicts, but encodings vary.

Proof format

Proof format

No verif.

+ verif.

Solved

GRAT

0.578 ×

0.697 ×

41/53

GRAT

No verif.

+ verif.

Solved

4.05·104 ×

2.98·104 ×

5/5

Core reconstruction is the active technique: the solver regenerates a bounded Pythagorean 2-core, refines incidence signatures to map the input, and applies an embedded two-coloring. It completes residual variables with non-monochromatic propagation and failed-literal DFS; the restricted recognizer does not validate numeric identities, and recognition or search failure returns UNKNOWN without an UNSAT certificate path.

A strict polarity/parity recognizer recovers three exact-cover dimensions, then tries DSATUR-guided min-conflicts and bounded smallest-column search. Failed recognition or search falls back to CDCL; SAT models are checked, while UNSAT emits textual DRAT for external elaboration and checking, and proof, verification, or parse failures can yield UNKNOWN.

4/5

A phase-shuffled exact-cover detector recovers three dimensions and tries bounded min-conflicts followed by a local walk. Failures use CDCL or, for larger recognized instances, exhaustive Algorithm X with state-clause additions for its proof search; SAT reports checked models, and parse, verification, or proof failures can yield UNKNOWN.

DPR

1.45 ×

1.45 ×

Stochastic local search leads the default path, combining break-based moves, clause weighting, and elite-agreement freezing. If it fails, high-degree decimation is followed by watched-literal CDCL repair on expanding neighborhoods and then the full formula; checked SAT assignments are emitted, while failed or internally unsatisfiable searches return UNKNOWN because no UNSAT certificate path is implemented. VeriPB

0.936 ×

0.936 ×

2/5

Weighted breakout local search is the active first stage, using the best coloring as a phase guide before expanding-neighborhood CDCL repairs and a full CDCL fallback with elimination and reconstruction. The recognizer accepts only this non-monochromatic structure; checked SAT assignments are output, while CDCL UNSAT invokes VeriPB logging and proof failure yields UNKNOWN.

DPR

VeriPB

0.618 ×

0.448 ×

0.702 ×

0.531 ×

41/53

37/53

Watched-literal CDCL drives search, with proof-logged variable elimination, first-UIP learning, and structural recognition used only for phases and preprocessing. UNSAT emits VeriPB-style RUP constraints and a final empty RUP; SAT reconstructs and checks a model but creates no final proof artifact, while malformed input or output setup failure yields UNKNOWN.

Harrison Green, Claire Le Goues, and Fraser Brown

railway-safety

131

ramsey-numbers

133

9 GBD instances · 2 validation · baseline 0/2

6 GBD instances · 2 validation · baseline 0/2

These benchmarks ask whether a bounded railway interlocking or transition-system scenario has a Boolean assignment satisfying its gate, state, and transition constraints together with a selected safety obligation. The encoding uses Tseitin-style CNF clauses over those variables.

The benchmark asks whether the edges of a complete graph can be colored with two colors while avoiding a monochromatic K_s in one color and K_t in the other. CNF uses one variable per edge and clauses for every forbidden clique. Proof format

No verif.

+ verif.

Solved

Proof format

No verif.

+ verif.

Solved

GRAT

1.98 ×

1.07 ×

1.00 ×

1.00 ×

0/2

1/2

GRAT

Preprocessing combines bounded Davis-Putnam elimination and binary-implication SCC reduction with watched-literal CDCL; very wide terminal clauses trigger failed-literal decomposition, while smaller cases use residual CDCL. Outside structural recognition it returns UNKNOWN; SAT models are checked against the original CNF, and UNSAT steps are logged as textual DRAT for external elaboration and checking. DPR

1.00 ×

1.00 ×

0/2

Failed-literal probing ranks literals in a syntactically wide clause; first-UIP CDCL searches branches under their negations and otherwise falls back to ordinary CDCL. Malformed input returns UNKNOWN; SAT models are checked against original clauses, while UNSAT logging records learned clauses and failed-literal units, can include witness-free clause additions, and closes with an empty clause. VeriPB

1.00 ×

1.00 ×

0/2

A final wide, all-negative clause activates assumption-based decomposition: watched-literal first-UIP CDCL tries each tail literal, retains learned consequences, and uses ordinary CDCL otherwise. Outside recognition it returns UNKNOWN; SAT models are checked against original clauses, while UNSAT learned clauses are emitted as VeriPB RUP records followed by an UNSAT conclusion.

Specialized comparator networks reduce selected (4,4,n) and (3,6,n) cases to guarded core cases, using ordering and branch exclusions for the former and selection/profile sorting before CDCL for the latter. It handles dispatched instances through an induced 18-vertex core, emits proof clauses and comments for GRAT checking, and otherwise returns UNKNOWN. DPR

1.00 ×

1.00 ×

0/2

An edge-indexed CDCL engine uses watched literals, clause learning, and restarts; selected recognized cases search an induced core. It accepts only the canonical encoding and supported range, serializes learned clauses and deletions for DPR elaboration and checking, and returns UNKNOWN when unsupported or failing; SAT from a core need not model the full input. VeriPB

1.00 ×

1.00 ×

0/2

An extrema-degree PB route targets (3,6;19), using symmetry and CDCL to establish units before summing degree bounds into a PB contradiction. If it does not finish, symmetry-guided CDCL is the fallback; PB/RUP records are logged, SAT models are checked, only exhaustive encodings are accepted, and unsupported or unresolved cases return UNKNOWN.

ramseycube

134

10 GBD instances · 2 validation · baseline 2/2

ramsey

132

6 GBD instances · 2 validation · baseline 1/2

These benchmarks ask whether the edges of a complete graph can be colored so that specified forbidden subgraphs are not monochromatic. Boolean CNF encodes the edge colors and may add auxiliary variables and clauses for local color constraints. Proof format

No verif.

+ verif.

Solved

GRAT

0.998 ×

0.988 ×

1/2

Line-graph reconstruction drives the K32 four-color path, which applies an F2^4 edge coloring; a direct K18 path instead uses symmetry-breaking sorting and a CDCL tail. SAT assignments are checked against the clauses; UNSAT paths emit textual DRAT traces for external elaboration, while generic Ramsey-shaped inputs use CDCL and rejected inputs return UNKNOWN. DPR

4,220 ×

91.0 ×

2/2

Structural recognition drives Ramsey-specific paths: four-color cases use bounded min-conflicts search on a recovered 3-uniform hypergraph, with a paired two-color fallback. SAT candidates are checked against the input CNF; exact recognized layouts alone receive UNSAT certificates from bundled DPR templates, while failed bounded search or other cases return UNKNOWN. VeriPB

1,990 ×

184 ×

2/2

Min-conflicts recoloring handles recognized four-color layouts, propagating gate variables after a candidate is found; a separate R(4,4;18) path reconstructs edge structure, adds symmetry reductions, and runs CDCL. Clause-checked SAT candidates are accepted; the R(4,4;18) path emits VeriPB red and RUP steps for UNSAT, while R(4,5;25) inputs and failed bounded searches return UNKNOWN.

These CNFs ask whether the edges of a complete graph can be colored with two colors while avoiding monochromatic subgraphs. Variables represent graph edges, and signed clauses encode forbidden 3-cubes, with some encodings also constraining a specified color on four-cycles. Proof format

No verif.

+ verif.

Solved

GRAT

1.04 ×

1.29 ×

2/2

Structural Q3/C4 recognition uses prescribed n=8-20 patterns and fixed-coloring witnesses for small SAT cases, verifying every clause before SAT. Asymmetric instances use a complete K9 core with 1-UIP CDCL, emitting learned clauses and an empty clause for external DRAT elaboration; unsupported inputs and symmetric K13 return UNKNOWN. DPR

3.06 ×

0.738 ×

2/2

Embedded colorings answer small SAT cases by mapping supported lexicographic K_n edge layouts to fixed assignments and checking every clause. Asymmetric inputs with an induced K9 core use 1-UIP CDCL and log learned clauses plus an empty clause; unsupported layouts, failed checks, and symmetric n=13 return UNKNOWN. VeriPB

230 ×

30.8 ×

2/2

Threshold-core detection searches for K13 or K9 complete cores, deletes clauses outside the core, and adds permutation-witnessed symmetry breaking before watched-literal CDCL. Only core UNSAT receives RUP-logged VeriPB output; direct-formula UNSAT and unsupported or unmatched inputs return UNKNOWN, while checked SAT models are returned.

The Case for Automated Hyperspecialization: Evidence from SAT

random

135

random-clustered

137

872 GBD instances · 175 validation · baseline 86/175

36 GBD instances · 8 validation · baseline 4/8

This family asks whether a Boolean assignment can satisfy every clause in a random-style CNF formula, or whether no such assignment exists. Instances encode clauses as sets of signed literals, often with a broadly uniform short width.

The benchmark asks whether a Boolean assignment can make every clause in a sparse 3-CNF formula true. Inputs use three-literal clauses and are intended to exhibit clustered variable co-occurrence, although recognizer filters describe only supported encodings, not the entire family.

Proof format

No verif.

+ verif.

Solved

GRAT

1.44 ×

1.43 ×

114/175

ProbSAT-style local search leads, using false-clause sampling, break/make counts, and restarts; its finite budget is not an UNSAT conclusion before watched-literal CDCL takes over. Verified SAT assignments are emitted; UNSAT requires a textual DRAT stream ending in the empty clause for external elaboration, while nonmatching or failed cases return UNKNOWN. DPR

1.13 ×

1.12 ×

96/175

Incremental probSAT-style local search samples false clauses and uses break/make scoring with random restarts, but is gated off for some narrow-width, dense, or high-ratio cases. Watched-literal CDCL then supplies a checked SAT assignment or, after a level-zero conflict, an UNSAT result with clause-addition proof; rejected inputs and proof or reconstruction failures yield UNKNOWN. VeriPB

1.50 ×

1.50 ×

117/175

Focused ProbSAT/WalkSAT local search leads, sampling variables from false clauses with break-weighted and occasional make-aware moves; default routing can give some recognized shapes an unbounded search, so they may not reach the exact fallback. The exact CDCL fallback verifies SAT assignments and emits VeriPB RUP constraints for UNSAT; solver or proof-file failures yield UNKNOWN.

random-circuits

136

21 GBD instances · 5 validation · baseline 4/5

The task is to choose Boolean primary inputs so an acyclic fixed-width lookup circuit produces a prescribed pattern on its final output bits. CNF uses truth-table clauses for the gates and unit clauses to fix the target outputs. Proof format

No verif.

+ verif.

Solved

GRAT

78.0 ×

61.0 ×

5/5

AVX-512 batched truth-table enumeration first tests compatible primary-input assignments, pruning gates outside the target cone. A CDCL fallback propagates table rows and learns first-UIP clauses when enumeration is unavailable or finds no model; verified SAT assignments are emitted, while unsupported inputs or failed search return UNKNOWN, with no UNSAT certificate path. DPR

1.24 ×

1.24 ×

4/5

Target-cone reduction replaces each relevant forbidden truth-table block with compact implications before a watched-literal CDCL search. The search uses a topological order portfolio, retrying with another order after a bounded attempt; verified SAT models are emitted, while unsupported layouts or failed search return UNKNOWN and no UNSAT certificate is produced. VeriPB

9.80 ×

7.53 ×

5/5

Batched truth-table enumeration with root bitset filtering first tests eligible recognized circuits. The fallback searches backward from fixed outputs, choosing a gate with the smallest compatible preimage set and propagating bitset domains; verified SAT assignments are emitted, while unsupported layouts, contradictions, or exhaustion return UNKNOWN, with no UNSAT certificate path.

Proof format

No verif.

+ verif.

Solved

GRAT

1.00 ×

1.00 ×

4/8

Break-weighted stochastic local search drives accepted 3-CNF instances, choosing variables from unsatisfied clauses and updating clause counts, occurrence lists, and unsatisfied-clause state incrementally. A found assignment is independently checked before SAT output; rejected inputs return UNKNOWN, while unsuccessful accepted searches continue restarting without an UNSAT proof path. DPR

1.00 ×

1.00 ×

4/8

Structural recognition selects supported clustered 3-CNF instances, then randomized break-weighted local search flips variables from unsatisfied clauses while maintaining clause counts and occurrence data incrementally. It independently checks and emits complete SAT assignments; unsupported inputs return UNKNOWN, while unsuccessful accepted searches restart without an UNSAT or DPR certificate path. VeriPB

1.00 ×

1.00 ×

4/8

Bounded ProbSAT searches recognized instances with break-weighted flips from unsatisfied clauses, then falls back to watched-literal CDCL with clause learning when it fails. SAT assignments are checked against clauses; CDCL UNSAT paths serialize a RUP-based proof for external VeriPB checking, while unsupported inputs return UNKNOWN and local-search failure alone does not establish UNSAT.

random-csp

138

24 GBD instances · 5 validation · baseline 3/5

These instances ask whether each object in a finite binary CSP can take one value while avoiding forbidden unary and pairwise combinations. CNF encodings often use exactly-one selector blocks or compact bit blocks defining legal states, with cross-object clauses forbidding state pairs. Proof format

No verif.

+ verif.

Solved

GRAT

13.7 ×

13.6 ×

5/5

State-level CDCL over the recovered finite-domain CSP leads the search, followed by arc-consistency backtracking and local repair or min-conflicts methods. Reconstructed assignments are checked against the original clauses before SAT output; failed finite phases may enter restart cycles until externally stopped, and no UNSAT certificate path is implemented. DPR

540 ×

534 ×

5/5

Structure recovery compiles recognized one-hot or compact-state CNF into domains with unary costs and forbidden pairs, then uses belief propagation, tabu or min-conflicts search, and consistency repairs. Feasible assignments are mapped back and checked against clauses before SAT output; unsupported or unsuccessful cases return UNKNOWN, with no UNSAT certificate path. VeriPB

19.4 ×

19.3 ×

5/5

Damped belief propagation seeds weighted breakout and tabu search on recovered CSPs; arc-consistency masks and dom/wdeg backtracking provide fallback. Original clauses are checked for SAT assignments; the unbounded WalkSAT route has no complete fallback, while unsupported or failed searches return UNKNOWN and no VeriPB logger provides an UNSAT certificate.

Harrison Green, Claire Le Goues, and Fraser Brown

random-hiddenmodel

139

random-mus

141

117 GBD instances · 24 validation · baseline 16/24

14 GBD instances · 3 validation · baseline 0/3

The benchmark asks whether a Boolean assignment satisfies every clause of a structured CNF formula whose clauses contain two or three distinct variables. A recurring subpattern translates to odd-parity XOR equations, but it is only a special subfamily, not the definition of the benchmark.

This family asks whether one Boolean assignment satisfies every disjunction of signed literals in a CNF formula. The recognized inputs use a restricted mix of two- and three-literal clauses; the label does not establish provenance or minimal unsatisfiability. Proof format

No verif.

+ verif.

Solved

Proof format

No verif.

+ verif.

Solved

GRAT

1.00 ×

1.00 ×

1.16 ×

1.16 ×

17/24

0/3

GRAT

Parity shortcuts and spectral-seeded NAE or focused local search first seek checked satisfying assignments, followed by watched-literal CDCL when those searches fail. The parity and heuristic paths have no UNSAT certificate; the fallback logs learned clauses and a root conflict as DRAT additions for external elaboration, while unresolved or unsupported cases return UNKNOWN. DPR

1.01 ×

1.01 ×

16/24

Validated structure sends odd 3-XOR blocks to packed GF(2) elimination; other recognized formulas use bounded focused search, width-2 inputs go directly to CDCL, and nonmatching inputs return UNKNOWN. Checked SAT models certify satisfiability; the fallback logs witness-free DPR RUP additions for learned clauses and an empty clause, requiring external elaboration and checking. VeriPB

1.28 ×

1.23 ×

18/24

Packed GF(2) elimination handles the recognized odd-XOR signature, then majority/minority seeds and focused ProbSAT search seek a checked model before watched-literal CDCL takes over. The fallback logs learned clauses and the root conflict as VeriPB RUP constraints, including deletions and a final UNSAT conclusion for external verification; malformed or unresolved inputs return UNKNOWN.

ProbSAT local search with break-weighted flips is the active first stage, and verified models terminate the search. Failure triggers bounded elimination, vivification, and watched-literal first-UIP CDCL; SAT assignments are checked, while UNSAT logs derived clauses and an empty clause in a DRAT trace intended for external elaboration, with unsupported inputs returning UNKNOWN. DPR

1.00 ×

1.00 ×

VeriPB

1.00 ×

1.00 ×

140

59 GBD instances · 12 validation · baseline 11/12

The benchmark asks whether a Boolean assignment satisfies every clause in a CNF formula, or whether no such assignment exists. Variables are grouped into communities, with many short clauses local to one community and fewer clauses linking communities. Proof format

No verif.

+ verif.

Solved

GRAT

0.203 ×

0.204 ×

4/12

Community decomposition leads on recognized fixed-size layouts: local DPLL handles contiguous blocks, followed by damped belief-propagation decimation for cross-block search and a bounded elimination followed by CDCL fallback. UNSAT proofs are emitted as textual DRAT for external elaboration; SAT models are checked, and unrecognized or unresolved inputs return UNKNOWN. DPR

0.229 ×

0.228 ×

5/12

Contiguous-community decomposition leads the search on recognized fixed-size layouts: deterministic DPLL handles sufficiently dense community blocks, followed by weighted coordinate repair and CDCL fallbacks for bridge clauses. Local contradictions and global CDCL refutations are recorded as DPR additions; checked SAT models are output, while unsupported or unresolved cases return UNKNOWN. VeriPB

0.331 ×

0.333 ×

7/12

Incremental ProbSAT-style search leads on recognized layouts, followed by community-product block search with projected DPLL repairs and an unbounded CDCL fallback. SAT assignments are checked; a completed CDCL refutation is written as VeriPB RUP steps, while formulas outside the recognized structure or unfinished refutations return UNKNOWN.

0/3

On recognized inputs, break-only probSAT with weighted flips leads the search, and every candidate model is checked against the input. After local search fails, bounded variable elimination and watched-literal 1-UIP CDCL continue until SAT or UNSAT; the fallback logs VeriPB RUP constraints, while probSAT failure alone produces no UNSAT certificate and unsupported inputs return UNKNOWN.

random-planted-solution random-modularity

0/3

WalkSAT with break-count choices leads the search before fallback. If no verified model is found, bounded elimination and watched-literal CDCL emit a textual DPR/RUP-style trace; unsupported inputs and some failures return UNKNOWN, and optional lookahead is inactive by default. External elaboration and checking are outside the solver; no certificate validation is claimed.

142

328 GBD instances · 66 validation · baseline 60/66

The task is to find a Boolean assignment satisfying every clause of a random planted-solution CNF. Recognized inputs have three literals over distinct variables per clause, with selected density ranges. Proof format

No verif.

+ verif.

Solved

GRAT

3.37 ×

3.37 ×

64/66

Triple-parity belief propagation seeds the exact 4.41-density branch, while other recognized densities use signed-pair spectral power iteration; both feed focused local search with break-score and polarity restarts. A verified full assignment is emitted on success, but malformed or out-of-band inputs return UNKNOWN, and recognized failures continue through unbounded restarts without an UNSAT certificate path. DPR

0.796 ×

0.796 ×

57/66

Signed spectral seeding and degree-normalized power iteration initialize break-weighted local search, with belief propagation and survey decimation added only in the exact 4.41-density branch. It verifies any complete satisfying assignment, but unsupported inputs return UNKNOWN and failed recognized searches can remain in unbounded random restarts; no UNSAT certificate path is implemented. VeriPB

0.655 ×

0.655 ×

54/66

Planted-likelihood belief propagation with confidence-guided decimation supplies seeds, followed by bounded Hamming and make-break repair and, on some paths, learned-clause CDCL completion. The solver independently checks a full satisfying assignment; unsupported or unsuccessful runs return UNKNOWN, no UNSAT certificate is produced, and a middle-density branch can remain in unbounded decimation and repair.

The Case for Automated Hyperspecialization: Evidence from SAT

rbsat

143

register-allocation

145

116 GBD instances · 24 validation · baseline 4/24

20 GBD instances · 4 validation · baseline 4/4

The benchmark asks whether each finite-domain object can take exactly one value while avoiding listed forbidden value pairs between objects. Its recognized CNF uses one signed choice block per object, one at-least-one clause, pairwise at-most-one clauses, and binary clauses forbidding cross-object pairs.

The benchmark asks whether an interference graph can assign one of k registers to each vertex so adjacent vertices differ. CNF uses x[v,c]: positive vertex clauses require registers, while negative edge clauses forbid adjacent vertices from sharing a register. Proof format

No verif.

+ verif.

Solved

Proof format

No verif.

+ verif.

Solved

GRAT

13.4 ×

1.75 ×

2.37 ×

2.37 ×

17/24

4/4

GRAT

Structure-aware finite-domain search leads with weighted min-conflicts, arc-consistent repair, and belief-propagation decimation before falling back to watched-literal CDCL with clause learning. SAT assignments are checked against the CNF; UNSAT is certified by a separate CDCL rerun that writes DRAT additions for external elaboration and checking, while recognition or certifying failure can return UNKNOWN. DPR

2.80 ×

2.80 ×

17/24

Signed one-hot recognition and bit-mask arc consistency drive the exact search after a bounded randomized min-conflicts pass, using dom/wdeg variable selection and least-damaging values. SAT assignments are checked; after failed searches, a separate unit-propagation DFS emits a tree-style proof of UNSAT with decision nogoods and an empty clause, without independent verification; unrecognized encodings return UNKNOWN. VeriPB

2.68 ×

2.68 ×

17/24

Structural one-hot recovery drives randomized constructive search, tabu or message-passing repair, weighted local search, and bounded MAC before a CSP-directed CDCL scout. A SAT assignment is clause-checked; only a scout UNSAT result enables a logged CDCL rerun with VeriPB RUP constraints for UNSAT certification, while failure enters unbounded diversified search rather than UNKNOWN.

reg-n

144

60 GBD instances · 12 validation · baseline 8/12

These structured formulas assign Boolean color choices to objects arranged by a hierarchy. Positive clauses require a choice for each object; tree or product-based exclusions, sometimes expanded into ternary and longer clauses, constrain which choices can coexist. Proof format

No verif.

+ verif.

Solved

GRAT

1.00 ×

1.01 ×

8/12

Parity-based canonicalization and rank-indexed checking recover tree coordinates and verify every required clause family exactly once. For supported fixed regimes, the solver remaps variables in a packed regime-specific DRAT template and reports UNSATISFIABLE only after replay succeeds; it has no general SAT search, and unsupported or malformed structures or missing proof data return UNKNOWN. DPR

1.00 ×

0.287 ×

8/12

Product-run detection and polarity-specific OR extensions compact recognized contiguous Cartesian-product layouts before a watched-literal first-UIP CDCL search. UNSAT results are logged as a textual DPR-style stream with RAT-style definitions, deletions, and CDCL lemmas, while SAT assignments are checked against the original CNF; failed recognition or internal failure returns UNKNOWN. VeriPB

1.01 ×

1.01 ×

8/12

Canonical variable recovery and exact, order-sensitive clause validation identify the hierarchical coloring structure without a SAT search. It introduces event, z, and a variables with red constraints, aggregates occurrences into macro conflicts, and emits RUP tree-resolution steps in a VeriPB text proof; recognition or output failures return UNKNOWN, and the source includes no verifier.

Maximum Cardinality Search drives chordal coloring for the exact direct encoding (3 <= k <= 20, at most 5000 vertices). Checked colorings provide SAT assignments; if they fail, DSATUR seeks a model or a bounded clique search can emit a pigeonhole DRAT proof, with unresolved or unsupported cases returning UNKNOWN. DPR

15.4 ×

1.34 ×

4/4

Bitset clique search leads the direct-coloring workflow, restricted to complete consecutive-block encodings with 1 <= k <= 63 and at most 8192 vertices; reverse-MCS and greedy coloring seek a checked SAT assignment. A rechecked (k+1)-clique yields a deletion-based pigeonhole DPR proof; otherwise exponential DSATUR emits rejection blocks, and unresolved cases return UNKNOWN. VeriPB

13.6 ×

14.8 ×

4/4

Bitset clique search and maximum-cardinality analysis guide the restricted contiguous-block recognizer for k >= 3; other inputs return UNKNOWN. Greedy coloring seeks a checked SAT assignment, while exact component-wise DSATUR handles unresolved cases. A rechecked clique yields VeriPB pseudo-Boolean proofs, and exact search emits RUP nogoods; model checks or proof-output failures return UNKNOWN.

relational-dependencies

146

20 GBD instances · 4 validation · baseline 4/4

The benchmark asks whether variables can satisfy Horn dependencies while limiting false primary variables and true auxiliary variables. Its CNF encoding uses clauses for a and b imply c plus complete subset blocks that impose the two cardinality limits. Proof format

No verif.

+ verif.

Solved

GRAT

26.0 ×

2.74 ×

4/4

A Horn-closure shortcut sets the first block true and second block false; if it fails, watched-literal CDCL uses implicit cardinality propagation and branches only on the first block. SAT assignments are checked against the input, while UNSAT paths emit textual DRAT; unrecognized formulas or unresolved failures return UNKNOWN. DPR

33.1 ×

1.69 ×

4/4

A closure-pruned primary-variable search uses Horn forward chaining, cardinality cutoffs, and lookahead branch selection; a SAT survey may supply a model or reusable tree. SAT models receive a raw clause check; on recognized instances, UNSAT searches serialize search trees as witness-free DPR additions, while construction or writing failure returns UNKNOWN. VeriPB

35.4 ×

0.504 ×

4/4

Certifying primary-variable search forward-chains Horn rules with bitmask closure and causal conflict clauses; it first tries a bounded natural-order search, then an unbounded degree-scored fallback. SAT leaves receive structural checks, while UNSAT traversals write RUP and cutting-planes resolution steps in VeriPB; unsupported inputs or failed checks or proof failures return UNKNOWN.

Harrison Green, Claire Le Goues, and Fraser Brown

relativized-pigeon-hole

147

risc-instruction-removal-subrv

149

55 GBD instances · 11 validation · baseline 8/11

21 GBD instances · 5 validation · baseline 0/5

The benchmark asks whether p first-layer pigeons can choose distinct objects, while each chosen object chooses a distinct one of only p-1 final holes. A CNF uses choice, activation, and object-to-hole variables, with existence, implication, and collision clauses encoding these two injective assignments.

The benchmark asks whether an acyclic Boolean netlist has an assignment satisfying its gate constraints and three terminal unit clauses. In CNF, each later variable is constrained as a Boolean function of up to three earlier variables. Proof format

No verif.

+ verif.

Solved

Proof format

No verif.

+ verif.

Solved

GRAT

1.00 ×

1.00 ×

1.51 ×

3.36 ×

9/11

0/5

GRAT

Extension-variable composition and a sequential at-most-one counter reduce complete recognized encodings to an ordinary pigeonhole contradiction. The UNSAT branch writes DRAT records containing extension clauses, derived clauses, deletions, and an empty clause; one omitted color-collision case instead gets a checked SAT model, while unsupported, over-limit, or proof/model failures return UNKNOWN. DPR

1.51 ×

0.720 ×

9/11

Structural recognition drives a square-case matching proof and a rectangular-case witness-backed repair with case splits that derive canonical selector units before an inner pigeonhole proof. Clause additions may be witness-free; one-clause-short forms receive only a checked SAT construction, limited to supported missing-tuple cases, and unsupported inputs or failures return UNKNOWN. VeriPB

2.94·104 ×

335 ×

11/11

Pseudo-Boolean clique counting is central: lower bounds force at least p active objects, while per-hole demand and clique capacities cap them at p-1, yielding a streamed UNSAT proof. One omitted second-layer collision instead uses a constructed and checked SAT model with no UNSAT certificate path; unsupported forms or failures return UNKNOWN.

It uses bounded-cone MiniCDCL and parity union-find to discover endpoint relations, then applies truth-table closure with local clause proving. It reports UNSAT only when closure derives a contradiction, emitting textual DRAT additions for external elaboration and checking; unsupported inputs or failed closure return UNKNOWN, and no SAT-witness path is implemented. DPR

1.00 ×

1.00 ×

VeriPB

1.00 ×

1.00 ×

148

9 GBD instances · 2 validation · baseline 1/2

These benchmarks encode relational checks on a processor circuit. Primary inputs and ordered auxiliary variables represent Boolean signals, while Tseitin gate clauses and fixed endpoint conditions ask whether the encoded circuit admits a consistent assignment. Proof format

No verif.

+ verif.

Solved

GRAT

0.667 ×

1.55 ×

0/2

Packed truth-pattern simulation groups equal or complemented variable signatures to propose units and equivalences, then checks each candidate by watched-literal RUP. Bounded propagation-only closure emits accepted RUP clauses and an empty clause on contradiction; unsupported encodings or no contradiction return UNKNOWN, and no SAT assignment path is implemented. DPR

0.667 ×

1.55 ×

0/2

Reduced ordered BDD construction uses interleaved primary-input ordering and unique-node/apply caching to evaluate the recognized circuit directly. It reports UNSAT only for a terminal-true target, emitting witness-free DPR additions for the derivation and final empty clause; strict width-8 and input-count checks, resource limits, and other outcomes return UNKNOWN, with no SAT witness path. VeriPB

0.667 ×

1.55 ×

Bounded Davis-Putnam elimination preprocesses the CNF by removing low-occurrence variables within resolvent bounds, then watched-literal CDCL solves the reduced instance. It logs BVE resolvents, learned clauses, and terminal conflicts as VeriPB steps, checks SAT assignments against the original clauses, and returns UNKNOWN when eliminated variables cannot be reconstructed.

0/2

0/5

It uses formula-seeded, 64-lane bit-parallel evaluation of recognized functional gates, filtering lanes by the unit clauses and selecting a surviving assignment for direct clause checking. It emits a checked SAT assignment when one survives, but has no UNSAT or VeriPB certificate path; unsupported inputs or no surviving lane return UNKNOWN.

rooks risc-instruction-removal-golcrest

0/5

It uses template-based gate propagation in CDCL, initially deciding on non-output variables and falling back to all variables when that queue empties. It reports UNSAT after a root conflict and logs learned clauses with a terminal empty clause, but has no SAT-witness path; unsupported, oversized, or completed satisfiable cases return UNKNOWN.

150

30 GBD instances · 6 validation · baseline 0/6

The benchmark asks whether n+1 squares can be selected on an n by n board while no row contains more than one selected square. CNF uses cell variables, row-conflict clauses, and an exact-cardinality constraint represented by a BDD. Proof format

GRAT

No verif.

+ verif.

Solved

1.10·104 ×

7.43 ×

6/6

Layered ordered-BDD recognition drives a frontier-state proof: sequential row-prefix variables track occupied rows, and minimal fixed-cardinality masks are propagated from deep BDD layers upward to emit the certificate. It is UNSAT-only for the recognized row-based encoding; unsupported structure, resource limits, or proof-generation failure yield UNKNOWN rather than invoking general SAT search. DPR

1.05·104 ×

5.30 ×

6/6

BDD recovery and frontier-limited dynamic programming drive the proof: occupancy variables and bitmask states summarize filled rows, while threshold-based minimal occupied subsets prune recursive BDD traversal; a cofactor fallback handles canonical ITE cases. It is UNSAT-only for the narrow recovered row-capacity encoding; recognition or generation failure returns UNKNOWN without general SAT search. VeriPB

584 ×

14.9 ×

6/6

Pseudo-Boolean cutting planes drive an UNSAT certificate: row at-most-one inequalities provide an upper bound of n, while proof-only one-unit flow through the ordered BDD and a rank-weighted potential derive the lower bound n+1. The solver accepts only the recognized square, row-based BDD structure; recognition or proof-emission failure returns UNKNOWN, with no SAT/model path.

The Case for Automated Hyperspecialization: Evidence from SAT

rubikcube

151

satcoin

153

20 GBD instances · 4 validation · baseline 0/4

20 GBD instances · 4 validation · baseline 0/4

Bounded Rubik’s Cube reachability asks whether legal face turns can transform an initial cube state into the solved state within a fixed move horizon. A CNF encoding represents move choices and intermediate states, constraining transitions and the final solved configuration.

The supported satcoin instances ask whether a 32-bit nonce in a specified range makes a Bitcoin-genesis header satisfy a double-SHA-256 difficulty target. The CNF bit-blasts the hashing circuit and the constraints selecting candidate nonces.

Proof format

No verif.

+ verif.

Solved

Proof format

No verif.

+ verif.

Solved

GRAT

1.00 ×

1.00 ×

0/4

GRAT

1.00 ×

0.557 ×

0/4

Watched-literal CDCL with activity branching, learning, restarts, and clause reduction is the main search, using recovered move-variable preferences only for a canonical cube layout. An optional two-phase search precedes it and falls back to CDCL; checked SAT assignments are accepted, while UNSAT uses DRAT-style traces for external certification and proof or certification failures return UNKNOWN. DPR

1.00 ×

1.00 ×

0/4

Recovered cube permutations drive two-phase cubie search with BFS pruning and move-word restrictions, but only on the recognized canonical layout and bounded horizon. A path gets auxiliary completion and full original-CNF checking; failed search or validation returns UNKNOWN, while optional proof writers emit DRAT-style traces for external DPR elaboration and checking. VeriPB

1.00 ×

1.00 ×

0/4

Generator-numbered gadget validation, coordinate pruning tables, and depth-fixed two-phase DFS search only the recognized layout and phase schedule. On failure, ProofCDCL probes scoped suffixes for SAT, then runs watched-literal CDCL with RUP lemmas and PBC subproofs; checked models are accepted, while UNSAT requires a level-zero conflict and otherwise returns UNKNOWN.

sat-x

Bit-parallel unit propagation tests up to 64 interval candidates at once after a bounded two-equality recognizer identifies the 32-bit range. When every lane is refuted, an MSB-first interval tree emits propagation-checked DRAT additions and deletions for external elaboration; a surviving lane, recognition failure, or proof-generation failure returns UNKNOWN. DPR

1.00 ×

VeriPB

1.00 ×

1.00 ×

0/4

152

scheduling

The benchmark asks whether fixed-width integer words satisfy a cube-sum equation, typically X^3 - Y^3 - Z^3 = 3, with optional bounds or range restrictions. CNF represents the arithmetic with bit vectors and Tseitin auxiliary variables, plus clauses enforcing constants, targets, comparisons, and exact-width arithmetic.

435 GBD instances · 87 validation · baseline 52/87

Proof format

No verif.

+ verif.

Solved

GRAT

1.01 ×

1.02 ×

1/4

Bounded signed cube identities and 2-adic cube-root lifting generate candidate words; propagation and residual DPLL complete and validate each candidate before SAT output. Failures return UNKNOWN, and the cube path has no UNSAT certificate; only a narrow 16-bit Brocard branch uses CDCL and logs learned clauses plus a contradiction for external DRAT elaboration. 0.758 ×

0.764 ×

0/4

An embedded arithmetic witness fixes the inputs of recognized cubic layouts; flat propagation and residual watched-literal DPLL complete the assignment and rescan all clauses before SAT output. Failed searches return UNKNOWN; the cubic path has no UNSAT certificate, while a narrow Brocard fallback uses CDCL and logs learned clauses for external DPR elaboration and checking. VeriPB

0/4

Double-SHA-256 evaluation of extracted nonces drives the search after recognizing a restricted CNF layout and interval. A hash hit returns UNKNOWN because no CNF model is reconstructed; otherwise a prefix tree emits VeriPB blockers and RUP parent steps for UNSAT, with structural leaf checks and UNKNOWN on recognition or proof-writing failure.

20 GBD instances · 4 validation · baseline 1/4

DPR

1.00 ×

64-lane bit-parallel propagation scans recognized interval candidates, with scalar clause propagation as a fallback for unresolved groups. An all-clause-satisfying candidate yields a checked complete assignment. When all candidates are refuted, it adds blocking clauses for covered cubes, including witness-free batch additions, and uses a high-bit-first closing tree; unresolved, recognition, or resource-limit failures return UNKNOWN.

0.758 ×

0.764 ×

0/4

For recognized cubic layouts, an embedded witness fixes the cubic inputs; occurrence-indexed propagation and false-first residual backtracking complete the assignment and validate clauses before SAT output. A small Brocard fallback uses watched-literal CDCL and emits VeriPB RUP steps for UNSAT; failures return UNKNOWN, with no generic UNSAT certificate or in-process proof check.

154

These scheduling CNFs ask whether games or jobs can be assigned to rounds, dates, venues, or workers while respecting one-assignment, conflict, home/away, break, and exclusion constraints. Variables represent assignments and auxiliary conditions, and clauses enforce the allowed choices and incompatibilities. Proof format

No verif.

+ verif.

Solved

GRAT

0.645 ×

0.687 ×

25/87

Circle-method one-factorization and layered orientation dynamic programming construct recognized double-round-robin schedules while tracking venue alternation and rejecting forbidden break runs. Watched-literal propagation completes auxiliary variables and validates SAT assignments; bounded MiniCdcl or Hall and matching refutations cover selected cases, while failed or unsupported cases return UNKNOWN and UNSAT paths log DRAT additions. DPR

0.572 ×

0.609 ×

17/87

Structural fingerprints dispatch narrow encodings; circle factorization fixes Break schedules and derives home/break variables, followed by unit propagation and residual auxiliary-only DPLL. SAT valuations are clause-scanned, but one CDCL path can leave entries unassigned and validate them as true; selected UNSAT cases emit tree or learned-clause additions, while unsupported or failed cases return UNKNOWN. VeriPB

0.830 ×

0.881 ×

39/87

Symmetry-preserving red and RUP derivations exploit recognized round-robin structure, while matching uses augmenting paths and Hall-style deficiencies. A bounded monotone-choice search, with stochastic construction only outside break encodings, precedes watched-literal first-UIP CDCL; SAT assignments are checked, specialized UNSAT derivations require proof output, and unresolved cases can return UNKNOWN.

Harrison Green, Claire Le Goues, and Fraser Brown

school-timetabling

155

sgen

157

60 GBD instances · 8 validation · baseline 6/8

128 GBD instances · 26 validation · baseline 4/26

The task asks whether courses with weekly loads can be assigned time slots and eligible teachers or resources while respecting daily blocks, occupied-day limits, availability, and non-overlap constraints. CNF encodings use course-period, day-use, resource-choice, and auxiliary cardinality or conflict variables.

The benchmark asks whether a Boolean assignment satisfies every signed CNF clause. Its instances use structured encodings of exact-cover, counting, matching, or counter constraints, represented by unit, binary, and wider clauses. Proof format

No verif.

+ verif.

Solved

Proof format

No verif.

+ verif.

Solved

GRAT

1.46 ×

1.47 ×

11/26

GRAT

0.945 ×

0.945 ×

5/8

A narrow translated-layout recognizer extracts durations, daily limits, educator choices, and availability, then schedules contiguous blocks with legal day allocations and bit masks. Min-conflicts with tabu repair and random restarts seeks a model; watched-literal completion and full clause checks precede SAT output. Unsupported layouts return UNKNOWN, while scheduling failures retry without an UNSAT certificate path. DPR

1.42 ×

1.42 ×

6/8

A structural recognizer extracts the timetable grid, resource choices, and cardinality constraints, then uses a weighted Hall/max-flow gate followed by randomized min-conflicts repair and bit-mask interval placement. Propagation and full clause checking precede SAT output; failed construction or checking returns UNKNOWN, and no UNSAT certificate path is implemented. VeriPB

0.793 ×

0.790 ×

5/8

An encoding detector identifies grid, loads, options, and intervals, then uses Horn/implication closure, min-conflicts repair, and source-sensitive moves. Small residual neighborhoods use watched-literal CDCL completion with full clause checks before SAT output. Parse or detection failures return UNKNOWN; search may restart indefinitely without an UNSAT certificate path.

set-covering

156

Signed five-variable exact-cover recognition leads to Algorithm X/DLX, followed by potentially unbounded stochastic repair, while separate partition-gap and counter-gap patterns use watched-literal CDCL. Checked SAT assignments are printed; CDCL UNSAT logs learned clauses, deletions, and an empty clause as textual DRAT for external elaboration and checking, whereas unsupported or unresolved cases return UNKNOWN. DPR

1.57 ×

1.58 ×

12/26

Signed exact-cover reconstruction drives fixed-seed min-conflicts with tabu moves, followed by smallest-constraint Algorithm X; compact or all-ternary cases can fall back to watched-literal CDCL. Verified SAT assignments are printed, while CDCL UNSAT runs log learned clauses, deletions, and an empty clause for external elaboration; failed searches and unresolved larger cases return UNKNOWN. VeriPB

3,900 ×

261 ×

26/26

Pseudo-Boolean counting, binary-clique, and sequential-counter recognition handle structured UNSAT cases first, while SAT search uses augmenting-path matching or bit-parallel exact-cover DFS. UNSAT certificates come from counting or counter derivations, or from restricted exact-cover refutation with inequalities and RUP nogoods; checked SAT models are printed, and unsupported, failed, or uncertified paths return UNKNOWN.

sgen-balanced

158

40 GBD instances · 8 validation · baseline 7/8

4 GBD instances · 1 validation · baseline 1/1

Choose a subset of columns that hits every positive requirement while never choosing both endpoints of a declared conflict. In CNF, requirements are positive clauses and conflicts are distinct binary negative clauses, yielding an independent hitting-set decision problem.

The benchmark asks whether a Boolean assignment satisfies a structured 3-CNF formula. Each clause has three distinct variables, while occurrences are distributed nearly evenly across variables and polarities, often through ordered rounds or layers; the recognizers accept only particular such layouts.

Proof format

No verif.

+ verif.

Solved

GRAT

71.5 ×

58.9 ×

8/8

A conflict-aware local search first builds an independent hitting set through randomized restarts and replacement moves that evict conflicting selections. It then falls back to uncovered-row CDCL with learned clauses; checked SAT assignments are reported, while UNSAT is logged in textual DRAT for external checking, and unsupported or failed cases return UNKNOWN. DPR

81.8 ×

52.0 ×

8/8

A bounded focused local search scores flips by coverage and conflict effects, then falls back to complete fail-first search when it misses a model. The exact path tracks row coverage and forbidden neighborhoods, checks SAT assignments, and emits witness-free RUP clauses in textual DPR for UNSAT; unsupported forms return UNKNOWN. VeriPB

1.73 ×

1.77 ×

7/8

An independence-preserving min-conflicts pass first seeks a cover by evicting conflicting selections, then falls back to exact bitset search on the smallest compatible domain. It uses conflict-directed backjumping and sibling exclusions, checks SAT assignments, and emits VeriPB RUP no-goods for external UNSAT checking; unsupported input or proof-file failure returns UNKNOWN.

Proof format

No verif.

+ verif.

Solved

GRAT

9.27 ×

24.6 ×

1/1

On recognized formulas, shallow decisions use clause-pressure scores and choose polarity toward lower pressure; deeper search falls back to activity-based CDCL with first-UIP learning, restarts, and learned-clause reduction. SAT models are checked before printing, while UNSAT runs emit learned-clause additions, deletions, and final empty clause for external elaboration; unsupported formulas or proof-file failure return UNKNOWN. DPR

1.09 ×

2.61 ×

1/1

Exact ordered ten-layer inputs, optionally with a partial layer, are recognized; others return UNKNOWN. Dense pivots trigger independent-set elimination with resolvents and reconstruction; CDCL is the fallback. SAT checks precede assignment output; UNSAT records learned clauses, resolvents, deletions, and an empty clause; reconstruction failure aborts and model-check failure returns UNKNOWN. VeriPB

0.998 ×

2.72 ×

1/1

After balanced-shape recognition, it first runs bounded probSAT local search with break-count weighting, then independently falls back to watched-literal CDCL with phase saving and first-UIP learning. Verified SAT models are printed; UNSAT runs emit parent-logged clauses and a final empty-clause RUP for external checking. Unsupported inputs, proof-file failure, or failed model checks return UNKNOWN.

The Case for Automated Hyperspecialization: Evidence from SAT

sliding-puzzle

159

software-bmc

161

62 GBD instances · 13 validation · baseline 0/13

18 GBD instances · 4 validation · baseline 4/4

The benchmark asks whether a 5x5 board with one blank and 24 numbered tiles can reach the ordered goal arrangement within a bounded number of moves. Its CNF encodes initial states, move transitions, and goal checks for this reachability question.

The benchmark asks whether a bounded execution of software can satisfy its Boolean transition and assertion constraints. Bit-blasting program and state bits, with Tseitin variables for intermediate gates, turns this question into CNF whose satisfiability represents a possible execution.

Proof format

No verif.

+ verif.

Solved

GRAT

1.86 ×

1.86 ×

6/13

IDA* search combines Manhattan distance with immediate-reversal and parity pruning, then a linear-conflict bound, after recognizing the fixed CNF layout. A found path fixes the encoded state variables, and residual DPLL completes and checks the SAT assignment; no UNSAT certificate path exists, so unsupported or unsolved cases return UNKNOWN. DPR

1.86 ×

1.86 ×

6/13

IDA* search updates Manhattan distance, suppresses immediate reversals, and advances bounds by two on the fixed 5x5 layout. Found paths yield checked SAT assignments; after failure, a boundary-crossing construction with PHP-style symmetry clauses may emit DPR proof through horizon 31, reporting UNSAT only on success; other cases return UNKNOWN. VeriPB

1.86 ×

1.86 ×

6/13

Bounded depth-first A* search combines Manhattan and linear-conflict bounds with immediate-reversal pruning, child ordering, and a failed-state table. For a found path it tries direction mappings, fixes selector bits, propagates and checks a complete SAT assignment; no UNSAT certificate path exists, so unsupported or no-path cases return UNKNOWN.

social-golfer

160

Proof format

No verif.

+ verif.

Solved

GRAT

2.94 ×

2.96 ×

4/4

Watched-literal CDCL with 1-UIP minimization and topology-biased decisions drives accepted instances; binary implication specialization serves large bounded-width cases, while local repair is inactive. Checked SAT models are printed, while root-level UNSAT emits learned clauses and the empty clause as textual DRAT for validation; unsupported inputs or failed checks return UNKNOWN. DPR

1.58 ×

4/4

1.54 ×

Bounded variable elimination with non-tautological resolvents precedes watched-literal CDCL and first-UIP learning on selected formulas; guarded cases use CDCL alone. The aggregate-only recognizer sends unrecognized or resource-failing cases to UNKNOWN rather than falling back, while SAT reconstructs assignments and UNSAT logs clause additions for external DPR elaboration and checking. VeriPB

1.09 ×

0.649 ×

4/4

Polarity-biased watched-literal CDCL drives recognized inputs, using first-UIP learning, phase saving, and geometric restarts; oversized instances are rejected. SAT assignments are checked against original clauses, while UNSAT emits learned clauses as VeriPB RUP steps plus an empty RUP and UNSAT conclusion; malformed, unrecognized, or storage-failing cases return UNKNOWN.

39 GBD instances · 8 validation · baseline 3/8

The task arranges players into G groups of P each week for W weeks, using every player once weekly and preventing any pair from meeting twice. CNF clauses encode these group and pair constraints, sometimes with auxiliary variables. Proof format

No verif.

+ verif.

Solved

GRAT

1.27 ×

1.27 ×

4/8

Conflict-directed coloring first extracts matching structure from the Table encoding, then falls back to Context or classic reconstructions with specialized schedules, watched propagation, and 1-UIP CDCL. A full SAT assignment is checked against every original clause; unsupported inputs, failed construction, and exhausted search return UNKNOWN, with no UNSAT certificate path implemented. DPR

2.55 ×

2.55 ×

6/8

Structural CNF recognition drives schedule reconstruction: table, context, and classic layouts are decoded by incidence components, pair fingerprints, and supported schedule constructions. The resulting assignment is checked against every original clause, but unsupported or failed cases return UNKNOWN and no UNSAT result or DPR trace is produced. VeriPB

2.55 ×

2.54 ×

6/8

Strict Table and Context recognition seeds embedded or finite-field affine schedules, with symmetry alignment and recursive clique-factor search as bounded construction fallbacks. Unit propagation and false-phase completion fill the assignment before every original clause is checked; unsupported or unsuccessful constructions return UNKNOWN, and no UNSAT certificate path is implemented.

software-verification

162

277 GBD instances · 56 validation · baseline 36/56

These CNFs encode bit-blasted software-verification conditions. Boolean variables represent circuit or state signals, while clauses express gate relations, fixed values, and other verification constraints. Proof format

No verif.

+ verif.

Solved

GRAT

0.950 ×

1.71 ×

35/56

A bounded reverse-circuit probe evaluates grouped gate definitions in parallel with 64-bit lanes and checks candidate assignments against the original CNF. On probe failure, watched-literal CDCL with first-UIP learning and restarts provides the fallback, logging textual DRAT clauses; the probe cannot establish UNSAT, and unrecognized or oversized inputs return UNKNOWN. DPR

0.708 ×

1.25 ×

24/56

Two bounded low-variable probes with opposite phases try SAT or UNSAT before unbounded CDCL; exhausted probes are discarded. The fallback uses watched-literal CDCL with VSIDS and first-UIP learning, checking SAT models and logging proof-minimized clauses that may be witness-free, including the empty clause for a root conflict. Parse, model-check, and resource failures return UNKNOWN. VeriPB

0.638 ×

1.01 ×

20/56

A pre-conflict bias toward low variable IDs exploits the expected primary-before-derived allocation. The sole search engine is watched-literal CDCL with VSIDS, phase saving, first-UIP learning, geometric restarts, and RUP logging. Checked SAT assignments are emitted; UNSAT produces an empty RUP contradiction and VeriPB conclusion, while malformed, oversized, proof-output, or model-check failures return UNKNOWN.

Harrison Green, Claire Le Goues, and Fraser Brown

sorting-networks

163

st-connectivity-principle

165

23 GBD instances · 5 validation · baseline 1/5

2 GBD instances · 1 validation · baseline 0/1

These instances ask whether a bounded-depth comparator network can be chosen to sort every encoded Boolean input on its channels. CNF variables select disjoint comparator pairs in each layer, while auxiliary simulation variables and clauses enforce channel use and sorted outputs.

The benchmark asks whether two edge-selected path systems on identical rectangular grids can connect alternating terminal pairs without sharing a vertex. Boolean edge variables and CNF clauses enforce local degree or parity rules, plus cross-color clauses forbidding selected edges from meeting at one vertex.

Proof format

No verif.

+ verif.

Solved

Proof format

No verif.

+ verif.

Solved

GRAT

1.00 ×

1.00 ×

1/5

GRAT

1.00 ×

1.00 ×

0/1

Canonical comparator-block recognition only supplies phase and activity hints for selected layouts, while watched-literal CDCL performs the search with learning, restarts, and clause reduction. SAT assignments are checked against original clauses and printed; UNSAT emits learned additions and an empty clause as an external DRAT stream, while unsupported inputs return UNKNOWN. DPR

1.34 ×

1.34 ×

2/5

Structural conflict/OR recognition admits supported layouts, then hard-coded comparator cores seed selected 13- and 16-channel searches before watched-literal CDCL fallback. Fallback SAT assignments are checked against the CNF, while UNSAT emits learned witness-free DPR/RUP additions and an empty clause; seeded UNSAT or unsupported input returns UNKNOWN. VeriPB

2.01 ×

1.99 ×

3/5

Min-fill bucket elimination is the active refutation strategy, using literal masks and subsumption pruning while resolving variables and dynamically rebucketing when needed. It emits retained clauses and the empty clause as textual DRAT for external elaboration, reports UNSAT only on a refutation, and returns UNKNOWN if recognition or refutation fails. DPR

1.00 ×

1.00 ×

0/1

Bounded frontier-ordered Davis-Putnam elimination is attempted first, using compact masks and subsumption before an in-process watched-literal CDCL tail handles the residual formula. It logs resolvents and learned clauses, including witness-free additions, and reports UNSAT only after an empty-clause conflict; restrictive recognition, width bounds, or complete assignments return UNKNOWN. VeriPB

10.9 ×

6.59 ×

1/1

Odd blocks use embedded comparator portfolios, occurrence propagation, and residual DPLL; even blocks use watched-literal CDCL, with the seven-block fallback also using CDCL. Checked SAT assignments are printed; even-branch UNSAT logs RUP constraints and an empty clause, while unsupported, even-SAT, or other unsuccessful cases return UNKNOWN; the seven-block fallback has no UNSAT proof path.

Column-wise transfer maintains a subsumption-minimal boundary clause summary, while watched-literal local DPLL searches each column and records propagation and branch-resolution clauses as VeriPB RUP steps. The final unassumed search emits UNSAT only on conflict; recognition or height limits, and a satisfying search with no witness path, return UNKNOWN.

ssp-0

station-repacking

164

5 GBD instances · 1 validation · baseline 1/1

The benchmark asks whether a subset of integer-weighted items sums to a fixed target, using Boolean selectors for the choices. A CNF encoding adds auxiliary variables for arithmetic and a ripple-adder circuit constrained to equal that target. Proof format

No verif.

+ verif.

Solved

GRAT

0.319 ×

0.320 ×

0/1

It recognizes a rigid ripple-adder layout and uses 64-bit bit-sliced propagation of singleton selector assignments to reconstruct the encoded weights and target. Packed meet-in-the-middle subset-sum search supplies candidates, which are propagated and checked against every CNF clause; failed recognition, search, or validation returns UNKNOWN, with no UNSAT certificate path. DPR

0.319 ×

0.320 ×

0/1

It uses fixed-cardinality meet-in-the-middle search over complementary selector halves, after recovering weights and the target from bit-lane evaluations of the ripple-carry circuit. Candidate assignments are expanded and checked against the full CNF; unsupported, negative, or candidate-free inputs return UNKNOWN, and no UNSAT or DPR certificate path is implemented. VeriPB

0.319 ×

166

9,842 GBD instances · 1,969 validation · baseline 1673/1969

0.320 ×

0/1

It uses bit-parallel evaluation of the recognized topological circuit on singleton selector assignments to recover weights and target bits. A low-64-bit meet-in-the-middle filter generates candidates, each checked against the complete CNF to reject collisions; unsupported or uncertifiable searches return UNKNOWN, with no UNSAT or VeriPB certificate path.

Choose exactly one channel for each station while avoiding incompatible pairs of selected station-channel options. A CNF encoding uses variables for station-channel choices, positive clauses to require a choice, and negative clauses to forbid multiple choices or conflicting pairs. Proof format

No verif.

+ verif.

Solved

GRAT

0.544 ×

0.556 ×

1313/1969

Peeling with guaranteed extensions and station-level TabuCol-style search, including breakout weighting, leads to watched-literal CDCL with EVSIDS, restarts, and first-UIP learning. SAT candidates are extended and checked; only CDCL establishes UNSAT, recording learned clauses and deletions in a textual DRAT trace ending in the empty clause while unsupported inputs return UNKNOWN. DPR

0.615 ×

0.622 ×

1402/1969

Randomized min-conflicts search with tabu moves, restarts, and bounded residual repair first builds channel assignments using compact masks and station ordering, then falls back to watched-literal CDCL with learning and witness-free DPR/RUP clause additions. Only CDCL can establish UNSAT; checked SAT assignments are certificates, while unsupported encodings or failed checks return UNKNOWN. VeriPB

0.458 ×

0.466 ×

1184/1969

Bit-mask propagation with randomized greedy/min-conflicts construction and a tabu, adaptive-weighting retry leads to component-wise CDCL with watched clauses, first-UIP analysis, and nonchronological backtracking. SAT models are independently checked; CDCL UNSAT results are replayed with RUP steps into a VeriPB certificate, while unsupported inputs or replay failures return UNKNOWN.

The Case for Automated Hyperspecialization: Evidence from SAT

stedman-triples

167

subgraph-isomorphism

169

44 GBD instances · 9 validation · baseline 2/9

194 GBD instances · 39 validation · baseline 12/39

These CNFs ask whether choices of change-ringing calls, orientations, and successor states can satisfy all constraints of a lifted Stedman state system. The choices use construction variables, one-hot or compact state variables, and transition variables, with CNF clauses enforcing allowed states and propagated changes.

Subgraph isomorphism asks whether vertices of a pattern can be injectively assigned to host vertices while preserving edges. CNF encodings can represent these choices as one-hot domains or bounded bit blocks and add clauses forbidding incompatible pairs. Proof format

No verif.

+ verif.

Solved

Proof format

No verif.

+ verif.

Solved

GRAT

1.07 ×

0.924 ×

0.861 ×

0.861 ×

1/9

14/39

GRAT

Explicit six-state block recognition leads into watched-literal CDCL with first-UIP learning, activity-guided branching, and phase-randomized restarts. SAT assignments are checked against every original clause before printing; UNSAT emits textual DRAT clauses and a terminal empty clause for external GRAT elaboration, while unsupported inputs or internal failures return UNKNOWN. DPR

0.889 ×

0.889 ×

1/9

Graph-projected search extracts oriented transitions, applies breakout walks and subtour cuts, then seeds exact projected CSP search before CDCL; recognized compact or signed layouts go directly to CDCL. It emits checked SAT assignments only; no DPR proof path exists, and unsupported or unsuccessful searches return UNKNOWN rather than certified UNSAT. VeriPB

0.791 ×

0.791 ×

0/9

Direct and logarithmic recognition drives finite-domain CSP search with propagation, MRV branching, and recursive proof search; shortened logarithmic or auxiliary forms use watched-literal CDCL. Checked SAT assignments are reported; UNSAT branches emit textual DRAT traces with learned clauses or tree nogoods, while unsupported inputs return UNKNOWN. DPR

0.846 ×

0.362 ×

7/39

Graph-aware CSP search is distinctive: recognized direct or logarithmic forms with recovered alignment receive matching and neighborhood filters, while unaligned forms use only propagation. SAT assignments are validated; UNSAT search emits witness-free DPR additions, extending logarithmic blocks with one-hot clauses; unsupported inputs or proof branches not establishing UNSAT return UNKNOWN. VeriPB

0.878 ×

0.892 ×

8/39

Bounded recognition enables call-first branching in watched-literal CDCL with first-UIP learning and all-variable fallback; no specialized transition propagator is active. Checked SAT assignments are printed without a completed VeriPB conclusion; UNSAT logs learned clauses as augmented RUP constraints and writes an UNSAT conclusion unless proof logging is disabled, while unsupported inputs return UNKNOWN.

Graph-aware bitset search distinguishes this solver: eligible one-hot inputs use matching filters, min-conflicts, and clique search, while bit-vector, auxiliary, or failed cases fall back to watched-literal CDCL. SAT models are checked; bounded one-hot UNSAT cases try TreeProof, while other UNSAT cases emit learned-clause RUP certificates and unsupported inputs return UNKNOWN.

stone

10 GBD instances · 2 validation · baseline 0/2

168

14 GBD instances · 3 validation · baseline 1/3

Stone benchmarks ask whether markers can be placed on a two-parent directed acyclic graph, with each marker assigned a red or non-red status. CNF clauses enforce one placement per vertex, red sources, a non-red sink, and two-parent propagation. Proof format

No verif.

+ verif.

Solved

GRAT

1,640 ×

11.2 ×

3/3

Source peeling and structural recognition recover mappings and graph orientation, accepting only an intact transition set or one canonical omission. A sink-reaching non-red omission yields a checked SAT assignment; otherwise it emits a DRAT-style derivation for external checking, while unsupported cases return UNKNOWN without general SAT search. DPR

1,500 ×

8.46 ×

3/3

Co-occurrence mapping and structural recognition recover phases, permutations, and the DAG, then identify missing combinations in the sink cone. The first sink-relevant hole yields a checked SAT assignment; otherwise it emits a proof ending in the empty clause for DPR checking, while unsupported inputs return UNKNOWN without general SAT search. VeriPB

2.18 ×

2.14 ×

2/3

Bit-mask recovery of polarity and color mappings, then DAG reconstruction, validates the shuffled three-source encoding and locates missing tuples. Supported complete instances emit a VeriPB derivation ending in UNSAT; recognized partial instances get checked SAT assignments, but oversized complete or unsupported inputs and proof failures return UNKNOWN without SAT search.

subset-cardinality

170

The benchmark asks whether a shared system of signed exact-two constraints can be satisfied. The CNF represents local blocks with complementary NAE-style clauses and adds endpoint constraints that impose a boundary parity or counting condition. Proof format

GRAT

No verif.

+ verif.

Solved

1.16·105 ×

430 ×

2/2

On recognized complete inputs, complementary NAE-pair recognition drives the strategy: XOR accumulators and ordered incidence swaps cancel shared block variables, with exhaustive boundary rows closing a DRAT proof. A weakened block gets a flow-based SAT-model fallback or modulo-four proof; an unused mod-three routine prevents the supplied source from compiling. DPR

1.23·105 ×

145 ×

2/2

XOR parity cancellation drives complete-instance solving: reconstructed exact-two blocks receive witness-free gate definitions and local lemmas, while incidence swaps eliminate internal variables into a boundary contradiction. With one missing NAE half, it tries a modulo-four proof, then uses enumeration and max flow for SAT; unsupported or failed cases return UNKNOWN. VeriPB

7.19·105 ×

1.06·105 ×

2/2

For the recognized degree-two layout, incidence-graph reconstruction and XOR propagation determine consistent local switching, then a counting-gap test compares required true literals with available variables. When the resulting contradiction is found, it emits one pseudo-Boolean cutting-planes proof; unsupported or noncontradictory cases return UNKNOWN, and no SAT-assignment path is implemented.

Harrison Green, Claire Le Goues, and Fraser Brown

subsumptiontest

171

sum-of-3-cubes

173

19 GBD instances · 4 validation · baseline 4/4

10 GBD instances · 2 validation · baseline 0/2

The benchmark asks whether a Boolean CNF is satisfiable after clauses are repeatedly split into two extensions, adding a variable in both polarities. The resulting clauses may be shuffled, while the family structure consists of complementary twins rather than an arbitrary CNF.

The benchmark asks whether a target integer can be written as the sum of three signed cubes. CNF instances bit-blast bounded cube circuits, using magnitude blocks, sign or selector variables, arithmetic auxiliaries, and fixed target bits.

Proof format

GRAT

GRAT

No verif.

+ verif.

Solved

323 ×

20.7 ×

4/4

A polarity-majority scan proposes a complete model while streaming clauses, with a full recount and rescan as fallback when conflicts appear. For a narrow one-false-clause case, exact sign-flipped pair contractions derive opposite units and emit a textual DRAT refutation; failed recognition, verification, or derivation returns UNKNOWN. DPR

93.6 ×

8.75 ×

4/4

A sign-pattern reduction reverses clause divisions by matching one-bit sign differences and repeatedly shrinking the formula to a core. Bounded propagation and recursive search can yield a checked SAT assignment; UNSAT requires an explicit empty clause and a derivation DAG that may add clauses without witnesses, while unsupported or exhausted cases return UNKNOWN. VeriPB

237 ×

6.00 ×

4/4

Polarity-imbalance scanning first proposes a complete assignment and verifies every original clause. If it fails, width-layered closure of exact complementary twins targets the least-imbalanced variable, then retries globally; a contradiction emits parent-linked VeriPB RUP steps, while only a fully checked recovered model yields SAT and other cases return UNKNOWN.

sudoku

172

30 GBD instances · 6 validation · baseline 0/6

The family represents Sudoku-style digit assignments in CNF, using variables for cell or related finite-domain choices and clauses for Sudoku or other instance constraints. The benchmark question is whether the resulting Boolean formula has a satisfying assignment; some encodings include multiple grids and auxiliary variables. Proof format

No verif.

+ verif.

Solved

GRAT

1.00 ×

1.00 ×

0/6

Watched-literal CDCL with activity-guided branching, 1-UIP learning, non-chronological backtracking, and restarts drives the active search. For SAT it prints an assignment; for UNSAT it writes learned-clause additions and an empty clause as textual DRAT for external GRAT elaboration and checking, while size, parsing, or proof-output failures return UNKNOWN. DPR

1.00 ×

1.00 ×

0/6

Order-prefix recognition prioritizes the first 17 objects, then uses activity-based first-UIP CDCL with clause-scanning propagation, restarts, and fallback branching. The recognizer checks structure, not Sudoku semantics; SAT assignments are checked against original clauses, while UNSAT emits learned clauses plus an empty clause as textual DPR for external elaboration; rejection or parse failure returns UNKNOWN. VeriPB

1.00 ×

1.00 ×

0/6

The compiled executable is an unconditional UNKNOWN fallback: it only checks invocation shape, neither parses the CNF nor searches. An exploratory source file contains row, column, and box filtering with bounded Sudoku backtracking, but it is not compiled or reachable, so no SAT assignment or UNSAT certificate is emitted.

Proof format

No verif.

+ verif.

Solved

2.00·104 ×

4,320 ×

2/2

Integer-construction search is active: the solver recognizes a narrow bit-blasted circuit shape and seeks bounded signed-cube identities through constructions, residue filtering, and cube-pair search. It completes and checks each candidate assignment; unsupported inputs or failed searches return UNKNOWN, with no UNSAT certificate path. DPR

2.00 ×

2.00 ×

1/2

Arithmetic shortcut search is active: a narrow circuit detector extracts three operand words and searches bounded cube identities using near-cancellation, modular roots, factorization, and divisors. Candidates are completed and clauses checked; rejected or exhausted searches return UNKNOWN, with no UNSAT or DPR certificate path. VeriPB

1.00 ×

1.00 ×

0/2

Modular arithmetic search leads: the solver recognizes a narrow cube circuit and tests bounded constructions with factorization, roots, CRT, and discriminant filters. If this fails, DPLL can emit a checked SAT assignment or a VeriPB RUP UNSAT log; unsupported cases return UNKNOWN, without proof verification.

summle

174

26 GBD instances · 6 validation · baseline 6/6

The benchmark asks whether a bounded read-once arithmetic expression can reach a requested target from the multiset 1, 2, 2, 4, 4, 8, 25, 100 using permitted binary operations. CNF bit-blasts operation choices, operands, intermediate values, and target conditions. Proof format

No verif.

+ verif.

Solved

GRAT

0.431 ×

0.436 ×

5/6

Subset DP, meet-in-the-middle checks, and cached models or witnesses guide expression construction on three layouts while watched-literal CDCL completes or searches the CNF. Satisfying assignments are checked against original clauses and emitted as DIMACS models; no independently checkable UNSAT certificate is produced, so refutations or unsupported inputs return UNKNOWN. DPR

1,050 ×

447 ×

6/6

Expression-table lookup and subset dynamic programming build a candidate operation sequence whose circuit is forced by unit assumptions before watched-literal CDCL completes remaining variables. Failed construction invokes generic CDCL with learned clauses logged as witness-free additions for external checking; unsupported shapes return UNKNOWN, and direct SAT construction emits no proof stream. VeriPB

815 ×

64.8 ×

6/6

Target extraction from repeated signatures and multiset dynamic programming construct a bounded arithmetic witness; forced selector and opcode bits are completed by watched-literal CDCL. Checked SAT assignments are emitted, but exhaustion, conflicts, failed validation, or unrecognized layouts return UNKNOWN because no UNSAT or VeriPB certificate path exists.

The Case for Automated Hyperspecialization: Evidence from SAT

tensors

175

test-configuration

177

40 GBD instances · 8 validation · baseline 2/8

44 GBD instances · 9 validation · baseline 4/9

The benchmark asks whether a three-way tensor over GF(2) is the XOR of at most R rank-one outer products. CNF encodings use AND gates for factor products, XOR constraints for tensor entries, and ordering constraints for interchangeable terms.

The benchmark asks whether n configurations of 46 options can obey fixed incompatibilities and implications while covering every compatible option pair. Its CNF has per-row variables and pair-conjunction variables; variants require each pair to be absent from a row or impose lexicographic row ordering.

Proof format

No verif.

+ verif.

Solved

GRAT

2.14 ×

2.14 ×

5/8

It searches for rank-one decompositions by choosing shared matrix updates that reduce residual slice ranks, with randomized diversification and pivot-style fallback. It completes ordering and auxiliary variables with dynamic programming and DPLL, checks every clause, and emits a SAT assignment; unsupported structure or failed search returns UNKNOWN, with no UNSAT certificate path. DPR

2.15 ×

2.15 ×

5/8

It builds a GF(2) slice-space span by choosing minimum-rank residuals across slicing modes, then factors residual matrices into rank-one terms with binary Gaussian elimination. It reconstructs and orders factors, propagates and checks the clauses, and emits a SAT assignment; bounded search failure or unsupported encodings return UNKNOWN, with no UNSAT certificate path. VeriPB

2.04 ×

2.04 ×

5/8

It contracts XOR components into tensor cells, then uses bit-mask tabu/min-conflicts seeding, slice switches, and Gaussian-elimination rank factoring to construct factors. Comparator completion uses subset DP and DPLL before a clause-checked SAT assignment; failed construction invokes internal CDCL that can emit an UNSAT RUP proof for external checking, while unsupported or out-of-regime encodings return UNKNOWN.

termination-analysis

No verif.

+ verif.

Solved

7.81·104 ×

814 ×

9/9

Randomized maximum-coverage search over enumerated maximal cliques seeks a fixed-size row cover, repairs negative-coverage variants locally, and clause-checks the resulting complete SAT assignment. On smaller recognized instances, an obstruction generator emits ASCII DRAT additions for external elaboration and checking; recognition or search failure returns UNKNOWN without a general SAT fallback. DPR

5.05·104 ×

394 ×

9/9

An exact recognizer for a capped fixed layout dispatches to a compatibility-cover construction, pads and orders rows as needed, evaluates auxiliaries, and checks every clause before returning a SAT assignment. Smaller instances use a hard-coded obstruction and DPR steps with witness-free clause additions; recognition or proof-generation failure returns UNKNOWN without a general fallback. VeriPB

8.25·104 ×

2,080 ×

9/9

Arithmetic-layout recognition selects a hard-coded compatibility cover for SAT, pads and orders rows as required, reconstructs auxiliaries, and checks every input clause before releasing the assignment. Smaller recognized instances use fooling-set counting with premise lookup to emit VeriPB RUP and cutting-plane proofs; unrecognized inputs or proof-generation failure return UNKNOWN without a SAT fallback.

testpattern-generation

The benchmark asks whether a Boolean assignment satisfies all clauses encoding termination-ordering constraints. Auxiliary variables may represent circuit gate outputs in a Tseitin-style CNF, while the implementations recognize only selected layouts, not every possible encoding. Proof format

No verif.

+ verif.

Solved

GRAT

0.389 ×

0.373 ×

9/14

Bounded randomized local search first targets recognized, short Tseitin-like CNF and accepts only a fully checked model; other accepted inputs undergo propagation, bounded elimination, probing, and watched-literal CDCL. SAT models are clause-checked; UNSAT uses a textual DRAT log of preprocessing, probing, and learned clauses, while unsupported inputs or missing proof output return UNKNOWN. 0.267 ×

0.316 ×

6/14

AND/parity DAG reconstruction drives bounded gate repair with phases, random flips, and restarts; large circuits skip repair, and failure falls back to CDCL, not UNSAT. Inputs outside unit-rich, mostly binary/ternary shapes return UNKNOWN; SAT models are checked, the CDCL fallback is uncapped, and UNSAT proof logs use root conflicts and learned, potentially witness-free clause additions. VeriPB

GRAT

176

68 GBD instances · 14 validation · baseline 13/14

DPR

Proof format

0.350 ×

0.413 ×

8/14

Conservative Tseitin recognition and bounded elimination of low-occurrence outputs preserve semantic variables before watched-literal CDCL; recognized one-hot cases receive a bounded random-walk SAT probe, while unmatched layouts return UNKNOWN. A non-logging scout precedes a fresh VeriPB RUP run when needed; checked SAT assignments are accepted, but failed or incomplete UNSAT certification returns UNKNOWN.

178

52 GBD instances · 11 validation · baseline 11/11

The benchmark asks whether a Boolean assignment satisfies a CNF encoding of a test-pattern or fault-detection condition. In circuit-oriented instances, variables represent signals and clauses impose the relevant constraints, so SAT gives a test pattern and UNSAT rules out one for that encoding. Proof format

GRAT

No verif.

+ verif.

Solved

2.11·10 −4 ×

5.22·10 −4 ×

10/11

Statistical recognition selects two bounded-width ATPG-miter signatures without reconstructing topology. Watched-literal CDCL uses VSIDS, first-UIP learning, and restarts, falling back to a descending scan when the activity heap has no candidate; checked SAT models are printed, while UNSAT emits DRAT additions and an empty clause, and unsupported or failed cases return UNKNOWN. DPR

2.11·10 −4 ×

5.21·10 −4 ×

10/11

Syntactic envelope recognition accepts only CNF formulas matching bounded-width statistical signatures; it does not recover gates, boundaries, or a netlist. Watched-literal CDCL supplies first-UIP learning, activity-based branching, and restarts; checked SAT assignments are printed, while UNSAT produces learned clauses and an empty clause for DRAT elaboration, with other outcomes UNKNOWN. VeriPB

6.53 ×

1.30 ×

11/11

Narrow recognition precedes parity contraction of adjacent equivalence or inversion pairs and signed clause rewriting. Watched-literal CDCL uses first-UIP learning and restarts; SAT checks a reconstructed assignment, while UNSAT is certified by VeriPB RUP steps for learned and contracted clauses plus an empty conclusion; rejected, inconsistent, or proof-failed cases return UNKNOWN.

Harrison Green, Claire Le Goues, and Fraser Brown

theorem-proving

179

tree-decomposition

181

8 GBD instances · 2 validation · baseline 2/2

24 GBD instances · 5 validation · baseline 2/5

These benchmarks encode grid-based pebbling contradictions: they ask whether local Boolean rules can assign consistent states to grid vertices. Source or root conditions, predecessor implications, and a designated sink condition are expressed with small CNF gadgets, often using two Boolean variables per vertex.

The benchmark asks whether an undirected graph has an elimination ordering of bounded width that obeys precedence constraints. Structured CNF variables and clauses encode the graph, ordering, fill edges, and width bound, with supported layouts varying by solver. Proof format

No verif.

+ verif.

Solved

Proof format

No verif.

+ verif.

Solved

GRAT

1.72 ×

1.71 ×

33.4 ×

8.85 ×

2/2

3/5

GRAT

Support-hash grouping and DSU recovery reconstruct the paired-variable DAG, validating a unique source, acyclicity, and topological order without checking rectangular dimensions. The recovered order drives DRAT additions for external elaboration, including boundary and interior clauses, a sink unit, and the empty clause; unrecognized inputs return UNKNOWN, with no SAT-model path. DPR

14.4 ×

0.474 ×

2/2

Exact structural recognition recovers paired variables, predecessor DAG, and grid order from the two-variable OR-substituted rectangular grid form. It emits witness-free DPR additions for the recovered order, using head clauses and bridge clauses to derive the sink contradiction; unrecognized inputs and failures return UNKNOWN, with no SAT-search or model path. VeriPB

12.2 ×

2.93 ×

2/2

Two-pass counting sort reconstructs variable pairs and four-clause grid gadgets, then checks dimensions, polarity, clause ownership, and an acyclic root-reachable order. The solver emits a cutting-planes certificate with one derivation per non-root vertex and a final sink contradiction, while malformed or unrecognized inputs return UNKNOWN and no SAT-solving fallback is implemented.

Structural recovery and mixed elimination search distinguish this solver: it recognizes only direct or scrambled encodings, combines randomized min-fill/min-degree and annealing local search, and uses exact DFS on smaller graphs. A found order gets Horn or renamable-Horn completion and clause verification; failures return UNKNOWN, with no UNSAT certificate path. DPR

1.13 ×

1.13 ×

2/5

Structural recognition and randomized elimination search lead the solver: it recovers a narrowly supported graph encoding, tries greedy min-fill/min-degree orders, and uses capped beam search as fallback. Successful orders receive residual Horn completion and clause verification; unsupported or failed cases return UNKNOWN, and no UNSAT certificate path exists. VeriPB

0.872 ×

0.872 ×

1/5

Precedence-aware randomized min-fill/min-degree search leads the solver: it recognizes only the canonical layout, tries greedy orders, and falls back to heuristic repair stages. Successful orders receive Horn completion and clause verification before SAT assignment output; optional CDCL can generate RUP steps and an UNSAT conclusion, while normal failure returns UNKNOWN.

trigonometric-functions

182

6 GBD instances · 2 validation · baseline 1/2

tournament

180

16 GBD instances · 4 validation · baseline 0/4

These instances ask whether a partially fixed tournament can be completed so every seven-vertex subset contains a cyclic triangle, not a transitive ordering. The CNF uses edge orientations and triangle witnesses, with clauses linking witnesses to edges and requiring each seven-set to be covered. Proof format

No verif.

+ verif.

Solved

GRAT

1.00 ×

1.00 ×

0/4

Score-based local search is the active technique: it flips free edge orientations while incrementally tracking transitive seven-sets, using tabu moves, dynamic weights, focused randomness, and restarts. The narrow encoding is checked before search; found assignments are clause-checked, unsupported inputs return UNKNOWN, and the SAT-only path has no UNSAT certificate or exhaustion result. DPR

1.00 ×

1.00 ×

0/4

Restricted tournament blow-up construction is tried first, duplicating core vertices under capacity constraints and rejecting remaining transitive seven-sets. Only recognized encodings can yield checked SAT assignments; unsupported inputs return UNKNOWN, and no UNSAT derivation or exhaustion result is implemented. VeriPB

1.00 ×

1.00 ×

0/4

Stochastic min-conflict edge-flip search is active, maintaining selector values and seven-set violation counts while sampling unsatisfied sets and applying scored flips, noise, and restarts. A specialized prefix-completion branch checks constrained six-set-free extensions; found SAT models are checked, unsupported encodings return UNKNOWN, and no active UNSAT certificate path exists.

These instances ask whether Boolean assignments to input and internal variables can satisfy every clause in a circuit encoding of a trigonometric computation or relation. The benchmarks use gate-based bit-vector or Tseitin CNFs, while recognizers accept only restricted structural layouts. Proof format

No verif.

+ verif.

Solved

GRAT

1.01 ×

1.07 ×

1/2

Structural case splitting canonicalizes small functional gate groups in the piecewise-lines layout, deriving constant and equivalence clauses before watched-literal CDCL; Taylor-layout instances skip this preprocessing. UNSAT branches log DRAT clause additions, while checked SAT models are returned and unsupported inputs or a configured conflict cutoff yield UNKNOWN. DPR

2.34 ×

2.10 ×

2/2

Implication-graph equivalence substitution and blocked-clause elimination simplify recognized gate-heavy CNFs, with bounded variable elimination attempted only on larger accepted instances. First-UIP learning logs RUP clauses and a terminal empty clause for UNSAT; checked SAT assignments are returned, while unsupported inputs or failed checks yield UNKNOWN. VeriPB

0.953 ×

0.596 ×

1/2

A conservative topology heuristic identifies likely source variables and prefers them on activity ties, but all accepted CNFs use watched-literal, first-UIP CDCL search; no circuit-specific propagator or dispatch is active. UNSAT produces VeriPB RUP records, an empty-clause RUP, and an UNSAT conclusion, while checked SAT models are returned; malformed input or failed checks yields UNKNOWN.

The Case for Automated Hyperspecialization: Evidence from SAT

tseitin-formulas

183

unknown

185

197 GBD instances · 68 validation · baseline 51/68

118 GBD instances · 24 validation · baseline 12/24

These benchmarks ask whether a CNF encoding of XOR constraints has a Boolean assignment satisfying every clause. A common encoding groups clauses that forbid one parity of variables, with graph-shaped cases representing variables as edges between constraint vertices and vertex charges as parity requirements.

These benchmarks ask whether Boolean variables can satisfy a structured encoding of a finite combinatorial construction. Clauses represent choices and enforce exact-one, uniqueness, incidence, adjacency, forbidden-combination, or auxiliary-consistency conditions. Proof format

No verif.

+ verif.

Solved

Proof format

No verif.

+ verif.

Solved

GRAT

0.609 ×

0.611 ×

3.04 ×

2.77 ×

62/68

3/24

GRAT

Exact parity-block recognition drives graph charge-parity and spanning-forest solving, with bounded packed GF(2) elimination for non-graph systems and a narrow repair for selected quadratic encodings. Checked SAT assignments are reconstructed from recovered equations; selected graph contradictions receive attempted DRAT/extended-resolution proofs, while unsupported, oversized, or unproved cases, including dense inconsistencies, return UNKNOWN. DPR

0.506 ×

0.513 ×

32/68

Complete parity-block decoding drives GF(2) solving: graph systems use spanning forests, with bounded elimination and narrow repaired or quadratic-macro fallbacks. SAT assignments are clause-checked; supported UNSAT cases use graph or macro contractions to emit witness-free DPR clause additions, while rejected proof plans or failed stochastic searches return UNKNOWN, never establishing UNSAT. VeriPB

6.07 ×

4.45 ×

65/68

ANF/Mobius recovery yields XOR equations over signed variables or two-literal terms; graph components use charge propagation and tree reconstruction, while dense elimination produces SAT models only. Clause-checked models are reconstructed physically; graph contradictions can receive VeriPB proofs, while a bounded clique fallback emits cutting-planes proofs and non-graph inconsistencies return UNKNOWN without UNSAT certification.

uniform-random

184

The benchmark asks whether a Boolean assignment can satisfy every clause in a signed-literal CNF formula. Inputs are restricted random-width-shaped DIMACS formulas with controlled clause widths and related structural checks, which describe the encoding rather than establish statistical uniformity. Proof format

No verif.

+ verif.

Solved

GRAT

1.28 ×

1.27 ×

332/809

Pure 2-CNF instances use an SCC contradiction-path procedure; other accepted width classes first receive belief- or survey-propagation-seeded incremental local search with make/break and age-aware choices. On failure, watched-literal 1-UIP CDCL continues to a checked model or conflict, logging textual DRAT clauses for UNSAT; malformed or out-of-scope inputs return UNKNOWN. 1.09 ×

1.09 ×

241/809

Focused incremental local search uses clause-satisfaction state, break/make scores, and width-dependent probSAT- or Novelty-style choices. Only narrow complete-route instances enter watched-literal first-UIP CDCL; it can print a checked model or log witness-free DPR additions for UNSAT, while other local-search misses and out-of-scope inputs return UNKNOWN. VeriPB

DPR

0.933 ×

0.856 ×

11/24

Bounded structural recognizers try cyclic-isotopy and Dancing Links for direct orthogonal-Latin encodings, residue or TabuCol searches for five-colorings, and component CSP searches for signed-selector cases. Failures fall through to CDCL with failed-literal probing; SAT models are checked against input, while UNSAT logging yields textual DPR and proof or model-check failures return UNKNOWN. VeriPB

0.888 ×

0.696 ×

10/24

Choice-block recognition uses rare-polarity orientations, bitsets, singleton rules, MRV branching, and bounded DFS; narrow Latin search tries affine constructions before MRV. Failed attempts fall back to CDCL; SAT models are checked, pigeonhole cases emit direct PB proofs, while CDCL UNSAT traces end in RUP contradiction and proof-generation failures return UNKNOWN.

unknown-cases

186

20 GBD instances · 4 validation · baseline 4/4

4,043 GBD instances · 809 validation · baseline 201/809

DPR

Bit-mask recognition and bounded Dancing Links search target orthogonal Latin-square encodings; fixed balanced-design and bit-vector Latin constructions seed CDCL, with CDCL handling other inputs. SAT assignments are rechecked, and UNSAT paths log derived clauses and an empty clause for external DRAT elaboration; failed direct orthogonal-Latin searches or budget exhaustion return UNKNOWN.

1.09 ×

1.08 ×

241/809

Incremental probSAT-style local search maintains make/break, unsatisfied-clause, and sole-satisfier state, with weighted and deterministic seed strategies. A watched-literal CDCL fallback is available only for bounded instances; it can yield a checked SAT assignment or log RUP proof steps for UNSAT, while larger or failed searches can end UNKNOWN without an UNSAT certificate.

The benchmark asks whether a Boolean assignment satisfies every clause of a CNF formula. Structured instances may encode selections that cover requirements while forbidding combinations, or represent circuits and order constraints; these are recurring structures, not a universal encoding. Proof format

No verif.

+ verif.

Solved

GRAT

0.131 ×

0.131 ×

1/4

Four-incidence positive-cover recognition leads the search, followed by orientation, bounded graph/circuit and stochastic model attempts, all accepted only after checking a full CNF assignment. When CDCL proves UNSAT, it logs learned clauses and a final empty clause for external checking; failed structural or internal searches do not establish UNSAT and can return UNKNOWN. DPR

0.106 ×

0.106 ×

0/4

Polarity-normalized cover recognition leads to bitset propagation, randomized repair, local search, and minimum-coverage DFS, with fixed-layout circuit and repeated-key paths as bounded fallbacks. SAT candidates are checked against the original CNF; only monotone DFS within its 64-variable limit emits decision nogoods and may report UNSAT, while other failures or large-cover exhaustion return UNKNOWN. VeriPB

0.106 ×

0.106 ×

0/4

Monotone cover/forbidden-set search uses bitset filtering and smallest-uncovered branching, with local search and watched-literal CDCL fallbacks, accepting only clause-checked assignments. The complete monotone and CDCL paths replay learned or decision-nogood clauses as RUP steps with a final empty constraint; heuristic failures provide no proof, and unsupported or exhausted cases return UNKNOWN.

Harrison Green, Claire Le Goues, and Fraser Brown

waerden

187

xor_op

189

110 GBD instances · 22 validation · baseline 3/22

3 GBD instances · 1 validation · baseline 1/1

The benchmark asks whether positions 1 through N can be two-coloured so one colour avoids every three-term arithmetic progression and hits every K-term progression. Boolean variables encode these conditions as positive and negative CNF clauses; some instances use folded representations.

The benchmark asks whether a directed relation on vertices can be a strict ordering: it is asymmetric and transitive, yet every vertex must have a predecessor among three designated choices. CNF uses two XOR-related variables per relation atom and expands the logical constraints into clauses.

Proof format

No verif.

+ verif.

Solved

Proof format

No verif.

+ verif.

Solved

GRAT

1.04 ×

1.04 ×

4/22

GRAT

8,210 ×

375 ×

1/1

Bounded weighted local search first seeks a checked SAT assignment when a positive rank-2 clause exists, then complete DFS blocks positive edges and branches on the tightest unhit negative edge. Proof logging records parent, leaf, exclusion, and deletion clauses for external elaboration; restricted pure-polarity shapes or proof inconsistencies can yield UNKNOWN. DPR

1.01 ×

1.02 ×

3/22

Incremental weighted local search leads the SAT attempt, using breakout weights, flip scores, and age tie-breaking. Failure falls back to complete DFS with negative-clause branching; unsupported structure yields UNKNOWN, while checked SAT assignments are returned and DFS may log witness-free conflict clauses without an implemented UNSAT certification. VeriPB

0.944 ×

0.912 ×

2/22

Two-watched-literal CDCL is the active engine, using activity-based branching, 1-UIP conflict minimization, restarts, failed-literal probing, and learned-clause reduction. It checks SAT assignments, returns UNKNOWN for unsupported structure, and on UNSAT attempts to serialize learned clauses, deletions, and a final contradiction as VeriPB proof text, without invoking a proof verifier.

xor-chain

188

66 GBD instances · 14 validation · baseline 11/14

These instances ask whether a Boolean assignment satisfies a connected system of parity equations. A typical CNF encoding uses four clauses for each three-variable XOR and two for each two-variable XOR; related instances use sequential counters and graph constraints. Proof format

No verif.

+ verif.

Solved

GRAT

1.00 ×

1.00 ×

11/14

XOR-gate recognition and Gaussian elimination handle supported parity systems; persistent XOR-tree logging, fresh definitions, and RUP/resolution projections support UNSAT. On recognition failure, a triangle-graph fallback uses blossom matching and sequential counters for clause-checked SAT models; strengthened forms skip this path, and unsupported or failed dispatch returns UNKNOWN. DPR

1.00 ×

1.00 ×

11/14

Degree-two incidence-graph recognition selects only supported, clause-complete binary and ternary parity groups in a connected structure. Bit-packed GF(2) elimination yields SAT assignments checked against every original clause, while UNSAT uses spanning-tree XOR expressions, fresh definitions, chord cancellation, and local RUP derivations; unsupported or failed cases return UNKNOWN. VeriPB

3.00 ×

2.99 ×

13/14

Direct XOR recognition uses BFS cycle reductions for parity instances, clause-checking SAT assignments and emitting VeriPB derivations for UNSAT contradictions; strengthened forms have only the latter path and may return UNKNOWN. Counter variants reconstruct triangles and use blossom matching for SAT or pseudo-Boolean proofs for conflicting bounds; matching failure or unsupported cases return UNKNOWN.

Exact width-2 XOR decoding with fallbacks builds an embedded-gadget proof: it names atoms, derives ordering clauses, and greedily eliminates vertices to an empty clause. It has no SAT fallback: only recognized three-predecessor schemas are handled, unsupported inputs or failures return UNKNOWN, and DRAT is sent for external elaboration and checking. DPR

2.34·104 ×

128 ×

1/1

Greedy elimination over predecessor bitsets replaces each removed predecessor by its predecessors, emits canonicalization additions for used XOR gadgets, and ends with an empty clause when successful. Recognition covers only the exact bounded three-predecessor macro layout; no general SAT or SAT-witness path exists, and unsupported or failed constructions return UNKNOWN. VeriPB

1,660 ×

121 ×

1/1

XOR-block recovery followed by in-process CDCL searches the recovered logical clauses rather than the expanded raw formula. On UNSAT it writes a VeriPB proof with XOR reductions and RUP steps, but does not invoke a verifier; recognition mismatches, satisfiable outcomes, or proof failures return UNKNOWN.

The Case for Automated Hyperspecialization: Evidence from SAT

K

Extended: Solver taxonomy

We used GPT-6 Astra (xhigh) and GPT-5.6 Luna (high) to iteratively generate and refine a taxonomy for seven different questions about solver behavior. Here we describe our methodology and potential limitations (§K.1). We then present the full taxonomy (§K.3). K.1

Methodology

Given a query about a specific aspect of solver behavior, such as “What solving strategies are used in the solvers?”, we first built out a suitable taxonomy using GPT-6 Astra (xhigh) in Codex, providing the full suite of 567 generated solvers and their source code. We iteratively refined the taxonomy to ensure it covered all variation and properly distinguished between closely related concepts (both LLM-mediated and with some human feedback). Given a frozen taxonomy, we then tasked GPT-5.6 Luna (high) with annotating each solver independently for each query and set of labels. We report results for all non-zero counts. A given solver only contributes once to a given label (even if it implements several variants) but each solver can contribute to multiple labels. The full listing of frozen taxonomies is provided in §K.3. K.2

Potential limitations

LLMs are fallible and our taxonomies may be incomplete or inaccurate. For the use of these taxonomies in our paper (observing broad variation patterns) we believe this is an acceptable risk. The exact counts of solvers for each label and the specific techniques listed, if changed slightly, would not alter the conclusions we draw. K.3

Full taxonomy labels

Æ Note that the content in this section is LLMgenerated. These full taxonomy labels were generated with GPT-6 Astra (xhigh) with some human feedback. K.3.1 Solving strategies. Whole solving procedures, including systematic search, local search, algebraic and structural reasoning, inference, and direct answer construction. The algorithm audit counts default-reachable solving roles, excluding procedures used only for preprocessing or proof processing. Such procedures can still qualify in the other studies. For a local-search flip, make counts clauses newly satisfied, and break counts clauses made unsatisfied. Structural construction / theorem. Recognized instance structure supports a witness construction or a combinatorial contradiction, such as a pigeonhole or counting argument. The method must perform actual input-dependent reasoning. CDCL. Conflict-driven clause learning combines Boolean decisions and propagation with conflict analysis, then reuses

learned clauses during search. Watched literals, activity scores, or proof logging alone do not qualify. Greedy construction. An assignment or domain solution is built progressively without systematic backtracking or iterative local repair. Trivial fixed-phase candidates qualify only when actually tested as answer attempts. Exhaustive enumeration. Assignments or combinations are explicitly enumerated and checked, possibly in parallel using machine-word bits. Ordinary CDCL branching is not counted as a separate enumeration algorithm. Native-domain local search. Neighborhood moves repair recovered domain objects, such as changing a color or swapping permutation entries. The moves act on these objects beyond individual CNF variable flips. Constraint backtracking. Search assigns recovered finitedomain variables or domain objects, such as colors or permutations, and backtracks with constraint filtering. The search operates on domain constraints beyond ordinary Boolean DPLL. Circuit / equivalence reasoning. Recovered gate semantics, circuit simulation, equivalence reasoning, or circuit rewriting solve or reduce the problem. Identifying gates only to choose variable phases is insufficient. Cardinality / PB reasoning. Dedicated counting or bound propagation, cutting-plane deductions, or counting contradictions reason beyond ordinary CNF propagation. Merely printing a VeriPB certificate is insufficient. WalkSAT-style search. Search selects an unsatisfied clause, then flips one of its variables using low-break scores or noisy/random choices. An occasional random step within a different algorithm is insufficient. Other graph algorithms. A dedicated graph procedure, such as reachability, bipartiteness testing, or isomorphism reasoning, decides or constrains feasibility. Graph storage and ordering alone are insufficient. DPLL (no clause learning). Davis-Putnam-LogemannLoveland search branches on Boolean variables, propagates consequences, and backtracks without learning conflict clauses. A CDCL implementation alone does not receive this additional label. Weighted local search. Search adapts clause or constraint penalties to redirect subsequent local moves. Static weights and ordinary break counts alone are insufficient. Arithmetic / number theory. Arithmetic reasoning, factoring, modular algebra, or integer-feasibility reasoning directly constructs an answer or proves impossibility. Parsing numeric features alone is excluded. Propagation-only solving. A distinct answer procedure propagates consequences to a fixed point and completes an assignment deterministically without branching search. Propagation used only inside another solving algorithm is excluded.

Harrison Green, Claire Le Goues, and Fraser Brown

ProbSAT-style search. Candidate flips are sampled with explicit nonuniform probabilities derived from make/break scores, such as inverse powers of break counts. Uniform random noise alone is excluded. Tabu search / TabuSAT. A tabu list or tenure explicitly prohibits certain moves, potentially allowing exceptions when a move is sufficiently promising. Age-based tie-breaking or a preference against recent flips alone is insufficient. XOR / GF(2) elimination. Parity equations are solved or reduced through row operations or Gaussian elimination over GF(2), the two-element field with arithmetic modulo two. Recognizing exclusive-or (XOR) gates without elimination is insufficient. State-space search. Search traverses explicit domain states, using methods such as breadth-first or depth-first search, iterative deepening, heuristic search, bidirectional search, or beam search. Matching / flow. A recovered graph is solved using matching, augmenting paths, maximum flow, or a Hall witness: a set with too few neighbors to support the required matching. Resolution / elimination. The solver explicitly generates resolvents, eliminates variables, or repeatedly applies resolution toward saturation. Conflict-clause learning within CDCL alone does not receive this additional label. Meet-in-the-middle search. Search enumerates two partial solution spaces and joins compatible halves using a table, hashing, or sorting. Looking up a stored complete answer is excluded. Dynamic programming. A recurrence combines and caches subproblem results, for example over subsets, trees, or paths. Maintaining an ordinary assignment trail is insufficient. Simulated annealing. Search accepts worsening moves according to a temperature-dependent rule. Merely changing a random-noise rate over time is insufficient. Stored witness / proof / answer. A reachable answer route loads, embeds, or retrieves a precomputed assignment, proof, or answer. Runtime memoization and newly derived witnesses are excluded. Branch and bound. Search computes objective bounds and uses them to prune subproblems that cannot improve the result. Feasibility backtracking without such bounds is excluded. Large-neighborhood search. Search relaxes or removes a substantial part of a candidate assignment and reoptimizes the resulting neighborhood, often with an exact subsolver. Other algorithm. The source implements a concrete algorithm outside the named categories. The audit must give its precise name and explain the missing taxonomy category. Look-ahead SAT search. Tentative assignments and propagation repeatedly estimate the consequences of choices and

guide branching. Failed-literal probing used only in preprocessing is excluded. Belief / survey propagation. The solver iteratively exchanges probabilistic messages on a factor graph, potentially fixing variables according to the resulting beliefs or surveys. Input-feature answer rule. Names, hashes, sizes, or other input signatures select a prescribed answer or candidate without a demonstrated general structural derivation. This records the decision mechanism without judging its correctness. GSAT-style greedy flips. Improving variable flips are chosen using make/break scores over all variables or a maintained set of promising variables. Selection only within one unsatisfied clause is insufficient. Independent random sampling. The solver generates and checks fresh, independent complete candidates. This is distinct from neighborhood search and restarts of that search. 2-SAT / implication solving. A dedicated procedure decides 2-SAT, whose clauses have at most two literals, using strongly connected components of an implication graph or an equivalent method. Binary-clause propagation alone is excluded. Horn / tractable fragments. A dedicated algorithm solves a recognized tractable fragment, such as Horn formulas (at most one positive literal per clause) or dual-Horn formulas (at most one negative literal). Ordinary unit propagation alone is insufficient. Novelty-style search. Search selects between the best and second-best candidates using avoidance of the most recently flipped variable and a noise rule. Generic age-based tiebreaking alone is excluded. Local search (unspecified). Candidate assignments are repeatedly repaired through neighborhood moves. This residual category applies only when no more specific local-search label fits the component. Population / evolutionary. Search maintains and evolves multiple candidates through selection, mutation, crossover, or other population-based optimization. Repeated independent starts alone are excluded. Decision diagrams / compilation. The solver builds or manipulates a compiled Boolean representation, such as a binary decision diagram or zero-suppressed decision diagram, to decide satisfiability or construct a solution. Continuous optimization. The solver optimizes a realvalued relaxation using gradients or continuous dynamics, then decodes Boolean assignments from the result. Configuration-checking search. A variable becomes eligible to flip again when a neighboring variable or constraint configuration changes. Timestamps alone do not establish this mechanism.

The Case for Automated Hyperspecialization: Evidence from SAT

External engine, algorithm opaque. The solver invokes an external solving engine whose algorithm cannot be established from the supplied implementation. A known engine name alone does not establish CDCL, and proof checkers are excluded. LP / MIP optimization. An explicit linear-programming or mixed-integer-programming solver or relaxation uses optimization machinery. Native Boolean branching alone is insufficient. Pure random walk. A distinct search repeatedly makes unguided or clause-focused random variable flips. Random steps within a scored, noisy WalkSAT procedure alone do not receive this additional label. SMT / theory solving. Satisfiability modulo theories combines Boolean reasoning with a theory-specific decision procedure. Arithmetic preprocessing alone is excluded. K.3.2 Systems-level optimizations. Concrete mechanisms in input/output, CPU execution, memory management, incremental work, and parallelism. Conventional implementations can qualify; novelty and measured benefit are not requirements. Compiler / ISA flags. The default build explicitly requests compiler optimization, instruction-set features, link-time optimization, or related tuning. Such flags alone do not establish vectorized execution or a measured speedup. Preallocation / reuse. Capacity reservation, scratch buffers, free lists, pools, or arenas reduce repeated memory allocation. Ordinary container construction alone is excluded. Incremental state updates. Derived quantities, such as clause counts, move scores, or circuit values, are updated from local changes instead of recomputed globally. Assignment-trail writes alone are excluded. Flat / compact storage. Flat arenas, offsets, packed records, separate arrays for record fields, or compressed adjacency organize data more compactly than fragmented or pointerheavy storage. An ordinary array alone is insufficient. Watched / blocking literals. Watched literals, cached satisfying literals, binary implication paths, or occurrence lists restrict propagation to potentially affected clauses instead of scanning every clause. Specialized parsing. A custom byte or integer parser, bulk buffering, or explicit fast-I/O configuration avoids formatted input overhead. Ordinary token splitting, formatted scanning, or stream extraction alone is excluded. Scalar word parallelism. Multiple logical values are packed into scalar machine words or bitsets and processed together using bitwise operations. Ordinary scalar flags and randomnumber arithmetic are excluded. Lazy / deferred work. Generation stamps, lazy invalidation, deferred deletion or compaction, or delayed recomputation avoid recurring bulk work. Simply omitting cleanup is insufficient.

Buffered proof output. Proof bytes or lines are explicitly batched, or a larger output buffer is installed, to reduce output calls. Ordinary proof logging alone is excluded. Bit-count / scan intrinsics. Reachable computation uses dedicated bit-counting, bit-scanning, or rotation intrinsics or equivalent standard functions. Plain bitwise AND or XOR alone is insufficient. Reduced proof-output work. A concrete mechanism reuses serialized proof fragments, suppresses redundant emitted steps, or avoids repeated serialization. A mathematically short proof alone does not qualify. Specialized hot loops. Manual unrolling or dedicated fixedsize or short-clause kernels specialize a recurring computation. A constant loop bound or an inline annotation alone is insufficient. Lookup tables. Precomputed operation or transition tables replace recurring computation, for example with truth tables or byte population-count tables. Stored complete witnesses and ordinary search memoization are excluded. Memory-mapped input. The solver maps the input file into memory and accesses it through that mapping. An unused mapping wrapper or header alone is insufficient. Branchless updates. Explicit masks, arithmetic, or table lookups replace conditional selection in a recurring computation. A conditional expression alone does not establish branchless execution. OS memory hints. Explicit operating-system advice requests access-pattern, paging, or read-ahead behavior. Memory mapping alone is excluded; these hints are distinct from CPU prefetch instructions. Compact proof encoding. Proof records use an implemented compact serialization, such as binary records, variable-length integers, deltas, or compression. A proofformat name or ordinary text formatting alone is insufficient. Alignment / cache blocking. Explicit alignment, padding, tiling, or blocking organizes data or computation for locality. Contiguous storage alone is insufficient. Explicit SIMD. Packed vector operations execute through intrinsics, vector types, assembly, or a SIMD-specific library, such as SSE, AVX2, AVX-512, or NEON. Compiler flags and scalar bitsets alone are excluded. Branch-prediction hints. Explicit likelihood annotations or profile hints tell the compiler which branch outcome is expected. Ordinary conditional statements do not qualify. CPU software prefetch. The solver issues a software prefetch intrinsic or instruction for future memory accesses. Ordinary loads, buffering, capacity reservation, and operating-system advice are separate. GPU / accelerator kernels. A reachable route launches computation on a GPU or another accelerator. Headers, build support, and unused code alone are excluded.

Harrison Green, Claire Le Goues, and Fraser Brown

Worker processes. Multiple solver-worker processes execute concurrently. A single subprocess, an execution wrapper, or process-cleanup code alone is insufficient. Worker threads. A reachable route starts concurrent threads or a parallel region for solver computation. Threadlibrary dependencies and sequential restarts alone are excluded.

K.3.3 Preprocessing and inprocessing. Transformations of a remaining problem before search or while search is in progress. Before only and During only indicate exclusive observed stages; Before + during means both occur in the same solver. Stage unclear denotes unresolved timing. Ordinary per-decision propagation and general conflict learning alone are excluded. Root unit simplification. Root-level unit propagation fixes forced assignments and removes satisfied clauses or literals to simplify the residual problem before search starts or resumes. Ordinary propagation after each decision alone is excluded. Clause cleanup. Duplicate literals, tautologies, or duplicate clauses are removed, including normalization that enables these reductions. Sorting only to recognize structure is excluded. Cone / core reduction. A relevant cone, core, or proper subset is extracted and solved or transformed, with a mechanism transferring the result to the original problem. Merely ignoring clauses is insufficient. Structural strengthening. Implied constraints derived from recovered structure strengthen a residual problem, for example equivalences between duplicate gates or orderrelation closure. Ordinary conflict learning alone is excluded. Other preprocessing. A concrete transformation of the remaining problem falls outside the named categories. The annotation must identify the transformation and explain the taxonomy gap; ordinary search machinery is excluded. Resolution elimination. A variable is removed by resolving clauses containing its positive and negative literals and retaining a satisfiability-equivalent residual problem, possibly subject to growth limits. Branching and Gaussian elimination are separate. Self-subsuming resolution. Resolution produces a clause that subsumes a parent and therefore shortens it. Ordinary subsumption and general conflict learning alone are excluded; qualifying resolution-based learned-clause minimization may overlap with the search-tuning category. Gate / circuit reduction. Recovered gates are constantfolded, substituted, rewritten, or functionally eliminated to reduce a residual search problem. Recognition or standalone circuit simulation alone is insufficient.

Domain filtering. Explicit finite domains or their supported values are pruned using domain-consistency reasoning before search starts or resumes. Ordinary Boolean unit propagation alone is excluded. Failed-literal probing. Tentative literal assignments and propagation derive forced values or other simplifications. Probing used only to score the next branching decision belongs to search tuning. Cardinality / PB reduction. Recovered counting or pseudoBoolean constraints are simplified through bound reasoning, forced values, compression, or rewriting before further search. A direct counting refutation or the proof format alone is excluded. XOR / algebraic reduction. Parity or algebraic equations are eliminated or substituted to reduce a problem passed to further solving. A complete Gaussian solver without a residual reduction role is excluded. Literal substitution. Equivalent or complementary literals are merged and substituted to rewrite the residual problem, for example using implication-graph components. Detecting an equivalence only to solve directly is excluded. Problem decomposition. The original instance is partitioned into independent or explicitly coordinated smaller subproblems for separate solving. Ordinary branches of one search do not qualify. Variable compaction. Active residual variables or literals are renumbered to shrink solver data structures after or during reduction. Relabeling only for recognition or input permutation is excluded. Symmetry breaking. Canonical restrictions remove equivalent assignments, or a verified symmetry rewrites the problem, reducing the remaining search space. Recognizing a symmetric family alone is insufficient. Pure-literal elimination. A variable occurring with only one polarity is assigned or eliminated in a satisfiabilitypreserving simplification. Merely preferring that polarity during search is excluded. Clause subsumption. A clause is removed because another clause contains a subset of its literals. Removing only duplicate clauses belongs to clause cleanup. Clause vivification. Sequential literal assumptions and propagation establish that a clause can be shortened or is redundant. Generic failed-literal probing without clause shrinking is separate. Blocked-clause elimination. A clause is removed when some literal blocks it: every resolvent on that literal is tautological. Explicit generalized blocked or set-blocked rules also qualify. Implication reduction. Implied binary edges or clauses are removed by transitive reduction or another demonstrated redundancy test. Building an implication graph alone is insufficient.

The Case for Automated Hyperspecialization: Evidence from SAT

K.3.4 Search tuning. Policies controlling decisions, learning, restarts, local moves, diversification, and effort. The categories describe choices within search, rather than naming whole solving algorithms. Static / occurrence ordering. An explicit fixed variable or value order, occurrence score, or polarity count guides decisions. Ordinary container iteration without a decision policy is excluded. Scheduled restarts. A search restarts after a fixed budget or according to a predetermined schedule, such as Luby or geometric intervals. Switching to a different algorithm alone is separate. Input-dependent parameters. Input features determine search budgets, thresholds, heuristic parameters, or modes. Dispatch between unrelated algorithms without tuning a search is separate. Bounded search effort. An individual search attempt has an explicit step, conflict, flip, node, or time limit. The experiment’s overall timeout alone is excluded. Biased / seeded phases. Preferred Boolean values are initialized or steered using polarity, recovered structure, a candidate assignment, or a configured sign. Reusing past assignments alone belongs to phase saving. Conflict-activity branching. Conflicts or learned clauses update variable priorities, as in variable-state independent decaying sum (VSIDS) and its exponential variant. Static occurrence scores and clause activity alone are excluded. Phase saving. Previously assigned Boolean values are remembered and reused in later decisions. A single fixed preferred polarity is insufficient. Domain-aware branching. Semantic quantities such as remaining-domain size, graph degree, constraint slack, or gate role guide decisions. Generic literal occurrence ordering is separate. Learned-clause retention. A deliberate policy ranks, retains, or deletes learned clauses using quality, activity, size, recency, or related measures. Proof-format deletion bookkeeping alone is excluded. Diverse search starts. Search is deliberately retried with distinct seeds, shuffled orders, perturbations, or initial assignments. One random initial state alone is insufficient. Learned-clause minimization. Redundant literals are removed from a learned conflict clause beyond ordinary conflict analysis. General preprocessing vivification is separate. Random branch / tie choices. Randomness selects search variables or values or resolves decision ties. Random-number tests, identifiers, and independent initial sampling alone are excluded. Make / break scoring. Local-search moves are scored by their changes to satisfied or violated clauses or domain constraints. Make counts newly satisfied constraints; break counts newly violated ones.

Noisy move selection. Local moves use explicit noise or a nonuniform probability rule, as in WalkSAT or ProbSAT. Randomized branching in complete search is separate. LBD / glue quality. The number of distinct decision levels represented in a learned clause, called literal-block distance (LBD) or glue, is computed and used to manage clauses or control search. Look-ahead branch scoring. Tentative assignments, propagation, or predicted downstream reductions score the next branching decision. Probing solely to derive forced literals is preprocessing. Stagnation control. Detected lack of progress changes search behavior, perturbs a candidate, or stops an attempt. A fixed stopping budget alone is insufficient. Tabu / recency rules. Tabu tenure, move age, youngestvariable avoidance, or configuration eligibility constrain or prioritize moves. This broad category does not imply that every annotated solver implements TabuSAT. Adaptive constraint weights. Penalties on clauses or constraints change during search to redirect local moves. Fixed weights supplied by the problem alone are excluded. Temperature / noise control. Temperature, noise probability, or another exploration-strength parameter changes over time or with stagnation. A fixed noise probability alone is insufficient. Adaptive restarts. Observed search behavior, such as clause quality or stagnation, triggers restarts or changes their timing. A fixed restart schedule alone is excluded. K.3.5 Structure recovered from CNF. Semantic objects recovered or exploited beyond generic CNF clauses and literals. Recovering a representation does not by itself imply that it simplifies the problem or tolerates every encoding variant. Semantic components / cones. Meaningful components, cones of influence, layers, or decompositions are recovered and exploited. Ordinary recursion or adjacency storage alone is insufficient. Problem graphs. A native problem graph is reconstructed whose vertices and edges represent domain objects, for example in coloring or reachability. Generic clause-variable incidence and watch graphs are excluded. Gates / Boolean circuits. Gate semantics or a Boolean circuit are recovered from clauses and used in solving, reduction, or certification. This includes comparison circuits and AND-inverter graphs; suggestive variable names alone are insufficient. Finite-domain variables. Boolean literals are grouped into multi-valued variables and explicit domain constraints used in native search or propagation. Single Boolean variables or incidental clause groups are excluded. Cardinality / PB constraints. Counting constraints, weighted Boolean sums, or inequalities are recovered from

Harrison Green, Claire Le Goues, and Fraser Brown

CNF and used in reasoning. Coefficients appearing only in proof output are insufficient. Permutations / orders. A permutation, ordering, sequence, or positional combinatorial object is recovered and manipulated directly. Sorting input clauses or variables alone is excluded. Arithmetic / modular structure. Integer, modular, polynomial, subset-sum, or number-theoretic quantities and equations are recovered beyond generic Boolean counting constraints. Numeric parsing alone is insufficient. Other recovered structure. A concrete semantic representation falls outside the named categories. The annotation must identify it and explain the taxonomy gap; generic clause storage is excluded. Equivalence / implication. Semantic equivalence classes, strongly connected implication components, or a tractable implication fragment form a higher-level representation. Generic implication storage for clause learning alone is excluded. Matching / assignment. Matching, bipartite assignment, permutation-matrix, flow, or Hall-set roles and constraints are recovered. A generic graph representation alone is insufficient. XOR / parity equations. Recovered parity constraints are represented as equations or an explicit parity system used in reasoning or search. Treating XOR only as a circuit gate or bitwise operation is insufficient. Symmetry / group structure. A concrete symmetry, orbit, automorphism, or group action is identified and used in reasoning or construction. Generic sorting or a symmetry claim alone is insufficient. State-transition systems. Automata, transitions, traces, planning states, or temporal steps are recovered and used for reasoning in a state space. A purely combinational circuit alone is insufficient. Geometry / packing. Coordinates, shapes, intervals, placements, physical-board adjacency, or geometric constraints are recovered for solving. A matrix used only for storage is excluded. Statistical models. A factor/message, correlation, spectral, or other statistical representation supports inference or candidate generation. Ordinary clause incidence without statistical state is excluded. K.3.6 Combining solving methods. Ways to select, combine, or exchange information between solving procedures. Outer coordination is distinguished from policies internal to one search loop. Structural dispatch. Recognized problem structure or encoding selects a solving procedure. Syntax checking or changing one numerical heuristic alone is insufficient.

Sequential methods. Two or more different solving procedures are attempted in sequence. Repeated seeds, recognition variants, and proof emission after solving alone are excluded. Candidate / phase transfer. An assignment, incumbent, preferred phase, or partial candidate produced by one method is passed to another. Independent guesses and final result printing are excluded. Constraint / bound transfer. Constraints, algebraic relations, bounds, or deductions pass between distinct reasoning methods. Ordinary clause sharing within one engine alone is insufficient. Separate SAT / UNSAT routes. Distinct specialized procedures find satisfying assignments and derive unsatisfiability certificates within one submission. The two outcomes of one generic clause-learning engine alone are insufficient. Size / feature dispatch. Counts, estimated complexity, or other coarse instance features select different solving procedures. Changing a limit within one unchanged algorithm is insufficient. General SAT fallback. An unsupported, unsuccessful, or exhausted specialized route can invoke a generic SAT engine. The label records a reachable fallback and does not guarantee completeness. Repeated configurations. An outer coordinator runs repeated starts or configurations of a method and handles their success or failure. Restarts wholly internal to an unchanged search loop are excluded. Budgeted handoff. A per-method resource limit triggers another method or an escalation. A final timeout without an alternative, or a restart of the same method, is excluded. Residual subsolver calls. A higher-level procedure conditionally or repeatedly invokes a distinct solver on residual instances or subproblems and uses its answers. Ordinary recursive branches are excluded. Unlogged / certifying stages. A search or check runs without proof logging, then is repeated or transfers information into a certifying run. Identical search replay is not required; ordinary proof-format conversion alone is excluded. Refinement loop. Distinct procedures, or abstraction, candidate, and refutation steps, alternate in a feedback loop. One-way preprocessing followed by a single solve is insufficient. Adaptive effort allocation. Runtime progress or method outcomes change how effort is allocated among methods. A fixed sequence of predetermined budgets is excluded. Parallel portfolio. Distinct solving procedures run concurrently and their results are coordinated. Sequential calls and independent starts without different methods are excluded. K.3.7 Encoding assumptions and safeguards. Inputrepresentation assumptions, recognition mechanisms, scope guards, and answer checks. Recognition and safeguards can

The Case for Automated Hyperspecialization: Evidence from SAT

coexist with encoding dependence; these labels are not an empirical test of correctness or shuffle robustness. Unsupported-case rejection. A failed semantic or scope check intentionally reports UNKNOWN or rejects the specialized route. Generic timeouts and file-opening errors are excluded. Structural validity guards. A recovered template, invariant, or encoding condition is checked before specialization, with explicit rejection or fallback on failure. Basic input syntax checking alone is insufficient. Relational recognition. Semantic roles are inferred from incidence, gate or graph relations, or constraint patterns. This can coexist with variable-ID assumptions and does not establish invariance under input permutation. Exact templates / signatures. Fixed formula sizes, exact templates, hashes, or names select a semantic route or candidate beyond generic complexity thresholds. Signature matching is distinct from subsequent validation and is not an integrity judgment. Original-CNF model checks. At least one reachable SAT success route checks a completed candidate against every original clause. The label does not imply that all success routes do so; reduced-only and sampled checks are excluded. Multiple encoding variants. Multiple encodings, orientations, variable layouts, or clause-pattern variants of the same semantic task are explicitly handled. Supporting different input sizes alone is insufficient. Variable-ID / layout reliance. Arithmetic on variable identifiers, fixed offsets, or contiguous identifier blocks assigns semantic roles. Generic array indexing, renumbering, and literal encoding alone are excluded. Model reconstruction. Eliminated variables or a domain solution are lifted to a complete original-variable assignment using dependencies, elimination records, or an explicit mapping. Printing an already complete assignment is excluded. Canonicalized recognition. Signs, permutations, clause or literal order, or equivalent encodings are canonicalized to recognize structure across variants. Sorting only for deduplication or fast lookup is excluded. Rejected-candidate fallback. Failure of candidate, domain, or model validation resumes search or invokes another method instead of reporting SAT. Parsing or recognition failure alone is separate. Domain / reduced checks. A reconstructed domain solution or reduced or partial candidate is explicitly validated beyond ordinary search transitions, potentially before lifting to CNF. This is distinct from checking every original clause. Clause-order reliance. Input clause positions, adjacency, or order determine semantic roles. Ordinary traversal and sorting into order-independent groups are excluded.

Sampled / fingerprint checks. Bounded simulation, sampling, or fingerprints screen semantic recognition, equivalence, or candidate properties. Later exact validation is separate; generic random search and hash-table equality are excluded. Stored candidates / proofs. A reachable result route retrieves an embedded or external precomputed model, proof, or answer. Constants for a general construction, localoperation tables, and runtime memoization are excluded. Certificate / RUP checks. The supplied run path checks its own certificate or replays proof obligations before accepting an UNSAT certificate. This includes reverse-unitpropagation (RUP) obligation checks and does not imply an independent checker of a complete serialized proof.

Related documents

Record · ID 919496 · SHA-256 827f0c19330265af
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.