arXiv:2609.11180v1 [cs.AI] 10 Sep 2026
SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics Qibai Chen
Zeming Liu
Independent Researcher United States [email protected]
Brown University United States zeming [email protected]
A free, deterministic resolver answers every such query in microseconds, and the same model that fails the corner case aces the basic form. This is not an isolated glitch. Across three packaging ecosystems and six frontier models, we find that LLMs systematically and predictably misapply versionconstraint rules whose simple forms they handle perfectly. LLM coding agents resolve dependency versions all the time. They read a lockfile, judge whether an upgrade is admissible under ˆ1.2.3, pick a version that satisfies >=2.0,<3, or reason about whether a security fix falls inside a declared range. These are high-frequency operations in every package manager (npm, pip, Cargo, and beyond) and inside every coding assistant that touches dependencies. When an agent gets version semantics wrong, the consequence is not a cosmetic typo: it is an incorrect dependency decision—an unsafe version range, a missed patch, or a build that silently resolves to the wrong release. These are supply-chain-relevant errors: empirically, a substantial fraction of dependent releases are broken by upstream version/breaking-change issues [6]. Despite this, the semantics of version-constraint resolution—as opposed to the upstream task of inferring which dependencies a repository needs—has never been benchmarked. We close that gap with SemVerBench. The task is deliberately narrow and unambiguous: given a version V and a constraint C in ecosystem E ∈ {npm, PEP 440, Cargo}, decide whether V satisfies C. Every item has a unique, machine-checkable answer. This narrowness is a feature: it removes the confounds of open-ended generation and lets us attribute every error to constraint semantics rather than to formatting, taste, or ambiguity. The narrow surface has broad impact, because version resolution is universal. A first reaction might be that an LLM simply “cannot compute” such things. The evidence points the other way, along three converging lines. First, models are at or near ceiling on the basic forms of every grammar (plain membership, epochs, exact pins), so they demonstrably command the underlying concepts; the failures are concentrated in corner cases of rules whose simple forms they handle perfectly. Second, a light, hedged hint—“I think the answer is . . . ”, not a forceful instruction—is enough to correct most previously-failed items,
Abstract—Large language model (LLM) coding agents resolve dependency versions constantly—deciding whether an installed version satisfies a declared constraint such as ˆ1.2.3 or >=2.0,<3—yet their grasp of version-constraint semantics has never been measured directly. We introduce SemVerBench, the first benchmark of LLM version-constraint resolution semantics across three ecosystems (npm, PEP 440, Cargo). SemVerBench contains 240 machine-checkable items with unique answers, built author-neutrally from four balanced sources (each ecosystem’s official test suite plus three frontier LLM proposers), and labeled by a non-circular two-implementation oracle rather than by humans. Evaluating a six-model panel, we find systematic, predictable per-mechanism blind spots rather than diffuse noise: a partial-comparator carry rule (>1.2 ≡ >=1.3.0) traps every model on Cargo (all near 60%), and although PEP 440 standard prefix matching is universal, on zero-pad/post-release corner cases GPT-5.1 collapses (0/26 zero-pad cases) while Claude stays at 97–100% (verified on a 67-item oracle-validated set). Opus significantly outperforms all other models, and Sonnet outperforms the OpenAI models (McNemar). Crucially, the failures look more consistent with an activation/application gap than with a knowledge gap: injecting the relevant rule recovers most errors and a light correct hint corrects models that already failed, whereas interval decomposition does not help, and models are at ceiling on the basic forms of the same rules. An authorstratified analysis finds no statistically significant self-favoritism (the largest own-item advantage, GPT-4.1 +10 points, is n.s.; the only significant author effect is Gemini scoring lower on its own items), and the blind spots hold on official-suite and cross-family items. Because the task is verifiable and a free, 100%-correct resolver exists, tool delegation reaches ≈100%. The takeaway: coding agents should delegate version resolution to a resolver, not reason about versions in-head. Index Terms—large language models, semantic versioning, dependency resolution, software engineering benchmarks, tool use, knowledge activation
I. I NTRODUCTION GPT-5.1, a frontier reasoning model, handles the standard PEP 440 prefix form ==2.* perfectly—and yet, asked the zero-pad corner case “does 2 satisfy ==2.0.*?” (PEP 440 says yes, because it zero-pads the release segment), it is wrong on all 26 such cases we tested, while Claude scores 97–100%. Accepted at the 38th IEEE International Conference on Tools with Artificial Intelligence (ICTAI 2026). This is the authors’ preprint version.
1
which is more consistent with nudging a latent competence than with teaching a missing one. Third, supplying the rule itself helps, whereas supplying additional reasoning structure (interval decomposition) does not: the missing ingredient is the availability of the right rule in context, not more reasoning steps. Together these are more consistent with an activation/application gap than with a capability or knowledge gap— a reading that aligns with work on model self-knowledge [10], though, as we discuss in Section VI, our design cannot fully separate activating latent knowledge from supplying missing knowledge. Our findings are also structured, not diffuse. Difficulty concentrates in specific mechanisms: basic membership, epochs, and plain ranges are at or near ceiling for all models, while a handful of mechanisms account for nearly all errors. Two stand out. First, a universal cross-ecosystem trap: a partial comparator rounds up, so >1.2 means >=1.3.0 and >1 means >=2.0.0; on Cargo, every model—including the strongest— falls to roughly 60% on this mechanism. Second, a vendor-specific blind spot: PEP 440 standard prefix matching (==N.*) is universal—all six models handle it—but its corner cases (zero-pad versions like 2 against ==2.0.*, and postreleases) form a sharp vendor-specific blind spot that GPT5.1 fails systematically while Claude does not. These are not random; they are reproducible signatures of how each model family internalizes (or fails to internalize) a documented rule. Contributions. 1) SemVerBench, the first benchmark of LLM versionconstraint / semver resolution semantics, spanning npm, PEP 440, and Cargo: 240 machine-checkable items with unique answers, constructed author-neutrally and labeled by a non-circular two-implementation oracle. 2) A characterization of systematic, per-mechanism blind spots across six frontier models, including a universal cross-ecosystem partial-comparator trap and a sharp, vendor-specific prefix-match failure, with statistical significance (McNemar) and an author-stratified analysis showing no significant item-proposer self-favoritism. 3) An actionable diagnosis: the errors look like activation/application gaps, recoverable by rule-injection and corrected by light accurate hints, and eliminated by delegating to the resolver (tool use reaches ≈100%).
was understood; DependEval [2] probes dependency-graph understanding (which modules depend on which), not the semantics of a single specifier. Both ask which dependencies are needed; neither probes whether a version correctly resolves against a constraint—our downstream, orthogonal task. To our knowledge no prior benchmark targets version/semver resolution semantics at this granularity. Constraint-following and constraint-reasoning benchmarks. A parallel line studies whether LLMs follow constraints. CFBench [3] covers natural-language instruction constraints (format, length, content) judged on free-form text, not a formal predicate with a single truth value; ConstraintBench [4] (Gurobi-verified, operations-research) concerns search over a feasible region, not specifier grammar; and ACS [5] uses an LLM as a soft judge of constraint-satisfaction in open-ended answers—the opposite of our deterministic oracle. SemVerBench instead targets a small, formally specified, verifiable semantics with a unique answer per item, labeled by code. Premise sensitivity and sycophancy. Our controlled falsepremise study (Section V-A) is adjacent to the sycophancy literature [9], which shows models capitulating to user opinions on subjective or free-form prompts. Our setting differs in a key respect: the task is objective and verifiable, so a “hint” is either correct or incorrect against ground truth. We find an asymmetric effect—models adopt correct hints strongly and resist weak incorrect ones—which is consistent with the activation/application account rather than with indiscriminate sycophancy. III. S EM V ER B ENCH A. Task An item is a triple (E, V, C): ecosystem, version, and constraint. The model must decide membership—does V satisfy C?, which we write V |= C—and end its response with ANSWER: true or ANSWER: false. Every item is a Boolean membership query: there is exactly one correct answer, computable by a deterministic resolver, and no item asks for free-form text, ranking, or generation. Scoring is exact match against the oracle label after extracting the final verdict. Parse failures (a response with no extractable verdict) are tallied separately rather than scored as wrong, so that comprehension is not conflated with output formatting. This minimal, verifiable framing is what lets us attribute every error to constraint semantics: a wrong answer cannot be excused as taste, ambiguity, or an unlucky phrasing, because a one-line resolver call settles the item.
II. R ELATED W ORK LLMs on software-engineering tasks. LLMs are increasingly evaluated on real software-engineering work—SWE-bench [7] measures whether they resolve actual GitHub issues endto-end—but such benchmarks bundle many sub-competences behind a single pass/fail outcome. SemVerBench instead isolates one specific, high-frequency, verifiable sub-competence: deciding version membership. Dependency benchmarks for LLMs. Recent work evaluates LLMs on dependency inference—which packages a repository requires. DI-BENCH [1] scores whether a generated manifest lets a repository build and pass its tests, a coarse end-to-end signal that does not isolate whether any version constraint
B. Ecosystems and Their Grammars SemVerBench covers three ecosystems whose constraint grammars overlap superficially but differ in subtle, consequential ways. npm uses the node-semver range grammar: a range is a disjunction of hyphen ranges and comparator sets, with sugar for caret (ˆ), tilde (˜), and x-ranges; the tricky parts are caret/tilde behavior when leading components are zero, partial comparators, and an opt-in prerelease policy.
2
Python PEP 440 defines version specifiers over a structured version (epoch, release, pre/post/dev segments) with operators including ==, !=, ˜= (compatible release), and the prefixmatch form ==N.*; its subtleties include epoch ordering, prefix matching on the release segment, and the rule that ordered comparators exclude adjacent pre/post releases of the boundary version. Cargo’s VersionReq resembles npm semantically (default caret behavior) but its own treatment of partial comparators is the dominant source of difficulty in our data. These three grammars give a small but genuinely diverse semantic surface.
the same prompt and topics—no vendor is favored—and any residual selection bias is measured post-hoc (Section IV-E); we do not claim the prompt is content-free. This four-way design diversifies difficulty (official suites anchor documented behavior; LLM proposers surface corner cases) and, since every item’s proposer is recorded, enables the conflict-ofinterest analysis in Section IV-E. Beyond these main 240, we additionally build a 67-item author-constructed, oraclevalidated prefix-match set used solely for the focused validation in Finding 3 (detailed there); it is not part of the main accuracy numbers.
C. The Two-Axis Taxonomy
E. Ground Truth: A Non-Circular Two-Implementation Oracle
Each item is tagged along two orthogonal axes. The syntax axis records the surface grammatical form of the specifier— caret/tilde sugar, a bounded comparator (e.g., >=1.2,<2.0), a wildcard/prefix form (e.g., 2.*), or an exact pin—i.e., how the constraint is written. The mechanism axis records the underlying semantic rule the item exercises— e.g., the partial-comparator carry or ordered-comparison exclusion— i.e., what resolving the constraint requires the model to know. The two are orthogonal: the same syntactic form can exercise different mechanisms, and the same mechanism can surface under different syntax. We organize the entire analysis by the mechanism axis, because the mechanism is what isolates the semantic difficulty; the syntax axis is recorded for coverage and balance but is not the unit of analysis, since two items that look identical syntactically may be trivial or hard depending on which rule they hinge on. Key mechanisms include: • Caret/tilde and magic-zero (ˆ, ˜): compatibility ranges whose semantics change when leading components are zero (e.g., ˆ0.2.3). • Partial-comparator carry: a partial comparator rounds up, so >1.2 ≡ >=1.3.0 and >1 ≡ >=2.0.0. • Prerelease admission: when a prerelease such as 1.2.0-rc.1 is or is not admitted by a range. • PEP 440 ordered-comparison exclusion: >V does not match a post-release of V , and <V does not match a prerelease of V . • Epoch: PEP 440 epoch-prefixed versions (e.g., 1!1.0). • Prefix match: PEP 440 ==N.* / !=N.* matching on the release-segment prefix.
We use no human labels. Each item’s answer is computed by two independent reference implementations per ecosystem and retained only when they agree: • npm: node-semver × semantic_version. • PEP 440: packaging × pep440_rs. • Cargo: the Rust semver crate × semantic_version. The oracle is non-circular in two distinct senses. First, the two implementations per ecosystem are independently authored projects, written in different languages and by different maintainers—for PEP 440, packaging is the reference Python implementation while pep440_rs is an independent Rust reimplementation, so an item is labeled only when two codebases that share no implementation lineage agree on its answer; this filters out single-library quirks and bugs. Second, the oracle is independent of the subjects: no panel model labels any item, and three of the four item sources are the panel models themselves only as proposers, never as graders. We additionally apply, for PEP 440, a prereleasepolicy-robustness filter: items whose answer would depend on a resolver’s prerelease-inclusion policy (rather than on the specifier semantics themselves) are dropped, so that every retained PEP 440 item has a policy-independent answer and cannot be “wrong” merely because of a defensible policy choice. F. Evaluation Protocol We evaluate a six-model panel: Claude-Opus, ClaudeSonnet, GPT-5.1, GPT-4.1, GPT-4o, and Gemini-2.5-Pro. All models are queried at temperature 0 with three independent runs over the full 240 items. At temperature 0 the three runs are nearly identical (run-to-run spread 0.4–3.3 points across models), so across-run variance understates true uncertainty; we therefore report Wilson 95% confidence intervals that treat each model’s 3-run mean accuracy as a binomial proportion over the n = 240 items, reflecting benchmark sampling uncertainty. For paired model comparisons we use the McNemar exact-binomial test, computed on per-item correctness after taking the majority vote across the three runs (n = 240 paired items). There were 0 API errors and a parsefailure rate of 0.30%. The 240 items span 9 mechanisms and 18 syntax tags, and are mildly imbalanced toward “false”: 103 true (42.9%) and 137 false (57.1%). A majority-class
D. Author-Neutral Construction SemVerBench has 240 items, 80 per ecosystem, drawn from four author-neutral sources, 20 items each. The first source is each ecosystem’s official test suite—the regression tests shipped with the canonical resolver implementation. The other three sources are frontier LLM proposers (Claude, GPT, Gemini), each given the identical prompt: propose challenging version/constraint membership items within fixed perecosystem topic areas (e.g., prerelease admission, caret/tilde ranges), optimizing for difficulty. A proposer’s suggested answer is discarded; every item is independently labeled by the oracle, so no model labels its own items. We call the construction author-neutral in that all three proposers receive
3
p < 0.05–.01), so the gap is not a confidence-interval artifact. Notably, competence here does not track model recency: the newest OpenAI model, GPT-5.1 (80.4%), ranks numerically below GPT-4.1 (82.2%), though this within-vendor gap is not statistically significant (the OpenAI models and Gemini are mutually n.s. by McNemar, with overlapping Wilson CIs). We therefore read it only as evidence that recency does not predict competence on this task, not as a reliable ordering. We cannot attribute the cross-vendor gap to specific training differences—the models are closed—and we do not need to: the point is not “Claude wins” but that the spread itself is unjustified, because the task is machine-checkable and a free, 100%-correct resolver exists, so any in-head error rate—for any vendor—is avoidable.
TABLE I OVERALL ACCURACY WITH W ILSON 95% CI S TREATING THE 3- RUN MEAN ACCURACY AS A BINOMIAL PROPORTION OVER n = 240 ITEMS . M C N EMAR : O PUS > ALL (p < 0.01–.001); S ONNET > EACH O PENAI MODEL (p < 0.05–.01); S ONNET VS . G EMINI N . S .; O PENAI & G EMINI MUTUALLY N . S . Model Claude-Opus Claude-Sonnet Gemini-2.5-Pro GPT-4.1 GPT-5.1 GPT-4o
Acc. (%)
95% CI
90.6 86.8 82.9 82.2 80.4 80.1
[86.0, 93.5] [81.8, 90.4] [77.6, 87.2] [76.7, 86.4] [74.9, 84.9] [74.5, 84.6]
predictor that always answers “false” therefore scores 57.1% (per-ecosystem false rates: npm 52.5%, PEP 440 58.8%, Cargo 60.0%); this is the chance floor against which all accuracies below should be read.
B. Finding 2: Partial-comparator carry is a universal trap A partial comparator rounds up: >1.2 is equivalent to >=1.3.0, and >1 is equivalent to >=2.0.0. On npm, all six models land at exactly 80% on this mechanism (16 of 20 items correct under 3-run majority, for every model). That crossmodel uniformity is itself informative: when six independently trained systems converge on the identical score, the difficulty is a property of the mechanism, not of any one model. On Cargo the effect is more severe: every model falls to roughly 60%— Opus 63%, Sonnet 59%, and the remaining four in the 62– 63% band. This is the clearest cross-ecosystem, cross-vendor blind spot in the benchmark, and notably it does not spare the strongest model. A representative failure: 1.0.5 |= >1 is False (because >1 carries to >=2.0.0), yet models answer True, apparently treating >1 as >1.0.0. The rule is short and documented; the models simply do not apply it.
IV. F INDINGS Table I reports overall accuracy. Claude-Opus leads at 90.6 [86.0, 93.5], followed by Claude-Sonnet at 86.8 [81.8, 90.4]; the OpenAI models and Gemini cluster between 80% and 83%. By McNemar, Opus is significantly better than every other model; Sonnet is significantly better than each OpenAI model; Sonnet versus Gemini is not significant; and the OpenAI models and Gemini are mutually not significant. Marginal CIs may overlap, but our significance claims rest on the paired McNemar test (per-item majority vote), which controls for item difficulty and detects differences that overlapping marginal CIs can mask. These overall accuracies of 80–91% sit well above the 57.1% majority-class floor (and far above a 50% coin flip), confirming that the models have genuine competence on the bulk of items—the benchmark is not being solved by exploiting label imbalance. This is only the aggregate picture, however: as Findings 2–4 show, several individual blind-spot buckets do not clear the majority-class floor at all. Indeed, aggregate accuracy hides the structure that is the point of this paper. Table II breaks accuracy down per (ecosystem × mechanism). Plain membership, epochs, and basic ranges sit at or near 100% for all models; the difficulty is concentrated in a few mechanisms. Crucially, the headline blind-spot buckets fall close to the chance floor: Cargo partialcarry at ≈60% is barely above the 57.1% overall majority rate, and at or below Cargo’s own 60.0% majority rate, so on this mechanism the models retain almost no genuine signal beyond guessing; Claude-Sonnet’s 59% there is at or below Cargo’s 60.0% majority rate—essentially no signal beyond guessing. The PEP 440 ordered-exclusion bucket (60–70%) is similarly close to chance for the weaker models. We discuss each headline blind spot in turn.
C. Finding 3: Prefix-match corner cases are a sharp, vendorspecific blind spot PEP 440 prefix-matching corner cases produce the single most dramatic failure in SemVerBench. On the main benchmark’s prefix-match bucket (n = 7 items), only Claude handles it (Opus 95%, Sonnet 90%); everyone else drops sharply—GPT-5.1 to 19%, GPT-4.1 to 38%, GPT-4o and Gemini to 52%. These 7 items are predominantly corner-type (zero-pad, post-release, multi-segment), which is exactly why GPT-5.1 scores ≈19% here—consistent with the 21% it scores on the focused corner set below; the ordinary standard form is at 100% for all models. Focused validation. The main prefix-match bucket is small (n = 7) yet carries the paper’s most dramatic claim, so it warrants targeted scrutiny; the other small main-benchmark buckets (e.g., cargo.bare_caret n = 6, npm.plain n = 9, pep440.plain n = 6) are near-ceiling— conservative, high-accuracy results that do not hinge on a few items— so they do not warrant the same treatment. We therefore built a 67-item prefix-matching validation set, separate from the main 240. The items are author-constructed: we programmatically enumerated and hand-designed (version, constraint) pairs covering two regimes—standard prefix (the version has at least the prefix’s release segments and no post-release)
A. Finding 1: A cross-vendor gap Claude models (Opus, Sonnet) significantly outperform the OpenAI and Google models by paired McNemar testing (Opus > all others p < 0.01–.001; Sonnet > each OpenAI model
4
TABLE II P ER -( ECOSYSTEM × MECHANISM ) ACCURACY (%, POOLED OVER 3 RUNS ). n IS THE NUMBER OF ITEMS IN THE BUCKET ( SUMMING TO 240). B OLD ROWS ARE THE HEADLINE BLIND SPOTS . C ELL SHADING : GREEN ≥90, YELLOW 75–89, ORANGE 50–74, RED <50. ecosystem.mechanism
Opus
Sonnet
GPT-5.1
GPT-4.1
GPT-4o
Gemini
n
cargo.bare caret cargo.magic zero cargo.partial carry cargo.plain cargo.prerelease npm.magic zero npm.partial carry npm.plain npm.prerelease pep440.compatible pep440.epoch pep440.ordered exclusion pep440.plain pep440.prefix match
100 100 63 100 85 100 80 100 100 100 100 84 94 95
100 100 59 100 94 96 80 100 98 89 100 65 100 90
94 93 62 92 75 96 80 100 88 75 90 70 100 19
100 94 63 100 88 95 80 100 86 86 95 63 94 38
100 87 63 97 81 81 80 100 85 83 97 62 94 52
83 94 63 92 96 98 80 100 86 92 98 60 100 52
6 18 27 13 16 19 20 9 32 12 20 35 6 7
and corner cases (zero-pad, post-release, multi-segment). They are deliberately not produced by the main benchmark’s LLM proposers, since a targeted validation is meant to probe specific sub-cases. Labeling uses the identical oracle (packaging under three prerelease policies, kept only if robust across all three, × pep440_rs consensus). The set comprises 34 standard + 33 corner = 67 items; the 33 corner items are 26 zero-pad/post-release “true” cases (e.g., 2 |= ==2.0.*, 2.0.post1 |= ==2.0.*) plus 7 boundary “false” cases. The blind spot is not a small-sample artifact (Table III). On the 34 standard items all six models score 100%: every model handles the ordinary ==2.* form. The blind spot lives entirely in the corner cases, where it is sharp and vendor-specific: GPT-5.1 scores 21%—and 0/26 on the zero-pad/post-release “true” subset specifically—while Claude stays at 97–100% and GPT-4o/4.1/Gemini sit in a 58–67% middle band. The diagnosis is concrete: GPT-5.1 systematically treats a version with fewer release segments than the prefix (2 vs. ==2.0.*, which PEP 440 zero-pads) and post-release versions as nonmatching, judging every such case False. Because the rule is well-specified, this is a vendor-specific internalization gap, not a hardness property of the mechanism: one vendor has clearly learned the corner cases and the others have not.
TABLE III F OCUSED PREFIX - MATCH VALIDATION (%, 3- RUN MAJORITY ): 67 AUTHOR - CONSTRUCTED , ORACLE - VALIDATED PEP 440 PREFIX ITEMS , SEPARATE FROM THE MAIN 240 (34 STANDARD +33 CORNER ). T HE LAST COLUMN IS THE 26 ZERO - PAD / POST- RELEASE “ TRUE ” CASES THAT FORM A subset OF THE 33 CORNER ITEMS ( THE OTHER 7 CORNER ITEMS ARE BOUNDARY “ FALSE ” CASES ). Model Claude-Opus Claude-Sonnet GPT-5.1 GPT-4o GPT-4.1 Gemini-2.5-Pro
Standard (n=34)
Corner (n=33)
Zero-pad/post (true, ⊂33, n=26)
100 100 100 100 100 100
100 97 21 67 58 58
26/26 25/26 0/26 15/26 12/26 13/26
each model’s accuracy by item source—items proposed by its own vendor family (n = 60), items proposed by the other LLM families (n = 120, official suites excluded), and the official test suites. We test each model’s own-vs-other-LLM gap with a two-proportion test. No model shows a statistically significant own-item advantage. The gap is numerically larger for the GPT family (up to +10 points for GPT-4.1), but even that is not significant (GPT-4.1 +10, two-proportion p ≈ 0.10; all other positive gaps p > 0.27); at only n = 60 own items these effects are indistinguishable from noise. The Claude family’s gaps are small (∼+3). The only significant author effect runs the opposite way: Gemini scores significantly lower on its own items (−15 points, p = 0.02)—its own items are the hardest for it. There is thus no statistically significant self-favoritism; if anything, the one significant effect is anti-favoritism. This does not threaten the headline findings, for three reasons. (a) The blind spots—partial-comparator carry, prefix corner cases, and ordered-comparison exclusion—appear on the officialsuite and cross-family items for every model, not only on self-authored ones. (b) The leaderboard ordering is not authordriven (Opus’s own-item gap is small and n.s.). (c) One full
D. Finding 4: Ordered-comparison exclusion cracks Sonnet The PEP 440 ordered-comparison exclusion rule states that >V does not match a post-release of V (e.g., 1.0.0.post1), and <V does not match a pre-release of V . This mechanism is broadly hard—most models land in the 60–70% range— and it specifically pulls Sonnet down from its 87%-class overall standing to 65%. Opus is more robust here at 84%. A representative failure: 1.0.0.post1 |= >1.0.0 is False under PEP 440, but models commonly answer True. E. Finding 5: Conflict of interest: no significant self-favoritism Because three of our four item sources are themselves LLMs, a natural concern is self-favoritism. Table IV stratifies
5
TABLE IV C ONFLICT- OF - INTEREST CHECK : ACCURACY (%, 3- RUN POOLED ) STRATIFIED BY ITEM SOURCE . “ OWN ” = ITEMS PROPOSED BY THE MODEL’ S OWN VENDOR FAMILY (n = 60); “ OTHER ” = ITEMS PROPOSED BY THE OTHER LLM FAMILIES (n = 120, OFFICIAL SUITES excluded); “ OFFICIAL” = THE OFFICIAL TEST- SUITE SOURCES . N O OWN - VS - OTHER GAP IS SIGNIFICANT ( TWO - PROPORTION TEST ); THE LARGEST, GPT-4.1 +10, HAS p ≈ 0.10. T HE ONLY SIGNIFICANT AUTHOR EFFECT IS G EMINI SCORING lower ON ITS OWN ITEMS (p = 0.02). Model Claude-Opus Claude-Sonnet GPT-5.1 GPT-4.1 GPT-4o Gemini-2.5-Pro
own-LLM
other-LLM
official-suite
90.0 84.4 79.4 84.4 80.6 70.0
86.9 81.7 74.7 74.4 72.2 83.9
98.3 99.4 92.8 95.6 95.6 93.9
TABLE V M ITIGATION ACCURACY (%). E0 = IN - HEAD BASELINE (= TABLE I 3- RUN MEAN ). E1 = RULE - INJECTION , E2 = INTERVAL DECOMPOSITION , E3 = TOOL - USE ( OFFICIAL RESOLVER ). E1/E2/E3 ARE SINGLE - RUN . Model
E0
E1
E2
E3
Claude-Opus Claude-Sonnet GPT-5.1 GPT-4.1 GPT-4o Gemini-2.5-Pro
90.6 86.8 80.4 82.2 80.1 82.9
100 99 91 95 92 95
89 88 79 84 82 85
100 100 100 99 100 100
(≥, >, ≤, <, or exact) under the ecosystem’s rules, (2) write the version’s full precedence form, (3) check membership and answer. E3 (tool-use)—the version-resolution instance of the general idea that LLMs can be equipped to call external tools [8]—gives the model a check_constraint tool that wraps the official resolver. E1/E2/E3 are single-run conditions. Results are in Table V. Three results follow. (i) E1 rule-injection produces a large lift for every model (e.g., GPT-5.1 80.4 → 91, GPT4.1 82.2 → 95, Opus to 100%): making the relevant rule available in context recovers most errors. (ii) E2 interval decomposition produces essentially no lift (several models move within noise, some slightly down). The missing ingredient is therefore the right rule in context, not additional reasoning steps; this is inconsistent with a pure reasoningsteps deficit, though a single decomposition strategy cannot by itself exclude every reasoning-based account. (iii) E3 tooluse reaches ≈100%: a check_constraint tool wrapping the official resolver makes the task trivially correct. In only 2 of 1440 tool-condition responses did the model invoke the tool, receive the correct result, and still emit a wrong verdict. The residual failure mode is thus not computation but ignoring or misreading a correct tool output: even with delegation, an agent must actually adopt the tool’s answer. The diagnosis: an activation/application gap. These mitigation results complete the three converging lines of evidence laid out in Section I: near-ceiling basic buckets, correction by a light hint, and the rule (E1) helping while reasoning structure (E2) does not—together more consistent with an activation/application gap than a capability or knowledge gap. Two caveats apply. The light-hint lift may partly be a generic re-check trigger rather than the hint’s content (Section V-A); and because E1 and the TRUE hint both place the rule in context, we cannot fully exclude supplying-missing-knowledge (Section VI). The activation reading is the more parsimonious one, not a proof of prior possession. The conclusion is direct. A free, deterministic, 100%-correct resolver exists for this task; any in-head error rate is therefore unjustified. Coding agents should delegate version-constraint resolution to that resolver rather than reason about versions themselves. Implications for coding-agent design. The recipe follows from Table V. An agent that touches dependency versions should gate every membership or range decision through a
source—the official suites—has no panel-model stake, and the two-implementation oracle labels every item independently of any model. F. Error analysis: recurring failure modes The mistakes are not random; Table II and the worked examples point to three recurring failure modes, each a specific misreading of a rule. (M1) Partial-comparator under-carry: a partial comparator such as >1 is treated as >1.0.0 instead of the correct >=2.0.0, so a version like 1.0.5 is wrongly admitted; this is the dominant Cargo error and the source of the universal ≈60% bucket. (M2) Prefix under-match: a version with fewer release segments than the prefix (2 vs. ==2.0.*, which PEP 440 zero-pads) or a post-release (2.0.post1 vs. ==2.0.*) is wrongly judged non-matching; this drives the GPT-5.1 prefix collapse (0/26 on these cases in the focused validation), while the standard prefix form is handled correctly by all models. (M3) Boundary pre/post admission: an ordered comparator is applied without the adjacent-release exclusion, so 1.0.0.post1 is wrongly judged to satisfy >1.0.0 (and symmetrically a pre-release is wrongly admitted by <V); this mode underlies the ordered-exclusion bucket that cracks Sonnet to 65%. All three modes are corrected by ruleinjection (Section V). By ecosystem, PEP 440 is the hardest in aggregate: it concentrates two of the three modes (M2 and M3) and has the largest mechanism count, so its blind-spot buckets dominate the lower half of Table II. Cargo is the easiest except for the partial-carry trap, which is precisely why it is a clean isolation of M1: its other mechanisms sit near ceiling, leaving the carry rule as the lone signal. npm sits in between, with a stable but non-trivial partial-carry and prerelease component. V. M ITIGATION : D ELEGATE TO THE R ESOLVER Having mapped the blind spots, we ask what fixes them. We compare four conditions per model. E0 is the in-head baseline—identical to the main three-run accuracy in Table I. E1 (rule-injection) prepends the relevant resolution rule to the prompt. E2 (interval decomposition) appends a three-step instruction: (1) rewrite the constraint as explicit version intervals
6
finds no statistically significant self-favoritism (the largest own-item gap, GPT-4.1 +10, is n.s.; the only significant author effect is Gemini scoring lower on its own items), and we further mitigate with an official-suite source authored by no panel model and a non-circular two-implementation oracle that labels every item independently. The blind spots hold on official-suite and cross-family items, so they are not a self-favoritism artifact. Second, Cargo is the easiest ecosystem in aggregate, with the partial-carry trap as its principal signal. Third, SemVerBench is a single snapshot tied to each ecosystem’s official resolver library, so resolver-/versionspecific policy nuances (we filter prereleases for PEP 440) are out of scope. Fourth, the main evaluation is three-run (item-level Wilson CIs plus McNemar), whereas the mitigation (E1/E2/E3) and false-premise studies are single-run and directional. Fifth, activation versus in-context teaching: because both E1 and the TRUE hint make the rule available in context, we cannot fully separate “activating latent knowledge” from “supplying missing knowledge”; the near-ceiling basic buckets and the sufficiency of a light hint make the activation reading more parsimonious but not decisive, and directly probing whether a model can state the rule unprompted would better separate the two—future work. Sixth, per-bucket sample size: some main-benchmark buckets are small (the PEP 440 prefixmatch bucket has n = 7 items), so bucket-level numbers are indicative; we address the most consequential case with the 67item focused validation (Finding 3, Table III), and otherwise rely on three-run stability and cross-ecosystem consistency. Seventh, external validity: SemVerBench measures versionmembership decisions in isolation, and we do not directly measure how an in-head error propagates through a coding agent into a real failure. The inference chain—an erroneous membership verdict yielding an unsafe range or a missed patch, and thence a supply-chain incident—is plausible and motivated by empirical evidence that breaking changes do manifest in dependent packages [6], but that end-to-end link is not established by our experiments. Eighth, scale and prompt robustness: the benchmark is 240 items (we prioritized machine-verified correctness and mechanism coverage over raw scale). On prompt phrasing, we additionally evaluated all 240 items under two alternative phrasings differing in wording and structure; per-model accuracy shifts by at most 4.2 points across the three phrasings (GPT-5.1 by only 0.8), so the findings are not artifacts of prompt wording. Ninth, training contamination: the official test suites are public code likely present in pretraining corpora, so the 92.8–99.4% official-suite accuracies (Table IV) may partly reflect memorization rather than generalization. This does not affect the main findings— difficulty comes from the adversarial LLM-proposed items, not the official suites—but the official-suite numbers should be read as an easy anchor, not as evidence of genuine generalization. Conclusion. On a machine-checkable task with a unique answer, frontier LLMs misapply version-constraint rules whose basic forms they handle perfectly. The failures are systematic and predictable, not diffuse: a universal partial-comparator
TABLE VI C ONTROLLED FALSE - PREMISE STUDY (%, HARD 61- ITEM SUBSET, SINGLE RUN ). A WEAK HEDGED HINT IS TRUE, NEUTRAL, OR FALSE. “ BASELINE ” IS THE NO - HINT ACCURACY ON THIS SUBSET ( LOWER THAN TABLE I; DO NOT CONFLATE ). Model Claude-Opus Claude-Sonnet GPT-5.1 GPT-4.1 GPT-4o Gemini-2.5-Pro
baseline
TRUE
NEUTRAL
FALSE
80 67 39 51 44 57
93 97 56 57 52 66
75 72 49 56 54 57
80 74 57 57 56 62
resolver tool (the ecosystem’s official library)—the E3 condition, essentially free and exact—rather than emit a verdict from its own weights. When no resolver is reachable (e.g., reasoning inline in chat), the second-best mitigation is to inject the governing rule (E1). What it should not do is the current default of in-head reasoning with no scaffolding (E0), the regime in which the blind spots occur. A. Controlled false-premise sensitivity To test whether the activation/application account is sound and to probe robustness to misinformation, we run a controlled false-premise study on a hard 61-item subset. The subset is chosen by an explicit, pre-specified rule—the lowest-baseline (E0) items per mechanism, i.e., the items models most often get wrong at baseline—so that a premise has room to change the answer; it is not hand-picked. Each item is presented with a weak, hedged hint of one of three kinds plus a no-hint baseline: TRUE (a correct “I think the answer is . . . ” hint), NEUTRAL (an uninformative hint), and FALSE (an incorrect hedged hint). Because the subset is deliberately the hardest items, its baseline accuracy is lower than the full-240 numbers in Table I; it is a subset baseline and should not be conflated with Table I. Results are in Table VI. The effect is asymmetric. A TRUE hint corrects models massively—Sonnet rises from 67% to 97% (roughly 18 of 20 previously-failed items corrected)— whereas a weak FALSE hint barely misleads them (typically 0–2 flips, and FALSE accuracy is often at or above baseline). For the weakestbaseline models, the mere presence of any hedged premise— even an incorrect one—appears to trigger more careful rechecking that a weak false premise is not assertive enough to override; so part of their lift may be driven by the presence of a hint rather than its content, an observation worth dedicated future study. Models thus adopt correct hints and resist weak incorrect ones: we do not claim they are “easily fooled by false premises.” That a light correct hint corrects so many items fits the activation/application reading, and the resistance to weak false hints shows robustness to mild misinformation when the answer is checkable. VI. L IMITATIONS AND C ONCLUSION Limitations. First, the LLMs under test are also three of the four item-proposer sources, so they are simultaneously subjects and authors. The author-stratified analysis (Section IV-E)
7
carry trap drags every model to ≈60% on Cargo, PEP 440 prefix corner cases (zero-pad and post-release) collapse GPT-5.1 while Claude is near-perfect, ordered-comparison exclusion cracks even strong models, and Claude leads significantly. The diagnosis is more consistent with an activation/application gap than with a knowledge gap (Section V). Since a free, 100%correct resolver exists and tool delegation reaches ≈100%, the recommendation is unambiguous: coding agents must not resolve version constraints in-head—they should delegate to the resolver. G ENAI U SAGE D ISCLOSURE LLMs play three roles in this work. (i) Subjects: the sixmodel panel (Claude-Opus, Claude-Sonnet, GPT-5.1, GPT-4.1, GPT-4o, Gemini-2.5-Pro) is the object of study. (ii) Item proposers: three frontier LLMs (Claude, GPT, Gemini) proposed part of the main 240 items from an identical difficulty-oriented prompt (their suggested answers discarded; every item labeled by the oracle); we record each proposer and analyze selffavoritism (Section IV-E), and the 67-item focused-validation set (Finding 3) is author-constructed. (iii) Assistance: the authors used LLM-based assistants to help with coding and to prepare an initial draft; all research questions, design, analysis, verification, and conclusions are the authors’ own. A RTIFACT AVAILABILITY The benchmark data and the two-implementation groundtruth oracle (for independent label verification) are available at: https://anonymous.4open.science/r/semverbench-60BF. R EFERENCES [1] L. Zhang et al., “DI-BENCH: Benchmarking Large Language Models on Dependency Inference with Testable Repositories at Scale,” in Findings of the ACL, 2025, arXiv:2501.13699. [2] J. Du et al., “DependEval: Benchmarking LLMs for Repository Dependency Understanding,” 2025, arXiv:2503.06689. [3] T. Zhang et al., “CFBench: A Comprehensive Constraints-Following Benchmark for LLMs,” in Proc. ACL, 2025, arXiv:2408.01122. [4] J. Tso et al., “ConstraintBench: Benchmarking LLM Constraint Reasoning on Direct Optimization,” 2026, arXiv:2602.22465. [5] L. Madmoni et al., “The Ability of Large Language Models to Evaluate Constraint-satisfaction in Agent Responses to Open-ended Requests,” 2024, arXiv:2409.14371. [6] D. Venturini et al., “I Depended on You and You Broke Me: An Empirical Study of Manifesting Breaking Changes in Client Packages,” ACM Trans. Softw. Eng. Methodol., 2023, arXiv:2301.04563. [7] C. E. Jimenez et al., “SWE-bench: Can Language Models Resolve RealWorld GitHub Issues?” in Proc. ICLR, 2024, arXiv:2310.06770. [8] T. Schick et al., “Toolformer: Language Models Can Teach Themselves to Use Tools,” in Proc. NeurIPS, 2023, arXiv:2302.04761. [9] M. Sharma et al., “Towards Understanding Sycophancy in Language Models,” in Proc. ICLR, 2024, arXiv:2310.13548. [10] S. Kadavath et al., “Language Models (Mostly) Know What They Know,” arXiv:2207.05221, 2022.
8