ConceptioArchivearXiv CS
arXiv CSopen access

Registry Descriptions Go Stale Unevenly: An 89-Day Measurement of Model Context Protocol Drift, and Why Drift-Ranked Re-Auditing Under-Covers It

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
cryptography, security, privacy, cybersecurity

Registry Descriptions Go Stale Unevenly: An 89-Day Measurement of Model Context Protocol Drift, and Why Drift-Ranked Re-Auditing Under-Covers It Gautam Bharti Independent Researcher ORCID 0009-0001-4448-1438 [email protected]

arXiv:2608.00997v1 [cs.SE] 2 Aug 2026

Preprint, 2026 Abstract Security studies of the Model Context Protocol (MCP) ecosystem have grown quickly, and they share a design: each audits a registry at a single point in time. None reports how long the registry descriptions those audits judged stay current - a necessary condition for any description-level finding to still apply, though not a sufficient one: we measure the shelf-life of the audited text, not the validity of a security finding itself, which turns on live tool behavior we do not observe (§7.1). We reconstruct 120 observations of the official MCP registry over 88.6 days, covering 19,099 distinct servers as it grew from 3,510 to 18,966. Our central result is a policy one: you cannot keep description-level findings current by re-auditing the servers that drift most. At a top-5% re-audit budget, ranking by prior drift catches only ~20% of the previously-seen servers whose description changes in a held-out window - versus ~27% for descriptor drift overall and only ~10% of all description changers. The limit is not that description drift is unpredictable - prior description change lifts next-period probability 4.8x (13.8% vs 2.8%) - but that the signal is exhausted almost immediately: only 5.0% of the population has any prior description-change history at all, against 16.4% for descriptors, so a top-5% budget consumes the entire ranking and caps near 20%, and every slot beyond it is filled by tie-break rather than by signal. Roughly half of all description changes then land on new arrivals a history ranking cannot reach by construction (the highest-drift group). The control that fits is content-binding - revalidate the moment a description's hash moves - plus a sized periodic full-catalog sweep for the new arrivals and the tail; a drift-history ranking is at best a partial, blind-to-new-arrivals control. This is scanner hygiene - keeping a description-level auditor's own findings current - not a runtime trust signal for an agent about to invoke a tool. The measurements behind this are the paper's second contribution. Description drift is concentrated and slow: of servers observed across at least ten intervals, three-quarters never change, the most active 5% generate 61% of all change events, and directly measured, only 11.9% of a cohort's descriptors change within 30 days (three non-overlapping cohorts span 11.4–12.3%; §3.2). Naively compounding the daily change rate predicts 35.8% at 30 days (73% at 89) - a ~3.0x overestimate at 30 days that widens to ~3.8x by 89 days, which we use only as a heavy-tail diagnostic, not a discovery; longer horizons rest on progressively fewer independent cohorts, so we report 19.1%-at-89-days as a trend, not a headline. We scope every claim precisely: this is a revalidation policy for the registry description text a description-level screen reads, and bounds neither tool-poisoning shelf-life nor the validity of any security finding - the live tool list and input schemas an attacker manipulates are a separate surface this data does not observe (§7.1). We release the panel, the figure generator, and the analysis code. (A deployment observation on our own scanner, §6, is secondary, single-deployment, and not among the paper's contributions.)

1

60 50 40 30

~27%

20 10 0

~20% blind to ~43-51% of changers (new arrivals, un-rankable)

5

Descriptor (any field) Description text

10 20 30 Re-audit budget (top X% by prior drift)

80 70 60 50 40 30 20 10 0

B Findings decay far slower than a naive rate implies Naive compounding (overestimate) Descriptor (measured) Description (measured)

100

73%

19%

30

60 Horizon (days)

89

Share of all change events (%)

A Drift-ranking under-covers description revalidation

Servers changed (%)

Rankable coverage of changers (%)

70

C Change is heavily concentrated 94%

80

79% 61%

60 40 27%

20 0

1

5 10 Most-active servers (top X%)

20

Figure 1. Three findings, one panel each. (A) Ranking servers by prior drift and re-auditing the top X% under-covers the description-revalidation surface: at a top-5% budget it catches ~27% of previously-seen descriptor changers but only ~20% of description changers, and is blind by construction to the ~43-51% of changers that are new arrivals it has never seen. (B) Directly measured drift decays far slower than naive compounding of the daily rate predicts (89-day descriptor survival 19% measured vs 73% compounded); description text, the surface a description-level screen must recheck, changes slower still. (C) Change is heavily concentrated: the most active 5% of servers produce 61% of all change events. Every value is read from the deposited figures.json and regenerated by make_paper_figure.py. CC-BY-4.0.

1. Introduction The Model Context Protocol gives agentic systems a uniform way to discover and call third-party tools, and its public registries have become, in under a year, a software supply chain with tens of thousands of entries. A security literature has formed around it at speed: large-scale registry audits [1], health-and-maintainability studies [2], attack analyses over harvested tool corpora [3], and client-side threat modeling of tool poisoning [4]. These works differ in scope and method but share one structural property: each observes the ecosystem once. The audit is a photograph. A photograph of a supply chain has a shelf life, and for MCP nobody has measured it. The question matters operationally - an integrator who consumed a description-level audit from [1] needs to know whether the text those findings were computed over is still the text being served in September (a necessary condition for the finding to still apply, though not a sufficient one - the live tool behavior can change under a stable description, and vice versa). To be unambiguous: this is a data-currency question for the auditor, not a runtime trust-TTL for an agent about to call a tool - nothing here licenses ”this server was audited N days ago, so it is safe to invoke now.” And it matters methodologically, because every claim of the form ”X% of descriptions exhibit property P” silently expires at a rate the literature has not characterized. This paper measures that rate, and finds that the interesting fact is not its magnitude but its shape. The two findings below are load-bearing and deposit-backed; a deployment observation on our own scanner (§6) is secondary and, being a single unreleased store, not independently replicable - we keep it out of the contribution claims. Drift is concentrated, and an audit ages slowly and measurably (§3–4). Three quarters of registry entries (of those observed across ≥10 intervals) never change over 88.6 days; the most active twentieth produces three fifths of all change events. Directly measured, only 11.9% of a cohort's descriptors change within 30 days - the three non-overlapping 30-day cohorts give 11.4 / 11.9 / 12.3%, which we take as the primary uncertainty; a moving-block bootstrap over overlapping cohorts gives a narrower 95% CI (11.6–12.2) that understates the between-cohort spread rather than corroborating it - so findings decay far slower than a naive rate implies. The naive rate is not merely imprecise: compounding the daily change rate predicts 35.8% at 30 days (and 73% at 89), because compounding an event rate into a population share assumes each event lands on a fresh server - the textbook heavy-tail failure, which we use as a diagnostic (§3.2) rather

2

than a discovery. We report the 30-day figure as the robust one and treat the 89-day survival (19.1%, one non-overlapping cohort) as a weak trend. Concentration is prospectively exploitable - but not on the surface a screen must revalidate (§5). At a top-5% re-audit budget, ranking by prior drift catches ~27% of the previously-seen servers that change (a 5–6x lift over random; ~13–17% of all changers, since half are new arrivals it cannot rank; 47% of change events); the same budget catches only ~20% of the previously-seen servers whose description changes, the surface a description-level screen actually has to recheck. Version churn is predictable; description rewrites are much less so. The operational consequence is specific and, we think, the paper's most useful result for operators of description-level registry screens: history-ranked re-auditing is the wrong control for description revalidation - the right one is content-binding (revalidate when the description hash moves) plus a periodic calendar sweep, neither of which a churn ranking provides. We scope this to description revalidation deliberately: the surface is the registry description text, not the live tool list or input schema an attacker ultimately manipulates (§7.1), which our data does not observe. Separately, §6 reports a deployment observation on our own scanner: applying the same lens to its verdict store surfaces a second staleness producer - the scanner's input pipeline lagging the registry, which mints verdicts wrong at birth rather than merely aged. We include it because instrument lag is absent from the threat models we know, but we do not count it among the paper's contributions: the store is a single unreleased deployment and §6's counts are not independently replicable (§7). We release the panel, the figure generator, and the analysis code; §6 also notes the three production changes the deployment observation forced.

2. Data and Method 2.1 The registry panel The official MCP registry is synced every four hours by our production pipeline; each successful sync commits the full snapshot to version control when it changed. The dated per-day snapshot files are ephemeral (they die with the CI runner), but the committed history is a faithful longitudinal record: 120 revisions from 2026-04-30 to 2026-07-28 (88.6 days, up to six observations per day), during which the corpus grew from 3,510 to 18,966 active servers; 19,099 distinct server names appear. We replay that history into a delta-encoded panel. For each observation and server we record f = SHA256 of the canonicalized server descriptor (any-field drift) and d = SHA-256 of the description alone (the description-revalidation surface), plus version and a parsed self-reported tool count where present. Because 75.2% of servers never change, delta encoding compresses 51 MB of per-revision digests to 1.3 MB. The panel is a deterministic function of the public snapshot history, pinned to snapshot commit e1d19d3 (see Data and Code Availability), so a replicator rebuilds it from version control and confirms an identical file. The released (pseudonymized) panel reproduces the named panel's full finding signature - every concentration tier, the 30/60/89-day survival curves, and an identical observation-timestamp hash - asserted by build_panel_deposit.py --verify. 2.2 The verdict panel The same repository's history yields 108 observations (2026-05-30 to 2026-07-28) of a production verdict store - a point-in-time LLM screen over registry descriptions, serving advisory verdicts on a public trust surface. Each verdict carries evaluated_at and, after 2026-07-19, content_hash = SHA-256 of the exact description judged. §6 joins this store to the registry panel. 2.3 Two traps for replicators publishedAt is per-version, not per-server. The registry rewrites publishedAt on every version publish verified directly: of 225 servers that changed version across one snapshot pair, 225 also changed publishedAt. Any survival analysis anchored on that field measures nothing. Our first attempt did exactly this and

3

produced a plausible, internally-consistent, wrong result; we caught it only because an independently derived quantity contradicted it. Hash derivations must mirror the writer byte-for-byte. The store's writer strips whitespace before hashing; an offline re-derivation that hashed raw registry text produced 34 false staleness accusations out of 781 (§6.1). The general rule: before comparing an independently derived hash to a stored one, reproduce the writer's exact normalization - or call the writer's own code. 2.4 Estimator cross-validation We compute rates more than one way and report where the estimates diverge as readily as where they agree - in this paper the divergence is the result. One agreement is load-bearing: (i) the per-day event rate - the direct interval-diff rate and the registry's publishedAt-implied version-publish intensity accumulate to the same ~36% of servers by 30 days (§3.2). A second, weaker consistency check (not load-bearing) is (ii) the post-judgment drift rate: verdict-dating gives ~26/day and the panel's independent description-change rate ~20/day (§6.2) - the same order of magnitude, which is all we claim of it; we do not treat 26 vs 20 as a precise match. The load-bearing disagreement is between those event-intensity estimators and direct cohort survival: compounding the event rate - and, separately, the version-recency population share - both put ~70% of servers changed by ~90 days, while direct survival of the cohort present at start measures ~19% (§3.2). That two independent routes to event intensity overshoot direct survival by the same ~3.8x is what isolates the error to the population-share conversion, not the rate and not the survival measurement - the latter's own reliability resting on its moving-block-bootstrap CI (§3.2), not a second estimator.

3. How Fast Does the Registry Change? 3.1 The rate Across 119 inter-observation intervals, the share of common servers whose descriptor changed, normalized per day: median 1.47%/day (mean 1.73; p10 1.01; p90 2.37). The mean is inflated by short-gap intervals a small absolute change divided by a fraction of a day - so we use the median throughout. Description-only change is several times rarer (§3.3). 3.2 The compounding error It is tempting to convert that rate into a shelf life: 1 − (1 − 0.01465)^N (the unrounded median of §3.1) gives 35.8% of servers changed by day 30, 58.7% by day 60, 73.1% by day 89 - an audit ”half-life” of about 47 days. Direct measurement refutes every one of those numbers: Horizon Compounded prediction Measured (cohort mean) 30 days 60 days 89 days

35.8% 58.7% 73.1%

Cohorts

11.87% (range 10.37–13.71) 60 16.30% 36 19.11% 9

Estimand. Under commit-on-change sampling (an observation exists only when the snapshot changed), we count a server as ”changed within N days” if its descriptor hash differs from its cohort-start value at any observation within the window - the ever-flipped estimand, which upper-bounds an endpoint-only definition (a change that later reverts still counts). A cohort start is each observation; a server is in the cohort if present at the start and observed with ≥0.9N days of follow-up. The 30-day figure is the most defensible single number in this paper, but its uncertainty must be stated honestly: the 60 cohort starts overlap - each shares its 30-day future window with its neighbours - so their range (10.4–13.7%) is pseudo-replication, not independent scatter. We therefore lead with the three genuinely non-overlapping 30-day cohorts (starts ≥30 days apart): 11.4 / 11.9 / 12.3%, and treat that scatter as the honest uncertainty. A moving-block bootstrap gives a 95% CI of [11.6, 12.2]% (full recipe in figures.json: B = 2000 resamples, block length = 40 cohorts ≈ the 30-day overlap horizon, seed 42, percentile method), 4

but that interval is narrower than the three-cohort scatter and excludes the 11.4% cohort - because the bootstrap resamples overlapping cohorts that share observations, it understates between-cohort variance rather than confirming the low end, so we do not read it as corroboration. The point estimate is tight and trend-free across a 3.5k→14k corpus. The 89-day figure, by contrast, admits only one non-overlapping cohort and rests on nine overlapping early ones; we treat it as a weak trend, not a measurement. The error is structural, not sampling noise. The daily rate counts events; compounding it assumes each event strikes a fresh server. On this population the assumption fails by design (§4), and the failure generalizes: registries, package indexes, and user-activity streams are typically heavy-tailed, so the compounding shortcut is biased upward almost everywhere it is used. The cheap detector: compare distinct entities that ever fired an event against total events. Here, 18,748 servers observed across ≥10 intervals produced 15,805 change events from only 4,652 distinct servers. A second, independent estimator confirms the event rate - not the survival. The share of servers whose current version was published within the last N days (from the registry's public publishedAt, no diffing) gives 36.7% / 56.5% / 70.4% at 30/60/90 days. These track the compounded prediction (35.8 / 58.7 / 73.1), not the direct survival (11.9 / 16.3 / 19.1) - as they must: version-recency counts new arrivals and per-version publishedAt rewrites (§2.3), so like compounding it reflects event intensity, not the survival of a fixed cohort. That an empirical intensity estimator and an arithmetic one overshoot direct survival together - by a factor that grows from ~3.0x at 30 days to ~3.8x at 89 days - is exactly what localizes the error to the population-share conversion rather than to the event rate or the survival measurement. (This estimator reads the registry snapshot, not the released panel, which carries no publishedAt.) 3.3 Description drift, the revalidation surface Measured directly, description-only change: 3.32% of servers by 30 days (range 2.58–4.17 across 60 cohorts), 5.52% by 60, 6.85% by 89. On an 18.5k corpus that is roughly 20 description rewrites per day - each one a mutation of exactly the text a description-level screen judged, and therefore must revalidate. We call this the description-revalidation surface deliberately, not ”the tool-poisoning surface”: the registry description is the input a description-level screen sees, but the behavior an attacker manipulates lives in the live tool list and input schemas, which the registry does not carry and this dataset does not observe (§7.1). The two surfaces can move independently; our claims are about the one we measure.

4. Who Changes? 4.1 A stable majority, a busy minority Of 18,748 servers observed across at least 10 intervals, 14,096 (75.2%) recorded zero descriptor changes in the whole window. The change-count histogram: 0 → 14,096 · 1 → 2,183 · 2–4 → 1,594 · 5–9 → 525 · 10+ → 350. ”Never changes” here means the registry descriptor is stable; a server whose live tool behavior mutates under a fixed description sits in this 75% and is precisely the blind spot §7.1 names. Descriptor stability is not tool-behavior stability, and this stable majority must not be read as a majority that is safe to leave un-probed. Concentration of the 15,805 change events: top 1% of servers = 26.7%, top 5% = 61.2%, top 10% = 78.7%, top 20% = 94.3%. Publisher-level concentration is milder but real: the top ten publishers account for 17.7% of events (the single busiest, 5.5%); 3,054 of 11,900 publishers ever produced a change event. (Publisher-level figures need the plaintext name and so are computed from the internal named panel, not the pseudonymized release - see §7.) The ≥10-interval eligibility filter - which by construction excludes late arrivals, the high-hazard newborns of §4.2 - does not carry these numbers. Varying the threshold from no filter (≥1) through ≥20 moves top-5% event share only within 61.2–61.5% and never-changed within 75.2–75.4% (never-changed: 75.4 / 75.3 / 75.2 / 75.2 at ≥1/5/10/20; top-5%: 61.5 / 61.3 / 61.2 / 61.4). The filter removes ~300 short-window servers without shifting either headline, so the concentration finding is not an artifact of censoring the newborns.

5

4.2 Drift is front-loaded - and age does not explain concentration Restricting to the servers first observed inside the window (so age is observed, not left-censored), the anyfield hazard falls from 38.0 changes per 1,000 server-days in a server's first week to 6.9 by week ten (5.5x); the description-only hazard falls 7.0 → 0.7 (10x). New servers are where drift lives. But age is not a substitute for drift history: the top 5% by change count is 80.8% newborn against an 81.3% newborn baseline - statistically indistinguishable. Age and concentration are close to independent axes, and §5 shows the history axis is the one that predicts. 4.3 Attack surface only grows 861 servers self-report a tool count in their description. Across the window: 248 increase events against 17 decreases (93.6% increases), net +2,435 tools. Single-day jumps include 8 → 27 tools. Self-reported counts are prose, not a probe of the real tool list (Limitation §7.2), but the direction is unambiguous: an audit's coverage claim decays faster than its findings do. 4.4 Delisting is negligible, but reappearance is not 885 servers vanished from at least one observation; only 133 (0.7% of the final corpus) were absent at the end. Removal is not a meaningful decay channel; mutation is. The complement is more interesting than the headline. Because almost nothing leaves permanently, almost everything that vanishes returns: we record 778 reappearance events across 764 distinct names, 14 of which vanished and returned more than once. A registry name is the referent an allowlist, a catalog entry, or a procurement approval keys on, so a name that leaves and comes back is the ecosystem’s analogue of the package-reclamation pattern familiar from npm and PyPI - and unlike ordinary mutation, it is invisible to a consumer who only checks whether the name is still listed. We can bound how often the returned artifact differs from the one that left. Of the 778 reappearances, 12 (1.5%) returned with a different description hash and 43 (5.5%) with a different descriptor hash. The overwhelming majority return byte-identical, which is consistent with transient sync or registryside availability blips rather than re-registration. So the reappearance channel is real, enumerable, and small; we report it because a null of this shape is what a threat model needs in order to stop worrying about a plausible attack path, and because no prior MCP registry study reports it at all. We stress the limit: the panel carries content hashes, not ownership or publisher identity, so these 12 cases establish that content changed across a disappearance, not that the name changed hands.

5. Can Staleness Be Predicted? 5.1 Retrospective lift Splitting the window at its midpoint (2026-06-29; 3,492 servers present in both halves): P(drift in H2 | drifted in H1) = 31.2%, against 4.4% for servers stable in H1 - a 7.1x lift. On description text alone the same split gives 13.8% against 2.8%, a 4.8x lift. Description rewrites are therefore substantially predictable from prior description rewrites; §5.3 shows why that predictability nonetheless buys little coverage. Two properties of this estimate constrain how far it generalises. First, the base is servers present at both the first observation and the midpoint, which is 99.5% of the panel-start corpus and only 18.4% of the corpus at the end. The conditional therefore describes the oldest tier and excludes by construction the new arrivals that carry roughly ten times the description hazard (§4.2) and about half of all description changes (§5.2). Second, it is one split of one window; we report it as a retrospective description of this corpus, not as a forecast. 5.2 Prospective targeting Retrospective concentration is not a policy, so we test the policy: rank servers by drift observed in a training window only, re-audit the top X%, and measure the share of held-out change events captured. We use two 6

training cut-dates (44-day and 60-day windows from panel start); with only two cut-points we report the spread across them as a sensitivity range, not a confidence interval: We report entity coverage, not event share, because §§3–4 warn that events ≠ entities under a heavy tail: a budget chosen by prior drift preferentially catches high-frequency servers, so the event share (~47% at top-5%) overstates how many changed servers are actually re-checked. But entity coverage itself has a denominator choice that must be stated. A history ranking can only rank servers that existed at ranking time, so rankable-population coverage - the share of the servers present at ranking time that change and are caught - is the measure of the policy's performance at its job. The whole-population view is lower, because roughly half of all distinct servers that change in a test window are servers that arrived after the ranking cut and are outside any history ranking's reach by construction. We give both. Budget

44d rankable cov. 44d all-changer cov. 44d event share 60d rankable

60d all-changer 60d event

top 5% top 10% top 20% top 30%

26.0% (5.2x) 39.1% 53.6% 61.3%

16.7% 24.4% 34.4% 36.6%

12.8% 19.3% 26.4% 30.2%

47.5% 60.5% 71.2% 74.6%

29.4% (5.9x) 42.9% 60.5% 64.4%

47.1% 59.3% 73.4% 75.4%

Read honestly: over the population it can rank, a top-5% budget by prior drift catches ~26–29% of the servers that will change across the two cut-dates - a 5–6x lift over a same-size random sample. Over all changers it catches only ~13–17%, because ~43–51% of test-window changers are new arrivals it cannot rank, and §4.2 shows new servers are exactly where drift concentrates. That blind spot is not a flaw in the measurement but a structural limit of the policy, and it is a second reason (alongside §5.3) that a history ranking is a partial control: it can only cover servers it has already seen. The event share is not a coverage number and we do not use it as one. 5.3 The description surface resists targeting Repeating the identical protocol on description-only changes, the top-5% budget catches only 19.3% / 21.5% of the previously-seen servers whose description will change (rankable coverage, ~4x lift; ~9–12% of all description changers), against 26–29% for descriptor drift (5–6x); the event-share figures are 25.3% / 27.5% versus 47%. The gap is real but smaller than the event share alone implied - the point is not that description drift is unpredictable, but that a drift-history ranking, tuned to the predictable version churn, buys measurably less on the surface a description-level screen must recheck. Version churn is habitual and publisher-driven; description rewrites are sparser, though not markedly less predictable per server. Why the ranking caps, and why the top-5% budget is the right place to read it. Measuring the conditional directly: prior description change lifts next-period probability 4.8x (13.8% vs 2.8%, 𝑛 = 3,492), against 7.1x for descriptors. So the description ranking is informative. What differs is how much of the population it can rank at all. In the 44-day training window only 613 of 12,193 servers (5.0%) carry any prior description change, against 2,000 (16.4%) for descriptors. A top-5% budget is 609 slots and consumes essentially the entire signal pool; taking the whole pool covers only 19.5% of the rankable changers. Every slot beyond 5% is therefore filled by tie-break, not by signal - 49.7% of a 10% budget, 74.9% of 20%, 83.2% of 30%. This is what produces the otherwise anomalous increasing marginal returns of the description curve (+0.72, +1.01, +1.78 coverage points per budget point across 5→10→20→30): the apparent improvement is not the ranking working better, it is uniform sampling replacing an exhausted ranking. The descriptor curve, whose signal pool is not exhausted until a 16% budget, shows the normal decreasing shape (+2.62, +1.45, +0.77). The 60-day window replicates both (description pool 829 / 5.9%, ceiling 23.4%). We therefore quote the top-5% figure not as a convenient budget but as the ranking’s ceiling: it is the point at which the method has spent everything it knows. Ranking is deterministic - descending training-window change count, ties broken by ascending server key - so the tie-break slots are arbitrary with respect to future change, not adversarially chosen. We consider this tension - not the headline concentration - the paper's most useful result for operators of description-level registry screens (not for tool-poisoning defense generally; §7.1), and it has a direct 7

policy implication: do not rank by descriptor churn to catch description change. The control that fits the data is content-binding - revalidate a verdict the moment its bound description hash moves - backed by a periodic full-catalog sweep for the long tail and the new arrivals a ranking cannot reach. Sizing the sweep concretely: content-binding re-screens only the ~20 descriptions that change per day (§3.3), while a sweep at cadence C days bounds worst-case description staleness to C days at roughly N/C re-screens per day - for the ~19k-server corpus, ~2.7k/day for a weekly sweep or ~630/day for a monthly one - a fixed, plannable budget that a drift-history ranking's variable, blind-to-new-arrivals coverage does not give. (This is a policy for description revalidation only; the live-behavior surface - the tool list and input schemas an attacker manipulates - is not observed here, §7.1, and none of this bounds tool-poisoning shelf-life.) 5.4 A null result Servers that received a public screening flag showed post-flag content churn of 0.93x baseline (2,504 flagged server-days of exposure against 230,610) - no detectable behavioral response to being flagged. Exposure is small and flags may simply not be publisher-visible; we report it to prevent its silent omission.

6. Consequences for a Deployed Scanner The scanner under study judges each server's description once at ingest and binds every verdict to content_hash, the SHA-256 of the exact text judged; until this work nothing re-compared that binding to the registry. These counts come from the production verdict store at the 2026-07-28 discovery snapshot; that store is not released (§7), so they are the one family of numbers a reader cannot recompute from the deposited panel. We report this as a deployment observation, not a contribution. 6.1 How stale were live verdicts? Joining all 18,543 live verdicts to the registry (99.6% resolve): 747 (4.04%) were bound to a description no longer published (an initial 781 included 34 whitespace-trap false positives, §2.3; the panel corroborates 747 of 747 datable cases). Only 13 were also past their clock-based expiry, so content-binding and wall-clock expiry are nearly disjoint signals. A minority of the 747 carried a failing dimension; we report that in aggregate, name no server, and publish no per-record re-screen tally - a description hash cannot distinguish remediation from rebranding from evasion, so the only defensible claim is categorical: a point-in-time store can serve a failing verdict bound to text no longer published, and here it did. 6.2 Two producers of staleness Dating the 747 against the panel splits them by producer: 379 (51%) were born stale - evaluated_at postdates the last day the judged text was published, the signature of the instrument reading a lagged input - and 368 drifted after judgment, judged against then-current text a publisher later changed (379 + 368 = 747). Post-remediation, drift runs at ~26/day, the same order as the panel's independent ~20/day description-change rate (§2.4) - a consistency check, not a precise agreement. Instrument lag - verdicts wrong at birth rather than aging - is invisible to any model that assumes the scanner sees the present, is the larger share here, and is absent from the MCP threat models we survey (§8). 6.3 Mitigations The findings forced three production changes (live 2026-07-29) that instantiate the §5.3 policy: read-time content-binding that renders a mismatched verdict STALE without ever softening a failing verdict; a failclosed input-freshness gate that refuses to judge a snapshot older than 24 hours, so an input outage ages the store honestly instead of minting born-stale verdicts; and a re-screen lane on the existing catalog-sweep timer - the periodic sweep §5.3 sizes. That sweep also backfilled server_id, a subject-binding guard present in the writer but unpopulated because no lane had ever re-run over judged records - a reminder that a point-in-time pipeline can leave a schema guard latent in the data until something forces a pass over the corpus.

8

7. Limitations §7.1 Registry metadata, not probed behavior. Descriptor drift is not tool-behavior drift; a server can change its real tool list or input schema - the surface an attacker ultimately manipulates - without touching its registry entry, and our data never observes that surface. This is the largest external-validity gap and the reason every defender claim in this paper is scoped to description revalidation; closing it (conformance probing of live tool lists) is ongoing work. §7.2 Self-reported tool counts are regex-parsed prose from 861 of ~19k servers. §7.3 Irregular cadence. Observations land only when the snapshot changed; intervals range from ~2.5 hours to over a day, and per-day normalization inherits that irregularity. §7.4 Single registry. The largest prior audit spans six registries [1]; our panel covers the official one. Generalization is unproven. §7.5 Censoring. ”75.2% never change” is right-censored at 88.6 days; the 89-day survival figure rests on nine cohorts drawn from the small early corpus. §7.6 The scanner is our own. §6 gains access no external auditor would have, at the cost of studying one deployment; the two-producer decomposition should be tested against other scanners, and its counts are not independently replicable.

8. Related Work [1] audits 67,057 servers across six registries (collected late June–early July 2025) and is, to our knowledge, the largest MCP security measurement; its limitations discuss tool-extraction and static-analysis precision but not temporal validity, though the authors observed change directly (”After repeating the data collection… we identified only one expired GitHub token”) without pursuing what repetition implies. [2] studies 1,899 servers (March 2025) and states the snapshot limitation explicitly - ”Our study is limited to a snapshot of MCP servers as of March 2025” - without quantifying it; this paper is, in effect, the measurement that caveat calls for. [3] analyzes 1,360 servers / 12,230 tools with no stated collection date and no temporal discussion. [4] threat-models seven MCP clients with tool poisoning as the central vector; its threats-to-validity section covers scoring subjectivity and client generalization, and its recommendations table calls for ”periodic rescanning” to catch tools that ”become malicious over time” - but it fixes no interval and measures no rate. We are careful about the bridge here: [4]'s target is the live tool, which our data does not observe (§7.1), so we do not supply its re-scan interval. What we supply is a decay/revalidation rate for the registry description text a description-level screen reads; [4]'s live-tool re-scan cadence remains unmeasured, by us and by the prior work. Across all four, none reports a revalidation interval or a decay rate for what it studies. Longitudinal measurement of software-supply-chain ecosystems is well established outside MCP: for the npm package graph and its security threats [5], for the corpus of open-source supply-chain attacks [6], and for the web-PKI certificate ecosystem [7]. Our contribution is bringing that lens to MCP: the first longitudinal panel of an MCP registry, and the finding that drift-history re-auditing under-covers the description-revalidation surface (§5.3), with the content-binding-plus-sweep policy that follows. The §6 instrument-lag observation is a deployment note, not a contribution claim - it comes from a single unreleased store and is not independently replicable - though it does illustrate a staleness component invisible without paired scanner-side data.

Data and Code Availability The pseudonymized MCP Registry Drift Panel v1 and the full analysis code are released under CC-BY4.0. Every panel-derived number in this paper regenerates from the deposit: paper_figures.py produces figures.json, paper_figures_addenda.py produces figures_addenda.json (§4.4, §5.1, §5.3), and make_paper_figure.py renders Figure 1. All three run against the deposited panel with no arguments and no dependencies beyond the Python standard library. The panel is deposited at Zenodo (concept DOI 10.5281/zenodo.21709945, which resolves to the current version) and pinned to the official MCP registry snapshot captured at commit e1d19d3 (data cutoff 2026-07-28; 120 observations); the exact bytes are fixed independently of the DOI by figures.json:_pin.panel_sha256. Two earlier datasets from this project 9

are also cited: MCP Drift v1 (DOI 10.5281/zenodo.21449150) and Source Liveness v1 (DOI 10.5281/zenodo.21501868). The production verdict store underlying §6 is a single unreleased deployment (§7).

References [1] Xiaofan Li and Xing Gao. ”A First Look at the Security Issues in the Model Context Protocol Ecosystem.” arXiv:2510.16558, 2025. [2] Mohammed Mehedi Hasan, Hao Li, Emad Fallahzadeh, Gopi Krishnan Rajbahadur, Bram Adams, and Ahmed E. Hassan. ”Model Context Protocol (MCP) at First Glance: Studying the Security and Maintainability of MCP Servers.” arXiv:2506.13538, 2025. [3] Shuli Zhao, Qinsheng Hou, Zihan Zhan, Yanhao Wang, Yuchong Xie, Yu Guo, Libo Chen, Shenghong Li, and Zhi Xue. ”Parasites in the Toolchain: A Large-Scale Analysis of Attacks on the MCP Ecosystem.” arXiv:2509.06572, 2025. [4] Charoes Huang, Xin Huang, Ngoc Phu Tran, and Amin Milani Fard. ”Model Context Protocol Threat Modeling and Analysis of Vulnerabilities to Prompt Injection with Tool Poisoning.” Journal of Cybersecurity and Privacy 6(3):84, 2026. DOI: 10.3390/jcp6030084. (Preprint: arXiv:2603.22489.) [5] Markus Zimmermann, Cristian-Alexandru Staicu, Cam Tenny, and Michael Pradel. ”Small World with High Risks: A Study of Security Threats in the npm Ecosystem.” USENIX Security Symposium, 2019. [6] Marc Ohm, Henrik Plate, Arnold Sykosch, and Michael Meier. ”Backstabber's Knife Collection: A Review of Open Source Software Supply Chain Attacks.” DIMVA, 2020. [7] Zakir Durumeric, James Kasten, Michael Bailey, and J. Alex Halderman. ”Analysis of the HTTPS Certificate Ecosystem.” Internet Measurement Conference (IMC), 2013.

10

Record · ID 423868 · SHA-256 0ea2354ee17c4a86
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.