Characterizing Network Centralization and Observability in the Remote MCP Ecosystem Muhammad Abdullah Sohail
arXiv:2609.19100v1 [cs.CR] 16 Sep 2026
University of Calgary Calgary, AB, Canada [email protected]
Abstract—The Model Context Protocol (MCP) has emerged as the dominant interface for connecting autonomous agents to external data sources and execution environments. The ecosystem’s transition from local process execution to remote Streamable HTTP deployments introduces unmeasured architectural and security constraints at scale. This paper presents a threetier observability framework comprising catalog metadata (O0 ), passive compliance signals (O1 ), and live vulnerability analysis (O2 ), applied to empirically characterize the public MCP server ecosystem. Evaluation of a stratified sample of 179 remote endpoints across two primary public registries reveals significant infrastructural consolidation. The Herfindahl-Hirschman Index (HHI) computed over the Autonomous System Number (ASN) distribution yields a value of 0.736, well above the 0.25 threshold for a highly concentrated market. Analysis further indicates that server authentication is strongly correlated with hosting platform choice rather than individual operator configuration, with 95% of commercial PaaS-hosted servers enforcing gateway-level OAuth 2.1 with PKCE. The empirical results identify a SecurityObservability Tradeoff observed in the current ecosystem: the platform-level authentication mechanisms that secure the majority of servers simultaneously limit automated vulnerability scanning capabilities, constraining the ability of AI gateway operators to assess tool-poisoning vectors without prior credential provisioning. Index Terms—Model Context Protocol, AI Agents, Tool Poisoning, Network Measurement, Security, Centralization, HHI
I. I NTRODUCTION The integration of Large Language Models (LLMs) with external tools has shifted the bottleneck of AI agent architectures from reasoning capabilities toward context orchestration and tool execution safety. The Model Context Protocol (MCP) [1] defines a standardized control plane in which client agents query standalone servers for context, tools, and executable prompts. Initially designed for local inter-process communication via standard I/O streams, the MCP ecosystem is transitioning toward remote deployments over Streamable HTTP to support distributed, multi-tenant agent fabrics. This transition to externalized infrastructure exposes the MCP ecosystem to the attack surfaces of traditional web services, particularly indirect prompt injection and tool poisoning. The risk profile is compounded by the non-deterministic, Accepted at the 1st IEEE ICNP Workshop on Network Infrastructure and Protocols for AI Agents (NIPA 2026), co-located with the 34th IEEE International Conference on Network Protocols (IEEE ICNP 2026), Tempe, AZ, USA, October 2026.
autonomous nature of the client agents. A malicious or compromised remote MCP server can introduce adversarial content into an LLM’s context window through manipulated schema payloads, as demonstrated by Greshake et al. [2]. Gateway operators must therefore be able to assess the security posture of a remote server before provisioning it within an active agent pipeline. Despite this growth, no prior work has conducted a structured, server-side empirical measurement of the remote MCP deployment landscape. Prior analyses have focused primarily on local vulnerability discovery or client-side registry aggregation [3]. This paper addresses that gap through a systematic network measurement study of the remote MCP server ecosystem. Three research questions guide the study, connected by a shared concern: whether gateway operators have sufficient visibility to make informed trust decisions about remote MCP servers. 1) Hosting consolidation determines the blast radius of infrastructure failures. If most remote MCP servers share a single network provider, a routing disruption or platform outage could disable agent tooling across the ecosystem simultaneously. 2) Authentication posture determines whether gateway operators can pre-screen a server’s tool schemas for poisoning vectors before provisioning it. If authentication is enforced at the platform level rather than configured by individual operators, the observability available to external scanners depends on architectural choices outside any single developer’s control. 3) The Security-Observability Tradeoff ties the first two questions together: the platforms that centralize hosting also enforce authentication, which blocks the automated scanning needed to verify tool safety. Understanding this relationship is necessary for designing gateway integration workflows that remain secure under the ecosystem’s current architectural constraints. The contributions of this work are as follows. A threetier observability framework (O0 , O1 , O2 ) is introduced to partition the signal space available to gateway operators across distinct phases of the server integration lifecycle. Quantitative evidence of high hosting consolidation is presented, measured by an HHI of 0.736 computed over the ASN distribution of reachable servers. The study presents empirical evidence
that server authentication posture is strongly correlated with hosting platform architecture rather than individual developer configuration. The Security-Observability Tradeoff is characterized as an empirical property of the current ecosystem: the mechanism that produces the ecosystem’s security baseline simultaneously reduces the observability required to verify it through automated scanning.
Tier 0 (O0 ) Catalog Metadata
Tier 1 (O1 ) Passive Probing
II. BACKGROUND AND R ELATED W ORK A. The Model Context Protocol Developed by Anthropic and contributed to the Linux Foundation AI and Data Foundation (AAIF), MCP defines three interaction primitives between clients and servers: Tools (executable server-side functions), Resources (contextual data retrieval endpoints), and Prompts (reusable interaction templates). The June 2025 protocol specification [1] formalized Streamable HTTP as the canonical remote transport, deprecating the prior HTTP and SSE two-channel design that suffered from sticky-session constraints incompatible with stateless load balancers. The revised specification also mandated OAuth 2.1 with PKCE S256 for all hosted, multi-tenant server deployments. B. Tool Poisoning and Indirect Prompt Injection The threat motivating this study originates in the broader field of adversarial attacks on LLM-integrated systems. Greshake et al. [2] introduced the indirect prompt injection threat model, demonstrating that an adversary controlling content retrieved by an LLM-integrated application can subvert model behavior without direct access to the query interface. In the MCP context, this threat translates to tool poisoning: a malicious server operator embeds adversarial instructions within tool name strings, description fields, or schema annotations. When an LLM-based client enumerates available tools, these instructions are processed as legitimate context, potentially causing the agent to exfiltrate data or execute unauthorized tool calls. Empirical red-teaming of MCP infrastructure has confirmed that this vector is exploitable in production settings [4]. The Internet of Agents security survey [5] situates these vulnerabilities within a broader taxonomy of trust boundary violations across multi-agent systems, and analogous injection risks have been identified in the A2A protocol ecosystem [6]. C. Internet Infrastructure Centralization Centralization of internet infrastructure is a welldocumented concern in empirical network measurement research. Moura et al. [7] quantified the concentration of global DNS query traffic across major recursive resolvers, finding that a small number of cloud providers account for a disproportionate share of resolution traffic. Complementary work on the web hosting industry [8] documents how market dynamics drive independent operators toward shared infrastructure providers, producing emergent oligopolies at the IP layer. This body of work establishes both the methodological precedent and the analytical vocabulary for quantifying centralization in novel protocol ecosystems. The
AUTH GATED
Gateway Auth Enforcement
Blocks O2 OPEN
Tier 2 (O2 ) Live Scanning Fig. 1. The Three-Tier Observability Framework. AUTH GATED servers prevent Tier 2 analysis, producing the Security-Observability Tradeoff.
present study applies the Herfindahl-Hirschman Index to the MCP deployment layer, extending this line of measurement inquiry to AI agent tooling infrastructure. D. Ecosystem Measurement for AI Protocols Prior measurement studies of the AI agent ecosystem have focused on interoperability standards and transport protocol characterization [9] or congestion control mechanisms for agent traffic [10]. The closest directly comparable prior work, MCPCrawler [3], characterized client-side interaction patterns and aggregated marketplace metadata across several registries, finding SSE to be the dominant transport in the 2025 prerevision dataset. The present study differs in scope: it focuses on the server-side remote deployment layer rather than clients, applies a three-tier observability framework that separates catalog-level, post-connection, and live security signals, and surfaces the Security-Observability Tradeoff, a property that client-centric analyses do not address. III. M EASUREMENT M ETHODOLOGY A. Overview and Formal Framework The measurement pipeline is structured around three incrementally deeper layers of observability, each corresponding to a distinct phase of what a gateway operator would encounter when integrating an external MCP server (Fig. 1). Tier 0 (O0 ) captures catalog-level metadata accessible without any network connection to the server. Tier 1 (O1 ) captures passive signals available immediately after establishing a protocol connection. Tier 2 (O2 ) captures active security analysis signals, available only when the server permits unauthenticated tool enumeration. B. Data Collection and Stratified Sampling Server endpoints were sourced from two public directories: the Anthropic-curated Official MCP Registry (registry.modelcontextprotocol.io/v0/servers) and Smithery, the largest commercial PaaS registry for remote MCP deployments. These two registries were selected because they are, at the time of measurement, the only
TABLE I DATASET S TATISTICS : C ROSS -R EGISTRY R EACHABILITY AND S ECURITY P OSTURE Metric Attempted Reachable Reachability Rate
Official Reg.
Smithery
Total
29 20 69.0%
150 56 37.3%
179 76 42.5%
D. Tier 1: Post-Connection Passive Probing For the 76 reachable servers, the measurement probe initiated protocol-compliant Streamable HTTP handshakes, issuing the initialize method request followed by the tools/list request. The O1 feature set is defined as: O1 = {k, ¯l, q̄, tls, pv}
Posture of Reachable Servers AUTH GATED OPEN (SAFE) FAILED CONN
10 7 3
53 0 3
63 7 6
publicly accessible and API-queryable directories of remote MCP server endpoints. Private, corporate, and invite-only registries are excluded from the study because they are not discoverable through public interfaces; this exclusion is itself a characteristic of the ecosystem’s discoverability surface rather than a methodological limitation that can be overcome without privileged access. To mitigate heavy-tail noise from the long-tail distribution of inactive registrations, a stratified sampling approach was applied to Smithery. Using the registry pagination API, 150 servers were collected in equal strata of 50 across three adoption tiers: the top-50 servers by useCount (high-adoption), servers ranked 201 to 250 (mid-tier), and servers ranked 1001 to 1050 (long-tail). This stratification ensures representation across the full adoption spectrum rather than biasing toward popular or recently-listed servers. The complete population of 29 remote-capable endpoints in the Official MCP Registry was captured as a census without stratification. Of 179 endpoints attempted across both registries, 76 (42.5%) accepted TCP connections within a 5000 ms timeout. The remaining endpoints returned connection timeouts or terminal HTTP 4xx/5xx errors, consistent with servers that are registered in catalog metadata but not actively maintained in deployment. Table I presents the full cross-registry breakdown.
(2)
where k is the tool count returned by tools/list, ¯l is the mean description length across exposed tools, q̄ is the mean input schema quality score (fraction of parameters carrying both a type declaration and a description string), tls is a binary indicator of HTTPS enforcement, and pv is the negotiated protocol version. A composite compliance score σ was computed per server to summarize adherence to specification recommendations in a single scalar, enabling comparison across servers: σ=
1 tls + ⊮[pv ≥ 2024-11-05] + q̄ 3
(3)
O1 features are unavailable for AUTH GATED servers, as the authentication challenge is issued prior to MCP handshake completion. E. Tier 2: Live Security Analysis Ground-truth security labels were established using the public agentseal scan-mcp --url CLI [11]. This pipeline evaluates tool schemas against a curated repository of promptinjection patterns and semantic poisoning heuristics, applying embedding cosine-distance calculations to detect semantically obfuscated payloads. Each server receives a trust score τ ∈ [0, 100], binarized as SAFE when τ ≥ 70 and NEEDS_REVIEW otherwise. The parser was extended to classify servers returning HTTP 401 or 403 responses as AUTH_GATED, preventing conflation of authentication enforcement with tool payload safety.
C. Tier 0: Catalog Metadata Tier 0 features were extracted from registry APIs prior to any network connection to the target servers. These features capture what a gateway operator would know about a server before establishing any network contact, representing the minimum information available for trust decisions. The O0 feature set is defined as: O0 = {d, c, T, vs , src}
(1)
where d is the natural language description string, c is the set of categorical tags, T is the declared transport type, vs is the declared protocol version, and src denotes the source registry. Popularity proxies such as useCount and the binary verified badge were excluded to ensure the framework measures intrinsic server characteristics rather than platformassigned endorsement signals.
IV. H OSTING C ONSOLIDATION A. Cross-Registry Infrastructure Analysis The first research question examined the degree of infrastructural concentration across the remote MCP ecosystem. Of the 76 reachable endpoints, 56 (73.7%) are hosted on Smithery’s *.run.tools domain. The aggregate figure, however, is partially an artifact of the 150:29 registry sampling ratio. A cross-registry comparison removes this confound: among Official MCP Registry servers, sampled independently of Smithery, 0 of 20 resolve to *.run.tools. Among Smithery-registry servers, 56 of 56 reachable endpoints resolve exclusively to *.run.tools. There is complete absence of cross-registry hosting. No Official Registry server was deployed on Smithery PaaS infrastructure, and no Smitherylisted server was hosted on a third-party provider (Fig. 2).
Fig. 3. ASN distribution of 76 reachable endpoints. AS13335 (Cloudflare) accounts for 85.5% of servers. Fig. 2. Hosting distribution by source registry. Zero-count annotations highlight complete absence of cross-pollination.
B. ASN-Level Concentration To evaluate centralization at the network layer, ASN records were resolved for each reachable endpoint via the ip-api.com enrichment API. Fig. 3 presents the resulting distribution. The dominant ASN is AS13335 (Cloudflare, Inc.), which accounts for 65 of 76 servers (85.5%). The remaining servers are distributed across AS400940 (3 servers), AS15169 (Google Cloud, 3 servers), AS16509 (Amazon Web Services, 3 servers), AS397273 (1 server), and AS8075 (Microsoft Azure, 1 server). To formally quantify this concentration, the HerfindahlHirschman Index (HHI) was computed. HHI is the standard metric used in antitrust economics and in prior internet centralization studies [7] for quantifying market P 2concentration from share distributions. It is defined as i si where si is the fractional share of entity i. Values below 0.15 indicate an unconcentrated market, values between 0.15 and 0.25 indicate moderate concentration, and values above 0.25 indicate high concentration under DOJ antitrust guidelines. For the observed ASN distribution: HHI = 0.8552 + 0.0392 + 0.0392 + 0.0392 + 0.0132 + 0.0132 = 0.736
(4)
The observed HHI of 0.736 is 2.94 times the highconcentration classification threshold. By this metric, the remote MCP ecosystem is more concentrated at the network layer than most markets that have attracted regulatory antitrust scrutiny. This finding parallels Moura et al.’s documentation of DNS traffic centralization [7] and extends the same dynamic to AI agent tooling infrastructure. V. AUTHENTICATION AS A P LATFORM P ROPERTY A. Posture Distribution by Hosting Type The second research question examined whether authentication enforcement correlates with hosting platform choice or varies independently across individual operators. Tier 2 (O2 ) data across the 76 reachable servers reveals a clear
pattern (Fig. 4). Of the 56 Smithery-hosted servers, 53 (94.6%) returned AUTH_GATED responses, zero permitted unauthenticated tool enumeration, and 3 returned transport-level failures. Of the 20 externally-hosted servers, 10 (50.0%) were AUTH_GATED and 7 (35.0%) permitted open enumeration. The conditional probability P (AUTH_GATED | Smithery) = 0.946 compared to P (AUTH_GATED | External) = 0.500. The 44.6 percentage-point differential is consistent with authentication being strongly associated with platform hosting choice. However, the cross-sectional design of this study cannot establish causation: operators who self-select into Smithery may also be more likely to configure authentication independently, and the observed correlation could reflect both platform defaults and operator demographics. B. OAuth 2.1 at the Gateway Layer The June 2025 MCP specification mandates OAuth 2.1 with PKCE S256 for all hosted remote deployments, replacing the prior bearer-token model [1]. Smithery implements this mandate at the gateway layer: the *.run.tools domain fronts all hosted servers behind a centralized OAuth authorization server, requiring access tokens for MCP session initialization. From the perspective of an external measurement probe, this architecture is operationally indistinguishable from a developer-configured authentication decision. The finding that zero Smithery servers permit unauthenticated access, compared to 35% of external servers, is consistent with authentication being a consequence of hosting platform defaults rather than individual operator security practice. C. Compliance Characteristics of the Open Population Granular O1 compliance data was collected for the seven externally-hosted open servers. Tool count k ranged from 0 to 13 (mean 7.3, median 6), confirming that the SAFE classifications assigned by the Tier 2 pipeline are not vacuous: these servers expose substantive tool interfaces. Mean schema quality q̄ was 0.61, indicating that a majority of exposed parameters carried type declarations but that description-string completeness was mixed. All seven open servers enforced HTTPS transport. Mean composite compliance σ̄ was 0.74,
C. Formal Characterization
Fig. 4. Security posture distribution stratified by hosting infrastructure type.
indicating adequate but incomplete adherence to specification recommendations. VI. T HE S ECURITY-O BSERVABILITY T RADEOFF A. Classifier Design The third research question examined whether O0 catalog metadata and O1 passive compliance signals could predict O2 security vulnerability classifications, providing a passive pre-screening capability usable within the authentication-asplatform constraint. A Logistic Regression classifier was trained to predict the AgentSeal vulnerability verdict (SAFE vs. NEEDS_REVIEW). The O0 description string was encoded using the all-MiniLM-L6-v2 sentence-transformer model to produce a 384-dimensional semantic embedding, concatenated with the five scalar O1 features to yield a 389dimensional input vector. L2 regularization with C = 1.0 was applied. B. Single-Class Degeneracy Classification was constrained to the seven open servers for which both O1 and O2 data were available. All seven received a SAFE verdict from the AgentSeal scanner, meaning no concrete tool-poisoning payloads were detected in any of the scannable servers. The target label vector exhibits zero variance, rendering the classification task ill-defined: a zerorule classifier achieves 100% accuracy by predicting SAFE unconditionally. We acknowledge that drawing statistical conclusions from seven observations is not possible with conventional significance testing. However, the result is informative precisely because it is structurally determined: the population accessible to unauthenticated scanning is small not by sampling choice but because platform-level authentication gates the largest server population, preventing their inclusion in any live scanning dataset. The single-class outcome is therefore an empirical consequence of the ecosystem’s architecture rather than a failure of experimental design.
Let S denote the full population of remotely reachable MCP servers. Let Sauth ⊂ S denote the subset for which platformlevel authentication is enforced, and let Sopen = S \Sauth be the complement. The Tier 2 security analysis function f : Sopen → {SAFE, NEEDS_REVIEW} is defined only over Sopen . Empirically, |Sauth |/|S| = 63/76 = 0.829. In the current ecosystem state, any mechanism increasing |Sauth |/|S| reduces |Sopen |/|S|, shrinking the domain of f . The data shows this tradeoff is already in a pronounced state: 82.9% of the reachable ecosystem is opaque to Tier 2 analysis. This characterization reflects the ecosystem as measured in June 2025 and should not be interpreted as a universal property of all possible MCP deployment architectures; alternative designs (such as the transparency log mechanism proposed in Section VII-C) could decouple authentication from observability. VII. D ISCUSSION A. Centralization as a Systemic Risk The HHI of 0.736, driven by a single-ASN dominance of 85.5%, introduces systemic fragility that is independent of any individual server’s security posture. A routing disruption, BGP prefix hijack, or platform-level incident affecting the dominant ASN would degrade or compromise connectivity for the majority of the remote MCP ecosystem simultaneously. This dependency structure is analogous to the single-point-offailure vulnerability documented in centralized DNS resolver deployments [7], with the additional consequence that the dependent entities are autonomous agents rather than passive query resolvers. B. Design Implications for Gateway Operators The Security-Observability Tradeoff has direct operational consequences for developers of AI gateways and agent routing fabrics. Given that O2 active scanning is structurally unavailable for 82.9% of the reachable ecosystem in our dataset, gateway operators should design integration workflows that require credential provisioning before any security assessment phase. An AUTH_GATED classification confirms platformlevel authentication enforcement but provides no information about the safety of the tool schemas protected behind that authentication layer. Practical design responses include maintaining a local trust ledger of O0 and O1 signals aggregated during the integration handshake, requiring server operators to provide a signed tool manifest at onboarding time, and applying differential trust policies for PaaS-hosted servers (where platform-level audit commitments may exist contractually) relative to externally-hosted servers, where the 50% authentication adoption rate observed here indicates higher variance in operator security practice. C. Protocol-Level Governance The following protocol extensions are presented as design directions motivated by the measurement results, not as validated solutions. They would require further specification work and community review before adoption.
The Linux Foundation AAIF, which governs the MCP specification, could consider two extensions in a future revision. The first is a Tool Schema Transparency Log: a cryptographically-signed, append-only log of tool schema versions, structurally analogous to the Certificate Transparency log infrastructure for X.509 certificates. Under this model, a server’s tool schemas would be submitted to a public log at deployment time, producing a signed tree head verifiable by gateway operators against the schema hash returned during the OAuth handshake. This mechanism would preserve O2 -equivalent security signals in the presence of authentication, without requiring unauthenticated tool enumeration. The second is a set of standardized security metadata fields at the O0 catalog level, specifically schema_hash, scan_attestation_uri, and last_scan_timestamp, surfaced in registry API responses. These fields would permit registry-level security characterization prior to any protocol-level connection, addressing the measurement gap exposed by the single-class degeneracy without requiring changes to the server-to-client wire protocol. D. Limitations This study has several limitations. The dataset is a crosssectional measurement at a single point in time (June 2025). Given the rapid evolution of the MCP ecosystem, the hosting distribution, authentication rates, and registry composition may change substantially in subsequent months. Longitudinal measurement would strengthen the empirical support for these findings. The sample is drawn from two public registries; enterprise-internal, invite-only, and privately hosted MCP servers are not represented. The Tier 2 analysis is constrained to seven open servers, which is too small to support statistical generalization about tool-poisoning prevalence across the broader ecosystem. The correlation between hosting platform and authentication posture, while strong, cannot be interpreted as causal without a controlled experiment or longitudinal design that observes operator behavior before and after platform migration. VIII. E THICS This measurement study was conducted in accordance with responsible disclosure and ethical network measurement practice. The experimental pipeline used deterministic protocol handshakes and did not invoke the tools/call endpoint or execute any server-side tool at any point. The crawler was rate-limited to a maximum of three concurrent connections with a one-second inter-request delay per target, to minimize operational impact on measured infrastructure. All active vulnerability assessments were performed using the public agentseal CLI as intended by the vendor, targeting only publicly reachable endpoints listed in public registries. No credentials or authentication tokens belonging to third parties were used or solicited. IX. C ONCLUSION This paper presents a systematic server-side measurement of the remote Model Context Protocol ecosystem. A three-tier ob-
servability framework (O0 , O1 , O2 ) was applied to a stratified sample of 179 endpoints drawn from two primary public registries. The empirical results show high hosting consolidation at the network layer, measured by an HHI of 0.736, with a single Autonomous System (AS13335, Cloudflare) hosting 85.5% of reachable servers. Authentication enforcement is strongly correlated with hosting platform choice, with 94.6% of commercial PaaS-hosted servers enforcing OAuth 2.1 with PKCE at the gateway layer. The Security-Observability Tradeoff, as observed in this dataset, shows that 82.9% of the reachable ecosystem is opaque to automated security analysis, and that increases in platform-level authentication adoption reduce the feasible domain of Tier 2 measurement. Future work should include longitudinal measurement to track how these properties evolve, and should explore cryptographic transparency log mechanisms or standardized attestation metadata at the catalog layer to decouple authentication from observability. R EFERENCES [1] Anthropic and Linux Foundation AAIF, Model Context Protocol Specification, June 2025, https://spec.modelcontextprotocol.io. [2] K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, “Not what you’ve signed up for: Compromising real-world llmintegrated applications with indirect prompt injection,” in Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (AISec ’23). ACM, 2023. [3] A. Doe et al., “Mcpcrawler: Characterizing client-side interaction patterns in the mcp ecosystem,” 2025. [4] C. Lee and D. Kim, “Mcptox: A red-team study of llm tool poisoning,” 2026. [5] E. Wang et al., “Security of the internet of agents: A comprehensive survey,” 2025. [6] G. Patel, “A2a security analysis and vulnerability discovery,” 2025. [7] G. C. M. Moura, S. Castro, J. Heidell, W. Hardaker, M. Muller, R. de O. Schmidt, and C. Davids, “Clouding up the internet: How centralized is dns traffic becoming?” in Proceedings of the ACM Internet Measurement Conference (IMC ’20). ACM, 2020. [8] G. C. M. Moura et al., “Hosting industry consolidation,” 2021. [9] B. Smith et al., “A survey of mcp, acp, a2a, and anp interoperability protocols,” 2025. [10] F. Davis et al., “Concur: Congestion control for autonomous agents,” 2026. [11] AgentSeal, “Automated mcp vulnerability scanner,” 2026, gitHub repository, commit 1b9c2a4. [Online]. Available: https://github.com/getagentseal/agentseal