Conceptio › Archive › arXiv CS
arXiv CSopen access

MCP-GRANITE Benchmark: GRANularity Interface TEsting for MCP-Based LLM Agents

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
clouddistributed-computingparallel-computing
distributed computing, parallel computing, cloud

Author copy of paper published at 34th International Symposium on the Modeling, Analysis, and Simulation of Computer and Telecommunication System (MASCOTS2026)

MCP-GRANITE Benchmark: GRANularity Interface TEsting for MCP-Based LLM Agents Demetris Paschalides, Moysis Symeonides, George Pallis, Marios D. Dikaiakos

arXiv:2609.24161v1 [cs.DC] 21 Sep 2026

Department of Computer Science University of Cyprus Nicosia, Cyprus {dpasch01, msymeo03, pallis, mdd}@ucy.ac.cy

Abstract—As LLM agents increasingly interact with external tools through standardized protocols such as MCP, tool-interface design becomes a critical yet underexplored factor. How functionality is decomposed into tools affects whether an agent can select the right tool and construct valid arguments. This choice is especially consequential at the edge, where resource constraints limit which models can run locally and scaling up is often not an option. We present MCP-GRANITE, an open-source extensible benchmark framework that treats tool-interface granularity as a controlled variable for MCP-based agents, evaluated under edge and IoT scenarios. It comprises 81 multi-step scenarios across 9 domains, instantiated at 4 granularity levels from fine-grained primitive tools to a single tool. We evaluate 9 locally deployed models (268M-20.9B parameters) across 8,748 trials using task completion, tool selection F1, argument accuracy, latency, and resource-usage metrics. Results show that a 4-tool interface offers the best trade-off, improving task completion by 16.4% over fine-grained primitives and 33.6% over a single monolithic tool, while nearly doubling argument accuracy. Model size is only weakly correlated with task completion and strongly with latency, while its association with argument accuracy is less robust, and a 3.2B model at the optimal granularity outperforms a 20.9B model at a mismatched one. These findings identify tool-interface granularity as a key design parameter for MCP-based agents. Index Terms—Model Context Protocol, tool-interface granularity, LLM agents, local AI, resource-aware benchmarking

I. I NTRODUCTION Large language models (LLMs) are increasingly used as the reasoning core of agentic systems. These systems combine LLM planning with external tool use, such as invoking APIs, reading sensor data, and interacting with services, to ground decisions in executable contexts [1]–[4]. This shift creates the need for standardized interfaces through which agents can discover, invoke, and coordinate external tools reliably. The Model Context Protocol (MCP) [5] has rapidly emerged as the de-facto standard for connecting LLM agents to such capabilities. MCP servers expose tools as named operations, and LLM-based clients discover them at runtime by loading their schemas into the model’s context. However, MCP leaves the interface design unspecified, since the same capability can be exposed as many primitive tools, a few task-level functions, a read/write split, or a single monolithic tool. Tool-interface decomposition is, therefore, an explicit design decision that every MCP server developer must make. This decision matters in practice because the functional decomposition of agentic tools, including their quantity and level

of abstraction, directly affects how the agent selects them and populates their arguments. In modern application design, the principle that interfaces shape user success is well established. Research on API-usability has shown that abstraction level, naming, and discoverability significantly affect how developers work with software interfaces [6]. Recent work confirms that the same principle holds for LLM agents, as purpose-built tool interfaces substantially improve their performance [7], and even minor changes to the available toolkit’s composition can destabilize function calling [8]. Yet despite this evidence, prevailing tool-use benchmarks mainly evaluate whether models can select and invoke tools correctly [9]–[11]. Similarly, MCPspecific benchmarks [12], [13] treat the tool interface as fixed and assess only whether the model can use it successfully. As a result, these benchmarks overlook the role of the functional decomposition of tools. They do not systematically vary the interface decomposition of the same capability set, such as the number of tools, their abstraction level, and their argument structure, as the primary experimental factor. This omission may be less visible for cloud-based LLMs, which can often achieve strong tool-use performance without careful interface design. However, when increasing model size is not feasible, the tool interface becomes one of the few design parameters available to developers. This is especially relevant at the edge, where LLM agents are increasingly used in IoT and cyber-physical applications under strict latency, energy, and memory constraints [14]. In these settings, tool calls often mediate interactions with local sensors, actuators, and devices, while fixed hardware typically limits the feasible LLM size. In such deployments, interface decomposition becomes a practical design knob for edge-hosted LLM agents. Finegrained interfaces expose many specialized tools with clearer semantics, but increase the number of choices the model must reason over and the call sequences it must compose. Coarsegrained interfaces simplify tool selection, but require the agent to encode richer operational intent through fewer tools with larger and more complex argument schemas. These trade-offs make tool-interface granularity a consequential factor for edgehosted LLM agents, with measurable effects on performance, robustness, and energy cost. This motivates a controlled benchmark in which researchers and MCP developers can compare alternative interface decompositions of the same capability set while holding the task semantics and execution setting fixed.

In this work, we study how tool-interface granularity shapes the relationship between model scale and performance in constrained environments, making three contributions: (i) We introduce MCP-GRANITE, an open-source [15], extensible benchmark for controlled comparison of MCP tool-interface decompositions. It comprises 81 scenarios across 9 edge/IoT domains, each instantiated at 4 granularity levels; (ii) We conduct a systematic empirical study across 9 locally deployed models (268M to 20.9B parameters), showing that the granularity Level 3 interface (4 tools) consistently outperforms both fine-grained (Level 4, 8-10 tools) and highly consolidated (Levels 1-2, 1-2 tools) alternatives, with over-consolidation causing 28.2% of L1 runs to complete without invoking any tool; and (iii) We provide a scaling and resource-aware analysis combining functional metrics with GPU utilization and power monitoring, showing that model size is weakly associated with task completion (ρ=0.285), strongly associated with latency (ρ=0.946), while its association with argument accuracy (ρ=0.686) is less robust. Our results show that a 3.2B model at the optimal granularity level outperforms a 20.9B model at a mismatched one, confirming that interface design can outweigh model scale.

Aggregated Results Evaluation Report

Configs

Configuration Parser Scenario Mock Tools

Mock Server Instantiator

API Mocked Responses

Scenario Prompts

Selected LLMs

MCP Mock Server

Monitoring Server

Workload Generator

Configs & Artificial Delay

Results

Evaluation Engine

MCP-GRANITE Utilization Metrics

Tools Invocation

LLM

Agentic Service

MCP Connector

Monitoring Module LLM Weights

LLM Online Repository

II. R ELATED W ORK

Fig. 1: MCP-GRANITE Architecture

As tool-augmented LLMs mature [3], [4], [10], [16], benchmarks develop along two parallel tracks. The first evaluates function-calling correctness, progressing from small tool sets with hierarchical metrics to the Berkeley Function Calling Leaderboard (BFCL) [9], which ranks over 100 models on parallel and nested invocations, and to ToolBench [10] and ToolSandbox [11], which scale multi-step reasoning to realworld and stateful APIs. The second track evaluates broader agent capabilities in interactive environments spanning OS interaction, web tasks, and software engineering [17]–[19]. More recently, MCP-specific benchmarks begin evaluating whether models can reliably discover and invoke tools exposed through MCP servers at runtime. MCP-Universe [12] reveals that even frontier models achieve only 43.7% success across 231 real-world MCP tasks, and MCPGAUGE [13] concludes that MCP augmentation does not uniformly improve performance. ToolPlanner [20] and MTU-Bench [21] vary what they call granularity, but in both cases this refers to the specificity of user instructions or the complexity of evaluation scenarios, not to how the tools themselves are structured. These efforts collectively raise evaluation realism, yet they share a common assumption, namely that the tool interface is mostly predefined and only model capability is measured against it. Moreover, API-usability research has long shown that abstraction level, naming, and discoverability shape how effectively human developers learn and use software interfaces [6]. These concerns become more important as pipelines for automatically generating MCP servers from existing APIs become available [22], since design problems in the original API may be transferred directly to agent-facing tool interfaces. Purposebuilt tool interfaces, such as SWE-agent [7], substantially improve LLM-agent performance compared with generic alter-

natives, while even the addition of semantically related tools can destabilize function calling [8]. Recent studies further show that tool descriptions are a critical design surface, with learned rewriting of descriptions improving agent accuracy on unseen tools [23], while small description edits can disproportionately change tool selection, revealing fragility in agent behavior [24]. These observations become more important under edge constraints, where memory, latency, and energy budgets limit the use of larger models [14]. As small language models can support function calling at the edge, the tool interface becomes a critical design knob when model scaling is not feasible, while the number of available tools affects agentic performance and energy consumption [25]. Thus, these results establish that what a server exposes, and how it presents functionality to the agent, is not a background condition but a design variable with measurable effects on agent behavior. Despite this, no existing benchmark systematically varies the granularity of the same MCP-exposed capability set as the primary experimental factor, particularly for locally deployable models under realistic resource constraints. III. T HE MCP-GRANITE F RAMEWORK MCP-GRANITE addresses the previously-mentioned gap by treating tool-interface decomposition as a core design variable. It is implemented as a controlled execution and evaluation pipeline (Fig. 1) that separates tool-interface decomposition and experiment specification, tool simulation, agent execution, monitoring, and result aggregation. A YAML configuration is resolved by the Configuration Parser into the target model, scenario prompts, and the domain- and granularity-specific tool definitions that determine the interface exposed to the agent. These definitions are forwarded to the

Mock Server Instantiator, which launches one containerized MCP Mock Server per granularity level, each exposing distinct tool schemas over the same domain logic and returning fixed responses, enabling controlled and repeatable evaluation. On the target edge device, the Agentic Service runs in a containerized environment that provides a portable and architecture-agnostic execution substrate. Based on the parsed configuration, the service retrieves the specified LLM weights from HuggingFace and loads the model as the reasoning core of the agent. An MCP connector routes tool invocations from the agent to the corresponding mock server. The Workload Generator then submits scenario prompts sequentially, recording the start and end timestamp and full response for each interaction. Throughout execution, the Monitoring Module collects metrics such as GPU utilization, memory usage, and power consumption. These are forwarded to a time-series backend supporting time-range queries during post-processing. After execution completes, the Evaluation Engine compares each trace against the gold-standard solution defined for the corresponding scenario and granularity level. Based on the observed tool calls, argument values, call ordering, and any failures or timeouts, it computes the functional metrics defined in Sec. IV. Using the recorded execution interval, the engine retrieves selected resource metrics from the monitoring backend and integrates them with the trace-level results. The engine emits an Evaluation Report capturing both the agent’s functional behavior and its underlying resource footprint, enabling MCP-GRANITE to support reproducible and resource-aware benchmarking across models, domains, and tool granularities. A. Granularity Modeling and Experiment Specification The central design variable in MCP-GRANITE is the tool-interface granularity level, which controls how the same domain operations are partitioned into MCP-exposed tools. The levels systematically sample the spectrum of granularity recognized in API and service design [26], defined along two dimensions: the number of tools visible to the agent and the amount of operational intent each call must encode. The tools for each level of granularity are created within the framework, as detailed in Sec. IV and Sec. V, with code available in our repository [15]. Users can extend the benchmark with additional domains, scenarios, and granularity designs while retaining the same controlled comparison methodology. Each level is implemented as an alternative tool-exposure module over the same underlying domain logic, rather than generated automatically. Then, they specify their scenarios through a YAML-based scenario model, shown in Fig. 2. Each scenario file includes a natural-language prompt and a per-granularity gold standard, which defines the expected sequence of tool calls and arguments required to complete the task. The same user prompt is used across all levels, isolating the effect of tool-interface granularity from task complexity. To illustrate the rationale behind the granularity levels, consider a smart-home scenario (Fig. 2), where the agent is asked to activate a night security protocol. This requires setting the living room to night mode, locking the doors, enabling

id: smarthome-007 domain: smarthome difficulty: hard user_prompt: Activate the full night security protocol. Set the living room to night mode, lock all doors, enable camera recording with night vision, check front door sensor, review recent events, and create a rule that sounds the alarm if any door opens after 11pm. gold_standard: L4: # 8 expected calls using Level-4 single-purpose tools - tool: set_device arguments: { device_id: "DV003", settings: {...} } - tool: read_sensor arguments: { sensor_id: "SN005" } - tool: create_automation_rule arguments: { name: "Night door alarm", ... } # ... 5 more calls L3: # 3 expected calls using Level-3 task-level tools - tool: set_room_mode arguments: { location: "living room", mode: "night" } - tool: get_room_status arguments: { location: "front door" } - tool: manage_automation arguments: { action: "create", ... } L2: # 5 expected calls using Level-2 query/control tools - tool: smart_home_control arguments: { action: "set_mode", ... } - tool: smart_home_query arguments: { action: "sensor_reading", ... } # ... 3 more calls L1: # 5 expected calls using the Level-1 monolithic tool - tool: smart_home arguments: { action: "set_mode", ... } - tool: smart_home arguments: { action: "sensor_reading", ... } # ... 3 more calls

Fig. 2: Smart-home scenario excerpt. camera recording with night vision, checking the front-door sensor, reviewing events, and creating automation rules that trigger an alarm if a door opens after 11pm. MCP-GRANITE spans four granularity levels (L4 to L1) that progressively consolidate the same functionality into broader tools. At Level 4 (L4), the developer designs one dedicated tool per operation (e.g., set_device, read_sensor, etc.). This granularity level exposes 8-10 single-purpose primitive tools (8 in the smart-home scenario), so the agent must select among many candidates but each call requires only a few, straightforward arguments. This is analogous to finegrained CRUD (Create, Read, Update, Delete) APIs, where each function performs only one action. Level 3 (L3) applies task-class grouping, organizing tools into task-level granularity such as set_room_mode and manage_automation. In our benchmark, L3 groups tools into four functional categories: status querying, action execution, configuration, and analysis. In this level, the agent selects among fewer tools, but each call must specify both the target entity and the desired operation, requiring more complex arguments. This is analogous to resource-oriented or aggregate endpoints, where related operations are grouped under a common interface. At Level 2 (L2), the interface reduces to two tools, applying read/write separation. For example, in smarthome scenario the interface has smart_home_query and smart_home_control, separating read from write operations. The operation formerly encoded in the tool name is moved into an explicit action parameter, such as action:

models: [llama-3.2, granite-4, gpt-oss, xlam-2, qwen-3, ...] domains: [agriculture, robotics, energy, industrial, ...] granularities: [L4, L3, L2, L1] # Level 4 to Level 1 repetitions: 3 # Agent-level controls max_agent_turns: 10 agent_timeout_seconds: 360.0 temperature: 0.0 max_tokens: 4096

Fig. 3: Experiment configuration. “set_mode", so most operational intent now resides in the argument schema, increasing argument complexity. This is analogous to Command-Query Responsibility Segregation (CQRS), a standard pattern in distributed systems that separates data retrieval from state modification. Lastly, Level 1 (L1) introduces a single-entry-point design, exposing one monolithic tool through which all operations are accessed. In our scenario, a single smart_home tool handles everything. The tool name carries no operational meaning, and the agent must encode the operation type, target entity, and all parameters entirely through arguments. This is analogous to the facade or API gateway pattern, where a single interface dispatches to internal logic based on request parameters. The four levels trace a principled path along the consolidation spectrum. L4 separates by individual operation, L3 by task class, L2 by data-flow direction, and L1 removes all separation. This produces a non-monotonic difficulty profile, as consolidation reduces the tool-selection burden but progressively increases the argument-construction burden. Additionally, in our configuration, gold standards are defined separately for each granularity level, enabling the benchmark to evaluate correct tool selection and argument accuracy at each level. For instance, the 8 fine-grained calls required at L4 may collapse to 3 calls at L3, while L1 may still require repeated invocations of the same monolithic tool. Thus, consolidation may reduce the number of exposed tools but not the number of calls required, distinguishing interface size from task execution length. Lastly, an experiment is specified in a single YAML file (Fig. 3) that declares the evaluated models, target domains, granularity levels, and repetitions, together with agent-level controls. The agent-level controls define the maximum number of tool-use steps allowed per task, the execution timeout, the sampling temperature that regulates output randomness, and the maximum generation length for each model response. Internally, the framework enumerates the full Cartesian product of models × domains × granularities × scenarios and executes each combination for the configured number of repetitions. IV. I MPLEMENTATION D ETAILS MCP Server Implementation: The MCP Mock Server is implemented as a FastMCP server [5], allowing the framework to simulate realistic MCP-based tool interaction while preserving experimental control. The implementation follows a layered design that separates domain representation, state management, and tool exposure. At the lowest layer, each domain defines a set of classes describing the entities manipulated by

the server, such as sensors in agriculture, patients in healthcare, or vehicles in fleet management. On top of this layer, each domain provides a set of functions that encapsulates its business logic and operates over an in-memory store maintaining the domain state (e.g., current sensor readings, device configurations, and operational statuses). Then, a tool-exposure layer is implemented through separate server modules, one per granularity level. Across all four levels, the underlying domain logic and data store remain identical. The only difference lies in the schema and abstraction level of the tools exposed to the agent, ensuring that tool-interface granularity is the sole experimental variable. To ensure reproducibility, the store is initialized with deterministic test data so that repeated executions of the same scenario observe identical conditions. To prevent cross-run interference, the agent harness launches a fresh MCP server container for each experimental condition, ensuring that no state persists between runs and that every scenario starts from a clean and deterministic environment. Agentic Service: The Agentic Service is responsible for instantiating the evaluated LLM and mediating its interaction with the MCP-based tool environment. A key property of this service is that tool discovery is performed dynamically at runtime: at the beginning of each session, the agent retrieves the list of available tools through MCP’s tools/list method, rather than relying on statically defined or hardcoded specifications. This closely reflects realistic MCP deployments, in which agents must adapt to the tool schemas exposed by the connected server at execution time rather than operate over a fixed, preconfigured interface. The service is implemented using Google’s Agent Development Kit (ADK) 1 , with McpToolset handling communication with the MCP server and LiteLLM 2 providing a uniform abstraction for model inference. This combination allows the same benchmark pipeline to evaluate diverse LLMs through a common OpenAIcompatible interface, without requiring modifications to the orchestration logic. As a result, different model families and sizes can be compared under the same interaction workflow and with identical dynamically discovered tool schemas, ensuring a reproducible execution protocol across all models. Evaluation & Utilization Metrics: To assess agent performance, each execution is compared against the gold-standard solution defined for the tool-granularity level. Let G denote the gold sequence of tool calls and Ĝ the sequence produced by the agent. Tool Selection F1 is computed over the multiset of invoked tool names, ignoring order. Precision is the fraction of predicted tool calls that match tools in G, recall is the fraction of gold tool calls recovered in Ĝ, and F1 is their harmonic mean. Argument Accuracy (ArgAcc) measures the fraction of required gold argument key-value pairs correctly produced after aligning predicted and gold calls of the same tool in sequence order. Missing calls, arguments, or mismatched values are counted as incorrect. Task Completion (TC) is a binary metric equal to 1 when the executed trace satisfies the scenario objective according to the gold-standard state and 1

https://google.github.io/adk-docs/

2

https://github.com/BerriAI/litellm

output conditions, and 0 otherwise. Confidence Intervals (CIs) use 2,000 model-level bootstrap resamples after per-model averaging. Comparisons against L4 use paired resampling with Benjamini-Hochberg correction over 9 tests. We also report robustness and efficiency metrics. The zerotool-call rate is the fraction of runs where the agent returns a final answer without invoking any tool. The error rate captures runs terminated by parsing errors, invalid tool calls, runtime exceptions, or schema violations, while the timeout rate captures runs exceeding the execution budget. Efficiency is measured through wall-clock time, number of tool calls, and energy per successful task, computed for aggregate results as total GPU energy, estimated from mean GPU power and execution time, divided by the number of completed tasks. As evaluation uses controlled mock servers, these measurements characterize agent-side execution rather than end-toend latency or energy of live IoT deployments. To capture the mean GPU power, we deployed a monitoring stack using Prometheus and NetData on edge node3 , supplemented by a custom nvidia-smi wrapper sampling at 5s intervals. Beyond the metrics reported here, the framework collects a broader set of trace-level and resource metrics including token counts, redundant-call rates, and node-level resource telemetry, such as GPU utilization, memory, power draw, temperature, and many more. Due to space constraints, we focus on GPU utilization, power consumption, and wall-clock time as the metrics most relevant to the granularity analysis (Sec. V-D). V. E XPERIMENTAL E VALUATION We instantiate MCP-GRANITE across 9 IoT domains organized into three application areas [14]: industrial automation (industrial, robotics, warehouse), smart environments (smart home, surveillance, energy), and field operations (agriculture, fleet, healthcare). Each domain is operationalized by an inmemory data store that holds the entities the agent interacts with, such as sensors, devices, patients, or vehicles, each identified by a fixed ID (e.g., SN005, DV003). For every scenario, a gold standard specifies the expected sequence of tool calls and their arguments at each granularity level. These gold standards were verified by running them against the store and confirming that the correct final state was reached. Scenarios are organized into three difficulty tiers based on how many tool calls and distinct entities they require: easy (1-2 calls, single entity), medium (4-5 calls, 2-3 entities), and hard (9-10 calls, 4+ entities). Each of the 9 domains contributes 3 scenarios per tier, yielding 81 scenarios in a balanced design. The complete scenario definitions and benchmark construction details are publicly available in the paper’s repository [15]. We evaluate nine models spanning 268M to 20.9B parameters (Table I), served through Ollama4 . The selection includes both general-purpose and function-calling variants. Quantization varies by model: FunctionGemma uses 8-bit, xLAM-2 uses 3-bit, Qwen3 uses 6-bit, GPT-OSS uses MXFP4 (blockscaled 4-bit), and the remaining five use 4-bit. These affect 3

https://prometheus.io/ & https://www.netdata.cloud/

4

https://ollama.com/

TABLE I: Models’ overall performance. TC with 95% CI. F1, ArgAcc as mean±std. G/G+T/G+C/FC = General / +thinking / +chat / Function calling. Calls = mean tool calls per scenario. Model Qwen3 G+T Llama3.2 G Mistral-Nemo G Ministral-3 G Granite4 G Hermes3 G+C GPT-OSS G xLAM-2 FC FuncGemma FC

Params TC [95% CI]

F1

ArgAcc

14.8B .76 [.74, .79] .86±.26 .53±.36 3.2B .58 [.55, .61] .76±.33 .29±.35 12.2B .49 [.46, .52] .68±.39 .34±.35 13.9B .42 [.39, .45] .50±.45 .26±.33 3.4B .40 [.37, .43] .63±.38 .27±.34 8.0B .39 [.36, .42] .75±.35 .31±.34 20.9B .34 [.31, .37] .66±.31 .36±.37 8.0B .25 [.22, .28] .63±.36 .35±.35 268M .19 [.16, .21] .37±.42 .10±.22

Time(s) Err% Calls 164.5 5.5 6.7 0.0 17.7 1.9 27.4 10.5 7.4 0.2 8.4 0.0 25.4 55.6 9.2 52.8 2.9 0.5

3.7 2.0 1.9 2.7 1.4 2.2 3.7 2.9 1.0

per-model performance, as higher-bit quantization preserves more weight fidelity but increases memory use and inference time. All experiments ran on a GPU-enabled edge workstation with an NVIDIA Tesla T4, 16GB VRAM, 128GB RAM, and a multi-core CPU, representing an edge-server or gatewayclass deployment [14]. Our full-factorial design combines 9 models, 9 domains, 4 granularity levels, 9 scenarios, and 3 repetitions, yielding 8,748 runs (972 per model). Lastly, each configuration had the same agent-level controls as Fig. 3. A. Overall Model Performance Table I summarizes mean performance across all scenarios. Qwen3 (14.8B) achieves the best performance, with 0.762 task completion, 0.857 F1, and 0.531 argument accuracy, but at a latency of 164.5s, driven by its large parameter count and its built-in reasoning capabilities, which produce extended thinking traces before each invocation. Llama3.2 (3.2B) is the second-best model, reaching 0.579 task completion and 0.760 F1 in only 6.7s with zero errors, making it 24.5× faster than Qwen3 while retaining about 76% of its task completion. Mistral-Nemo also performs competitively, with 0.488 task completion, 0.676 F1, and a low error rate of 1.9%. In contrast, GPT-OSS and xLAM-2 show very high error rates of 55.6% and 52.8%, despite their larger size. GPT-OSS is a Mixture-of-Experts model with 20.9B total parameters but only ≈3.6B active per forward pass. Its errors are dominated by malformed tool-call outputs and schema violations. xLAM2, which uses aggressive 3-bit quantization, shows similar structured-output failures, particularly when tool schemas are discovered at runtime rather than provided statically. At the other end, FunctionGemma is the fastest model at 2.9s, but its performance remains far below the stronger models. Key takeaway: Model size is not a reliable predictor of tooluse performance. Qwen3 achieves the highest accuracy, but Llama3.2 offers a better efficiency–robustness trade-off, while GPT-OSS and xLAM-2 show that larger or function-calling models do not generalize to dynamic tool interfaces. B. Effect of Tool-Interface Granularity Across nearly all models, L3, representing moderate consolidation, is the most effective granularity configuration (Fig. 4), achieving the highest task completion for 8 of 9 models. Mistral-Nemo is the only exception, where L4 performs slightly better. L3 reaches a task completion of 0.49, a +16.4%

L4

0.8

L3

L2

TABLE III: Granularity effects by model size class.

L1

0.6

Metric

<8B | ≥8B

L4

L3

L2

L1

0.4

TC F1 ArgAcc

<8B | ≥8B <8B | ≥8B <8B | ≥8B

.395 | .440 .540 | .619 .121 | .245

.465 | .506 .646 | .673 .313 | .451

.422 | .429 .675 | .739 .299 | .420

.347 | .385 .548 | .677 .168 | .316

0.0 a .2 m a3 em am FG Ll

n4

ra

G

M

A

xL

3

3 in

o

m

er

H

em -N

M

M

Q

n3 we

O

PT

G

Fig. 4: Task completion by model and levels L1-L4 TABLE II: TC, F1, ArgAcc as mean [95% CI] from 2,000 model-level bootstrap resamples. Bold = best. †: p < 0.05; ‡: p < 0.01 vs. L4 (BH-adjusted paired bootstrap). Granularity TC [95% CI] F1 [95% CI] ArgAcc [95% CI] Zero% / Time(s) L4 (8-10) L3 (4)‡ L2 (2) L1 (1)†

0.42 [.31,.52] 0.59 [.51,.67] 0.49 [.39,.59] 0.66 [.58,.74]‡ 0.42 [.32,.53] 0.71 [.58,.83]‡ 0.36 [.25,.50] 0.63 [.52,.73]

0.20 [.15,.27] 0.40 [.33,.48]‡ 0.38 [.28,.45]‡ 0.26 [.19,.34]‡

10.2% / 33.0 10.7% / 31.0 14.6% / 30.4 28.2% / 25.1

improvement over L4 and +33.6% over L1, computed from unrounded means (Table II). Its zero-tool-call rate is slightly higher than L4 (10.7% vs. 10.2%), indicating that L3 does not improve every individual metric. It also yields the highest argument accuracy (0.40) and an F1 of 0.66, while reducing execution time. Paired bootstrap tests confirm that the L3 improvements are significant (Table II). To verify this, we fit a logistic regression including model size, quantization, domain, and difficulty, which confirms that L3 independently improves completion odds by 1.36× (CI [1.12, 1.66], p=0.002), while L1 reduces them by 22% (p=0.009). In contrast, the fully consolidated L1 setting substantially degrades performance, with task completion dropping to 0.36 (−12.9% vs L4) and zero-tool-call rates rising from 10.2% at L4 to 28.2% (Table II). Collapsing the interface into a single monolithic tool forces the model to choose among many action types and populate a larger schema in every call. The effect is most pronounced for weaker models (FunctionGemma and xLAM-2), but also reduces task completion for Granite4 and Hermes3, confirming that excessive consolidation simplifies the interface at the cost of execution reliability. Argument accuracy benefits most consistently from consolidation. Moving from L4 to L3 nearly doubles mean accuracy, from 0.20 to 0.40 (a +99.5% improvement). L2 remains similarly high at 0.38 (+85.8% over L4). Moderate consolidation reduces tool-identification burden and helps models construct more accurate arguments. Task completion also improves for 8 of 9 models, showing that the L3 advantage generalizes. A more nuanced pattern appears in tool-selection F1. Although task completion peaks at L3, the highest mean F1 is achieved by L2 (0.71, +20.8% over L4). This likely reflects the smaller decision space at that granularity, where choosing between two broad tools is easier than selecting among four or more specialized ones. By contrast, the F1 advantage of L1 over L4 is small (+6.7%) and is not significant under the paired bootstrap comparison (Table II), so collapsing the interface to a single tool does not preserve the F1 gains of moderate consolidation. The higher F1 of L2 does not translate into the highest task completion, since successful execution depends

on accurately specifying arguments within the selected tool. This suggests that correct tool selection alone is insufficient and that argument construction remains a bottleneck. This is consistent with function-calling evaluations that distinguish selection correctness from argument accuracy and find that they can fail independently [9], [13]. Key takeaway: Moderate tool consolidation with L3 is the most effective design choice, delivering the best balance of task completion, argument accuracy, and latency. Overconsolidation increases failures and confirms that argument construction, not tool selection, is the main bottleneck. C. Scaling Analysis Fig. 5 plots each model at the granularity level where it achieves its highest task completion against the corresponding wall time. Three models define the Pareto frontier: FunctionGemma (3.1 s, TC = 0.21), Llama 3.2 (6.6 s, TC = 0.61), and Qwen 3 (167 s, TC = 0.79), with all others dominated on at least one axis. Among the dominated models, Hermes 3 and Granite 4 are closest to the frontier, offering competitive task completion at moderate latency. In contrast, GPT-OSS and Ministral 3 incur much higher wall times without proportional gains, reinforcing that model scale alone does not determine the best edge-deployment trade-off. Qwen3

Best Task Completion

0.2

0.8

Llama3.2 Granite4

0.6

M-Nemo Hermes3

0.4 0.2

Min3

GPT-OSS

xLAM

FGemma L4

L3

Pareto frontier

0.0 10

1

102

Wall Time at Best Level (s)

Fig. 5: Best task completion vs. time (log scale) Mean Wall Time (s)

Task Completion

1.0

Qwen3

102 Min3 Granite4

xLAM GPT-OSS M-Nemo

101 Llama3.2

FGemma 100

Hermes3

101

Parameter Count (Billions)

Fig. 6: Mean wall time vs. model size (log-log). Within the nine evaluated models, parameter count is more strongly associated with argument accuracy and latency than with task completion. Parameter count correlates positively with argument accuracy (ρ=0.686, p=0.041), indicating higher argument accuracy among larger models in this sample, and even more strongly with wall-clock time (ρ=0.946, p<0.001). For example, average execution time increases from 2.9s for FunctionGemma to 164.5s for Qwen3, a 56× increase. Fig. 6 further shows that latency grows roughly

L4

L3

L2

L1

GPU Power (W)

GPU Utilization (%)

100 75 50 25 0 a

m em

FG

n4

.2

a3

am

Ll

ra

G

M

A xL

m er

H

3

3 in

o

em

-N M

M

n we

Q

3

80 L4

L2

L1

40 a

O

PT

G

L3

60

m

em

FG

.2 a3

m

a Ll

n4

ra

G

M

A

xL

er

3 m

H

o em

-N

3

3

n we

in

M

M

Q

-O

PT

G

Fig. 7: Mean GPU utilization by model & level L4-L1.

Fig. 8: Mean GPU power draw by model & level L4-L1.

linearly with parameter count on the log-log scale, except for Qwen3, which is slower due to extended reasoning traces, and GPT-OSS, which is faster than its 20.9B parameters suggest because its MoE architecture activates only ≈3.6B parameters per forward pass. As a sensitivity check, excluding Qwen3 and GPT-OSS reduces the argument-accuracy association to ρ=0.360 (p=0.427), while the latency association remains strong (ρ=0.991, p<0.001). In contrast, parameter count is only weakly and non-significantly associated with task completion (ρ=0.285, p=0.458), showing that scale alone does not determine end-to-end agent success. To examine if tool-interface granularity depends on model scale, Table III groups models into small (< 8B) and large (≥ 8B). Both benefit from moderate consolidation, with L3 achieving the best task completion. The gain is slightly higher for small models, 17.7% versus 15.0% for large models. At L2, small models improve by 6.8%, while large models decrease by 2.5%, suggesting that smaller models benefit more from a reduced tool-selection space, whereas larger models better exploit the structure preserved at L3. L1 degrades both groups, confirming that over-consolidation is harmful across scale. These trends may also reflect quantization and training differences, since FunctionGemma uses 8-bit quantization while xLAM-2 uses aggressive 3-bit quantization, which may reduce structured-output reliability. Thus, these size associations are descriptive and should not be interpreted independently of quantization, architecture, and tool-use alignment. Key takeaway: Model size alone is a poor predictor of endto-end agent success. Its clearest association is with latency, while the observed relationship with argument accuracy is less robust. Smaller models appear more sensitive to tool-interface design and benefit more from moderate consolidation.

TABLE V: Efficiency (TC/second, mean ± std) by model and granularity. Bold = best granularity per model.

TABLE IV: Joules per task under L3. Power, Time, and J/task are mean±std. TC is shown with its 95% confidence interval. Model

Power (W)

Llama3.2 Granite4 Hermes3 FuncGemma xLAM-2 Mistral-Nemo Ministral-3 GPT-OSS Qwen3

58.2±14.5 6.5±2.9 58.1±13.5 8.5±7.3 60.1±13.2 8.4±4.2 41.1±10.0 3.1±0.7 59.6±12.9 10.2±9.3 62.4±11.0 17.6±18.5 63.6±10.8 28.4±33.5 63.5±9.2 28.6±20.4 68.8±1.6 167.0±94.9

Time (s)

TC [95% CI]

J/task

0.61 [.56, .66] 621±585 0.53 [.46, .59] 937±1,219 0.44 [.38, .50] 1,148±1,439 0.21 [.16, .26] 618±1,232 0.33 [.28, .39] 1,826±3,103 0.51 [.44, .57] 2,170±3,157 0.53 [.46, .59] 3,427±5,220 0.44 [.38, .50] 4,128±5,547 0.79 [.74, .84] 14,544±11,173

D. Resource Utilization and Energy Efficiency Fig. 7 and Fig. 8 summarize GPU utilization and power consumption across models and granularity levels. Both metrics are primarily driven by model size: GPU utilization ranges

Model

L4

L3

L2

L1

Llama3.2 .081±.101 .093±.099 .078±.094 .100±.126 FuncGemma .053±.129 .067±.134 .070±.152 .069±.156 Granite4 .046±.083 .062±.085 .060±.089 .046±.074 Hermes3 .049±.081 .053±.077 .043±.071 .040±.069 xLAM-2 .024±.062 .033±.066 .026±.054 .024±.048 Mistral-Nemo .028±.054 .029±.054 .026±.048 .027±.052 Ministral-3 .014±.048 .019±.063 .013±.035 .015±.056 GPT-OSS .011±.032 .015±.041 .014±.035 .013±.032 Qwen3 .004±.005 .005±.007 .005±.006 .005±.007

from 8-17% for FunctionGemma to 96-97% for Qwen3, while GPU power increases more modestly, from roughly 40W to 69W. Power scales sub-linearly with size, as Qwen3 is about 55× larger than FunctionGemma but draws only 1.7× more power, because smaller models leave GPU compute capacity idle. VRAM usage is dominated by the loaded model weights and therefore varies little across granularity levels. Granularity affects utilization mainly when the GPU is under-saturated. For large models, utilization changes by less than 3% across levels, while smaller models show larger swings (e.g., FunctionGemma rising from 8.3% at L4 to 17.6% at L2). GPU power changes by at most ∼5W per model. Hardware cost is thus dominated by model choice, with granularity modulating under-saturated edge models. To assess end-to-end efficiency, we compute GPU energy per successful task from mean GPU power and wall-clock time (Table IV). Llama3.2 is the most favorable operating point at 621J/task under L3, whereas Qwen3 requires 14,544J/task, a 23× higher cost. FunctionGemma appears competitive (618J/task), but completes only 20.6% of tasks, so much of its energy is wasted on failed tasks, while Llama3.2 consumes similar energy but completes 3× more tasks. Larger models improve capability at disproportionate time and energy cost, while midsized models offer the best trade-off (Table V). Key takeaway: Resource and energy efficiency are driven by model choice, with granularity as a secondary factor. Longer execution times dominate energy cost since power scales sublinearly with size. Granularity mainly affects under-saturated small and mid-sized models. E. Domain Difficulty and Failure Analysis Fig. 9 shows strong variation across both models and domains: Qwen3 performs best across most domains, while healthcare (0.202) and robotics (0.244) have the lowest mean completion and industrial (0.608) and energy (0.596) the highest. Despite these differences, tool-selection F1 remains stable (0.614-0.667), indicating that models generally identify

0.30

0.45

0.24

0.00

0.32

0.00

0.14

0.03

0.19

Llama3.2

0.65

0.77

0.68

0.22

0.83

0.20

0.64

0.31

0.56

Granite4

0.41

0.68

0.42

0.28

0.46

0.22

0.44

0.25

0.44

xLAM

0.18

0.36

0.25

0.14

0.44

0.29

0.21

0.07

0.31

Hermes3

0.37

0.46

0.52

0.22

0.69

0.17

0.45

0.18

0.42

M-Nemo

0.47

0.63

0.53

0.27

0.86

0.31

0.45

0.31

0.55

Min3

0.53

0.55

0.54

0.06

0.53

0.32

0.23

0.48

0.51

Qwen3

0.94

0.94

0.83

0.48

0.83

0.44

0.92

0.77

0.72

GPT-OSS

0.42

0.35

0.47

0.15

0.49

0.24

0.31

0.31

0.29

1.0

R EFERENCES

0.8

[1] G. Mialon, R. Dessì, M. Lomeli, C. Nalmpantis, R. Pasunuru, R. Raileanu, B. Rozière, T. Schick, J. Dwivedi-Yu, A. Celikyilmaz et al., “Augmented language models: a survey,” arXiv preprint arXiv:2302.07842, 2023. [2] Z. Xi, W. Chen, X. Guo, W. He, Y. Ding, B. Hong, M. Zhang, J. Wang, S. Jin, E. Zhou et al., “The rise and potential of large language model based agents: A survey,” Sci. China Inf. Sci., 2025. [3] S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao, “ReAct: Synergizing reasoning and acting in language models,” in ICLR, 2023. [4] T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom, “Toolformer: Language models can teach themselves to use tools,” in NeurIPS, 2024. [5] Anthropic, “Model context protocol: Specification,” https: //modelcontextprotocol.io/specification, 2024. [6] B. A. Myers and J. Stylos, “Improving API usability,” Commun. ACM, 2016. [7] J. Yang, C. E. Jimenez, A. Wettig, K. Narasimhan, S. Yao, and O. Press, “SWE-agent: Agent-computer interfaces enable automated software engineering,” in NeurIPS, 2024. [8] E. Rabinovich and A. A. Tavor, “On the robustness of agentic function calling,” in TrustNLP, 2025. [9] S. G. Patil, H. Mao, C. Cheng-Jie Ji, F. Yan, V. Suresh, I. Stoica, and J. E. Gonzalez, “The berkeley function calling leaderboard (bfcl): From tool use to agentic evaluation of large language models,” in ICML, 2025. [10] Y. Qin, S. Liang, Y. Ye, K. Zhu, L. Yan, Y. Lu, Y. Lin, X. Cong, X. Tang, B. Qian et al., “ToolLLM: Facilitating large language models to master 16000+ real-world APIs,” in ICLR, 2024. [11] J. Lu, T. Holleis, Y. Zhang, B. Aumayer, F. Nan, H. Bai, S. Ma, S. Ma, M. Li, G. Yin et al., “Toolsandbox: A stateful, conversational, interactive evaluation benchmark for llm tool use capabilities,” in NAACL, 2025. [12] L. Yang et al., “MCP-Universe: Benchmarking large language models with real-world model context protocol servers,” arXiv preprint arXiv:2508.14704, 2025. [13] W. Song, H. Zhong, Z. Ding, J. Xue, and Y. Li, “Help or hurdle? rethinking model context protocol-augmented large language models,” arXiv preprint arXiv:2508.12566, 2025. [14] Z. Lin et al., “A review on edge large language models: Design, execution, and applications,” ACM Comp. Surveys, 2025. [15] P. et al., “MCP-GRANITE Benchmark,” 2026. [Online]. Available: https://github.com/dpasch01/mcp-granite [16] S. G. Patil, T. Zhang, X. Wang, and J. E. Gonzalez, “Gorilla: Large language model connected with massive APIs,” in NeurIPS, 2023. [17] X. Liu, H. Yu, H. Zhang, Y. Xu, X. Lei, H. Lai, Y. Gu, H. Ding, K. Men, K. Yang et al., “AgentBench: Evaluating LLMs as agents,” in ICLR, 2024. [18] S. Zhou, F. F. Xu, H. Zhu, X. Zhou, R. Lo, A. Sridhar, X. Cheng, T. Ou, Y. Bisk, D. Fried et al., “WebArena: A realistic web environment for building autonomous agents,” in ICLR, 2024. [19] C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. Narasimhan, “SWE-bench: Can language models resolve real-world GitHub issues?” ICLR, 2024. [20] Q. Wu, W. Liu, J. Luan, and B. Wang, “ToolPlanner: A tool augmented LLM for multi granularity instructions with path planning and feedback,” in EMNLP, 2024. [21] P. Wang, Y. Wu, Z. Wang, J. Liu, X. Song, Z. Peng, C. Zhang, J. Peng, G. Zhang, H. Guo et al., “Mtu-bench: A multi-granularity tool-use benchmark for large language models,” in ICLR, 2025. [22] M. Mastouri, E. Ksontini, A. Barrak, and W. Kessentini, “From rest to mcp: An empirical study of api wrapping and automated server generation for llm agents,” 2026. [Online]. Available: https: //arxiv.org/abs/2507.16044 [23] R. Guo, K. Dong, X. Gao, and K. Das, “Learning to rewrite tool descriptions for reliable llm-agent tool use,” arXiv:2602.20426, 2026. [24] K. Faghih, W. Wang, Y. Cheng, S. Bharti, G. Sriramanan, S. Balasubramanian, P. Hosseini, and S. Feizi, “Tool preferences in agentic llms are unreliable,” in EMNLP, 2025. [25] V. Paramanayakam, A. Karatzas, I. Anagnostopoulos, and D. Stamoulis, “Less is more: Optimizing function calling for llm execution on edge devices,” in DATE, 2025. [26] O. Zimmermann, M. Stocker, D. Lübke, U. Zdun, and C. Pautasso, Patterns for API Design: Simplifying Integration with Loosely Coupled Message Exchanges. Addison-Wesley Professional, 2022.

0.6 0.4 0.2

Task Completion

FGemma

0.0

i t t t t gr ergy lee lth ar us obo ea Ind F Sm R H En

A

v

r Su

e ar W

Fig. 9: Task completion by model and domain the right tools but fail during multi-step reasoning and argument construction. This pattern holds at the scenario level, where task completion ranges from 0.852 (warehouse-001) to 0.009 (healthcare-003), yet even the hardest scenarios achieve moderate F1 (0.596-0.671). Across all 8,748 conditions, 13.6% result in errors and 1.3% in timeouts. Failures are highly model-dependent: GPT-OSS and xLAM-2 produce error rates of 55.6% and 52.8%, while Llama3.2 and Hermes3 produce none. Tool-interface granularity also affects robustness, with L3 yielding the lowest error rate (11.4%), followed by L4 (12.9%), L2 (14.1%), and L1 (15.8%). The degradation at L1 is mainly driven by zero-tool-call behavior (28.2% of L1 conditions), where models either hallucinate completion without invoking any tool or fall back to memorized tool names that do not match the MCP schema. Key takeaway: Agent performance is limited less by tool identification than by reliable action sequencing. L3 provides the most robust behavior, while L1 produces qualitative breakdown rather than gradual degradation. VI. C ONCLUSION AND F UTURE W ORK MCP-GRANITE is an extensible, open-source framework for studying MCP tool-interface granularity, evaluated across 9 models and 9 edge/IoT domains on a T4-based platform with controlled mock-server responses. Moderate consolidation (L3, 4 task-class tools) performs best, improving task completion by 16.4% over fine-grained tools and 33.6% over a monolithic interface, primarily through higher argument accuracy. Smaller models benefit most, while hardware cost is driven mainly by model choice: at L3, a 3.2B model uses 23× less energy per successful task than the best-performing 14.8B model. Over-consolidation reduces task completion by 12.9%, with 28.2% of L1 runs invoking no tool. Overall, our experimental results favor compact, task-semantic interfaces and evaluating granularity alternatives before deployment. Our future work will extend the benchmark to live MCP servers, broader model and edge-hardware coverage, and use the collected metrics to model how interface complexity affects agent behavior and formalize tool-granularity design. Acknowledgments: This work was co-supported by Pharos-CY, funded by the EuroHPC JU (GA: 101263007), with support from the EU’s Horizon Europe programme and the Government of the Republic of Cyprus, and by the EU Commission through the AI-DAPT project (HORIZON-CL4-2023-HUMAN01-01, GA: 101135826). Language was refined using ChatGPT; all content and ideas are the authors’ own.

Record · ID 1028658 · SHA-256 421718c6a1cc9852
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.