Conceptio › Archive › arXiv CS
arXiv CSopen access

Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe Dain Kim∗

Eungi Cho∗

Kyumin Kim∗ Shinyeong Noh∗ Kyuseong Lim∗ LG CNS {dainkim, eungizoa, kyuminkim, sy.noh, ks.lim}@lgcns.com ∗

Equal contribution.

USER

arXiv:2609.05395v1 [cs.AI] 4 Sep 2026

Abstract Data-sovereignty regulations increasingly require public institutions to deploy open-source, on-premise LLM agents that chain multiple tool-calls across live government APIs. However, open-source models consistently underperform in this multi-step setting, and no existing benchmark measures the gap. We introduce the Korean Open Public API Benchmark (KOPA-B ENCH), comprising 145 real-world tasks. To close this gap, we present EDGE, an Execution-grounded Dynamic Graph for toolcalling data synthEsis driven by live execution. EDGE builds a graph of how each tool’s output can feed another’s input, keeps only the links that succeed when actually called against the live APIs, and traverses these verified links to synthesize executable multi-step trajectories. Fine-tuned via GRPO on the resulting dataset, our 9B model nearly matches the untuned 27B model from the same family, improving substantially not only on KOPA-B ENCH but also on the BFCL benchmark.1

1

Introduction

Public institutions are increasingly adopting agents built on large language models (Tang et al., 2025; Mohammadi et al., 2025). Data-sovereignty regulations force these agencies to deploy open-source models (Qwen Team, 2026; Gemma Team, 2026) on local infrastructure rather than commercial cloud services. However, these models must operate over the laws, corporate disclosures, and transportation data that the institutions provide through open APIs, a context where open-source models consistently underperform. We study this by examining how tool-calling agents operate over live Korean public API endpoints. Under this context, we observe two distinct challenges: (1) frequent entity code-lookups that Code and data are available at: https://github.com/ dneirfi/EDGE-KOPA. 1

How many special-education students were enrolled in middle schools in Jeju as of 2021? AGENT · TOOL CALL

AGENT · TOOL CALL

A

✓

API RESPONSE

api_key "xxxx" — retrieved Output binds to the next call’s input.

✗ Agent needs to retrieve the API key before making the call.

! API RESPONSE

AGENT · TOOL CALL

If the agent successfully retrieves the API key and passes it to the tool, … A

class_and_student_status( KEY="xxxx" trgtYr="2021" cdoeNm="Jeju" scclNm="Middle school" numOfRows=10 pageNo=1) ✓

A

class_and_student_status( KEY="xxxx" class_and_student_status trgtYr="2021" KEY="xxxx" cdoeNm="Jeju" trgtYr="2021" class_and_student_status scclNm="Middle school" cdoeNm="Jeju") KEY="xxxx" numOfRows=10 scclNm="Middle trgtYr="2021" school“, pageNo=1) numOfRows=10 cdoeNm="Jeju") pageNo=2) scclNm="Middle school“, pageNo=3)

Invalid API-key error

AGENT · TOOL CALL

A

get_education_api_key()

class_and_student_status( KEY="get_education_api_key" trgtYr="2021" cdoeNm="Jeju" scclNm="Middle school")

✓

API RESPONSE

[{schlNm: "AA Middle School", [{schlNm: “AA40, Middle School”, totalStdCnt: [{schlNm: “AA40, Middle School”, totalStdCnt: speStdCnt: 3, totalStdCnt: speStdCnt:…}, 3, 40, …},{schlNm: speStdCnt: 3, …] …},{schlNm: …}, …] …},{schlNm: …},

API RESPONSE

[{schlNm: "AA Middle School", totalStdCnt: 40, speStdCnt: 3, …},{schlNm: …},

…]

…] AGENT · FINAL ANSWER

A

AGENT · FINAL ANSWER

“23 special-education students.”

“75 special-education students.”

✗ The agent prematurely answered

✓ Successfully aggregated records

despite remaining pages.

(a) Failure case of small models

A

from multiple pages.

(b) Trajectory of fine-tuned model

Figure 1: Comparison of tool-calling trajectories in Korean public APIs before and after fine-tuning. (a) Failure case (Before fine-tuning): The agent fails to retrieve the API key and, even after successfully retrieving a valid key, it answers incompletely based only on the first page of a paginated API response. (b) Success case (After fine-tuning): The agent successfully chains tool outputs, queries multiple pages, and aggregates the results to deliver an accurate answer.

transform simple queries into dependent chains, and (2) high-cardinality responses requiring data reduction or fan-out. Open-source models consistently stumble on both, either skipping prerequisite lookups or prematurely answering from partial results when facing multi-record outputs (Figure 1). We therefore construct the Korean Open Public API Benchmark (KOPA-B ENCH), 145 multi-step tasks over six domains of live Korean public APIs. Each task requires chaining dependent calls, where the output of one tool supplies the input of the next, and a lot of calls return multiple records. Improving

on these tasks requires training data under the same constraints, yet existing tool-calling datasets target neither Korean public APIs, live endpoint failures, nor high-cardinality handoffs. We therefore propose Execution-grounded Dynamic Graph for tool-calling data synthEsis (EDGE), a pipeline designed around these properties. EDGE builds a tool dependency graph and verifies each LLM-proposed dependency against the live APIs, keeping the graph dynamic by pruning those that repeatedly fail under execution. From the verified graph, EDGE synthesizes executable trajectories by labeling each output-toinput handoff between tools with the number of distinct returned values. This lets EDGE turn highcardinality responses into valid sequential trajectories through bounded fan-out or deterministic reduction. Trained on the resulting dataset, both small Qwen3.5 models improve substantially, raising KOPA-B ENCH pass@1 by +13pp (4B) and +10pp (9B), with the gains extending to the out-ofdistribution BFCL benchmark (Patil et al., 2025). Our contributions are as follows: • We construct KOPA-B ENCH, a multi-step function-calling benchmark on live Korean public APIs, where each task chains dependent calls through code-based lookups and high-cardinality intermediate results. • We propose EDGE, a synthesis pipeline that verifies tool dependencies by live execution and resolves high-cardinality handoffs into sequential trajectories. GRPO training on its data substantially improves small opensource models on KOPA-B ENCH and out-ofdistribution benchmarks.

2

Benchmark Construction

We introduce Korean Open Public API Benchmark KOPA-B ENCH, a multi-step function-calling benchmark over Korean Open Public APIs, comprising 145 tasks across 10 platforms and six domains. Each task requires chaining tool-calls, where one call’s output feeds the next’s input. Unlike prior benchmarks built on emulated services (Patil et al., 2025) or hand-authored simulations with an LLM user simulator (Barres et al., 2025), KOPA-B ENCH is grounded in real public APIs.

2.1

Platform Selection

We select platforms that satisfy three criteria: (1) domain coverage across traffic, finance, education, law, politics and district administration—domains where public institutions rely on open APIs; (2) API interconnectivity, enabling chained multi-step reasoning; and (3) data licensing permitting derivative works under the Korea Open Government License (KOGL). This yields 10 platforms spanning six domains. On average, a task requires five tool-calls (up to 14), with 59% involving parallel execution. Per-domain statistics and the full platform list are provided in Appendix D. 2.2

Tool Construction

For each platform, we parse the official documentation to implement a Model Context Protocol (MCP) server that exposes typed function signatures with natural-language descriptions, allowing direct invocation via standard LLM function-calling interfaces. The server integrates an API client handling authentication, session management, error handling, and retry logic. It also manages per-task environment state, initializing it at the start of each evaluation episode to support ENVIRONMENTbased evaluation (§2.4). In total, this yields a tool inventory of 2,318 functions across the 10 platforms, each backed by a live endpoint. 2.3

Task Generation

Based on the deployed tools, domain experts design chainable tool sets for each domain and formulate queries that necessitate multi-step tool invocations. Each query is annotated with golden actions, the ground-truth optimal trajectory, as well as the target response or environment state for evaluation. We validate the annotations in two stages. First, an expert who is not among the task authors audits every task against the tool descriptions, checking tool selection, arguments, and tool responses. Second, we execute each golden trajectory and give the resulting tool outputs to a frontier model (Claude Sonnet 4.6), then compare its response against the annotated target. The audit finds errors in 7 tasks (4.8%), and, excluding cases of model failure, 80% of tasks pass the execution check on the first attempt. We correct every task flagged by either stage and repeat the check until all 145 pass. 2.4

Evaluation Methodology

We evaluate along three dimensions—RESPONSE, ENVIRONMENT, and ACTION—corresponding

to the model’s final answer, the resulting server state, and the executed tool-calls. Following τ 2 Bench (Barres et al., 2025), each task carries reward_basis annotation that designates the subset of metrics by which it is scored. ACTION is always included, while RESPONSE and ENVIRONMENT follow the response-based and state-based paradigms of BFCL: RESPONSE applies when the task demands a definite answer, and ENVIRONMENT when it depends on the resulting system state. These complementary criteria can be applied individually or in tandem. Evaluation criteria. ACTION-only evaluation is insufficient (Appendix B.4), as a trajectory may contain the golden action yet leave the system in an incorrect state, or conversely reach the correct outcome via an alternative sequence. We therefore evaluate along two separate axes: RESPONSE and ENVIRONMENT verify the outcome, while ACTION scores the trajectory that produced it. Detailed procedures are in Appendix B, and evaluation prompts in Appendix J.4.

3

Methods

We introduce Execution-grounded Dynamic Graph for tool-calling data synthEsis (EDGE), a twophase framework that synthesizes multi-step toolcalling data from live APIs (Figure 2). Phase A (§3.2) builds a dependency graph, keeping only edges that execute successfully. Phase B (§3.3) assembles trajectories over the graph, typing each junction by response cardinality, and generates a Korean query and answer for each. 3.1

Problem Setting

We construct a tool inventory V = {v1 , . . . , vN } comprising N =2,318 tools across six domains. Each tool v ∈ V exposes a live endpoint and a JSON-typed signature σv = (descv , Iv , Ov ), which consists of a Korean description, an inputparameter schema and an output-field schema. Invoking v returns a set of records Rv whose cardinality ranges from zero to hundreds of thousands. Thus, a chain u → v (where u, v ∈ V) has no single output to bind, a challenge that we address in §3.3. 3.2

Phase A: Execution-Grounded Dynamic Graph Construction

Skeleton graph construction. We take the tool inventory V in §3.1 as the nodes of a directed dependency graph G = (V, E) (Figure 2a), where an

edge e = (u → v) asserts that an output field of u binds to an input parameter of v. Since scoring all O(|V|2 ) pairs with an LLM is infeasible, a dense retriever over signature embeddings restricts each source u to its top-K neighbors: 15 from the same domain and 10 from other domains, where bindings are rarer. For each candidate edge, a single LLM call returns a feasibility score se ∈ [0, 1], a binding set Be that maps bound parameters of v to the supplying output of u, and default candidates De for left unbound. These compile into a source plan over every parameter p ∈ Iv :   U PSTREAM ⟨Be (p)⟩ p ∈ dom(Be ), Πe (p) = D EFAULT ⟨dp ⟩ p generic,   C LARIFY ⟨De (p)⟩ otherwise, (1) where dom(Be ) denotes the set of parameters bound by Be . U PSTREAM binds p to its supplying output Be (p) from the upstream tool u, D EFAULT is a static lookup of generic parameters (API keys, etc.), and C LARIFY tries De (p) in the LLM’s proposals. We admit an edge only if its feasibility score meets the threshold. Each admitted edge receives a feasibility-anchored Beta prior αe = 1 + κse , βe = 1 + κ(1 − se ), where pseudo-counts αe and βe favor success and failure in proportion to se , with κ controlling the prior’s strength. Execution-grounded graph update. Each iteration samples a batch of paths, executes them against live APIs, and revises the edge posteriors and the graph from the outcomes (Figure 2a). The graph is dynamic, with both its posteriors and its topology changing as execution evidence accumulates. We select edges by Thompson sampling (Agrawal and Goyal, 2013), a posteriorsampling approach that balances exploitation and exploration. We model the success rate of each edge e as θe with a Beta posterior Beta(αe , βe ), whose pseudo-counts αe and βe accumulate the successful and failed executions of the edge. From the current node, we sample θ̃e ∼ Beta(αe , βe ) per outgoing edge and follow the largest, with an ε-greedy fallback to a random edge: ( arg maxe θ̃e w.p. 1 − ε, ⋆ e = a random outgoing edge w.p. ε. (2) Because each θ̃e comes from the full posterior rather than its mean, an edge with a high posterior

Figure 2: Overview of EDGE. (a) Phase A constructs an execution-grounded dynamic graph: from a skeleton of LLM-proposed candidate edges, it revises both edge posteriors and its topology using live execution evidence, yielding a refined graph G⋆ . (b) Phase B assembles trajectories over G⋆ by junction type and generates a naturallanguage query for each, producing the training dataset.

mean usually yields a large sample and is exploited often, while an edge with few trials has a wide posterior that occasionally yields a large sample and is therefore still explored. The ε-greedy term sets a floor on exploration, so cold-start edges that Thompson sampling alone might never select are eventually tried. To run a path, the first tool receives no upstream output, so its arguments come from a single LLM call at execution time (with a rule-based fallback), and every subsequent argument is filled by the source plan Πe of Eq. (1). Each outcome is structural when the binding itself fails or environmental when a transient server-side error occurs, and structural failures raise βe far more than environmental ones, because random server errors would otherwise weaken and eventually prune correctly bound edges. An edge is pruned when the posterior probability that its success rate exceeds τprune falls below εprune , after at least nmin trials. Iterating this loop to convergence yields the refined graph G⋆ , which retains only the admitted edges that survive pruning. The full update scheme, hyperparameters, and supporting analyses are provided in Appendix E. Phase B (§3.3) consumes G⋆ and a per-edge cache C of observed (input, first-record) pairs to ground query synthesis. 3.3

Phase B: Type-Based Trajectory Synthesis

Sequential trajectories. Given the validated edge graph (§3.2), we synthesize a sequential trajectory as a chain of tool-calls (T1 , . . . , Tm ). The chain starts from an initial tool-call and follows

field-level edges in the graph. For a junction ji = (Ti .fi → Ti+1 .ai ), values returned in the output field fi of Ti are used to instantiate the target argument ai of Ti+1 . The required arguments not determined by a junction are filled from the validated cache C. A key challenge with Korean Open Public APIs is that tools often return more than thousands of records. At this scale, a junction no longer specifies a unique downstream binding. Instead, it creates many possible continuations that can make the synthesized trajectory ambiguous, invalid, or unanswerable if left unresolved. Therefore, we introduce a junction typing scheme based on the cardinality of the connecting field. With fan-out budget ϕ = 5, we assign Pure Sequential (S EQ) if ni = 1, Fan-out (FAN) if 2 ≤ ni ≤ ϕ, and Derived (D RV) if ni > ϕ. S EQ passes the single value to the next tool, while FAN issues one downstream call for each value in Ui . For D RV, we insert an internal process node that selects a bounded subset before proceeding. We use max, min, ≥, and ≤ for numeric fields, equality for enum fields, date-after for date fields, and most-frequent as a type-agnostic operator. For threshold and equality filters, the pivot value is chosen from observed field values so that the filtered set contains at most ϕ distinct values. The structural pattern of a sequential trajectory is the ordered composition of its junction types, e.g., S EQ+FAN, which determines the query skeleton used later. While pipelines that assume single-value flows typically discard or mis-bind many-record responses, cardinality typing safely consumes them via bounded enumeration (FAN) or deterministic

KOPA-B ENCH Model Proprietary Claude Sonnet 4.6 (Anthropic, 2026) GPT-5.1 (OpenAI, 2025) Large open-source Qwen3.5-27B gemma-4-26B-A4B-it EXAONE-4.5-33B (Choi et al., 2026) Small open-source Qwen3.5-9B gemma-4-E4B-it Qwen3.5-4B Ministral-3-8B-Instruct-2512 (Liu et al., 2026) Qwen3.5-4B (Ours) Qwen3.5-9B (Ours)

BFCL

pass@1

pass@4

Action

MultiTurn

Single (non-live)

Single (live)

0.4655 0.3706

0.8207 0.4690

0.4269 0.5881

60.13 38.12

84.98 82.96

76.83 65.21

0.4482 0.3534 0.2586

0.5655 0.4690 0.3862

0.4213 0.5061 0.2670

65.38 55.38 53.00

89.65 81.23 87.52

83.35 80.16 80.16

0.3275 0.1862 0.1758 0.1299

0.4690 0.2207 0.3103 0.2308

0.3407 0.4069 0.2140 0.3797

52.25 20.75 49.08 24.50

82.40 84.79 79.81 82.38

78.09 63.43 78.02 74.54

0.3094 0.4310

0.4690 0.5517

0.3462 0.3655

53.12 58.12

79.96 87.65

76.61 81.57

Table 1: Results on KOPA-B ENCH and BFCL. KOPA-B ENCH — pass@k: fraction of the 145 tasks solved by at least one of k rollouts, so pass@1 is the fraction solved on a single attempt (averaged over the four rollouts) and pass@4 the fraction solved by at least one of the four; Action: tool-call accuracy over golden actions. BFCL scores are computed via the official evaluation repository. The bottom block reports our fine-tuned models.

reduction (D RV). Non-sequential trajectories. We additionally synthesize three non-sequential types from explicit templates, retaining only successfully executed paths. Semantic-parallel (S EM) calls independent tools with shared arguments and combines their outputs. Comparison (C MP) calls the same tool with different arguments and compares the results. Conditional (C OND) evaluates the condition over an intermediate result and proceeds along the branch that outcome specifies. These templates cover parallel, comparative, and conditional patterns that cannot be represented as a single linear chain. Query generation and validation. For each trajectory, an LLM query generator produces a Korean natural-language query and derives the answer from the cached execution results, guided by typespecific prompt templates that describe the trajectory structure. The resulting query, answer, and execution trace then pass through a three-stage validation pipeline. Throughout, only open-source model outputs become training labels, while proprietary models serve solely for verification. Appendix F details the pipeline with representative anomaly examples; all prompt templates are in Appendix J.

4

Experiments

We investigate whether training on EDGE dataset improves multi-step function-calling over Korean

public APIs and whether the gains generalize beyond our training distribution. We evaluate on KOPA-B ENCH to measure Korean public API tool use, and on BFCL to measure out-of-distribution performance. Appendix H verifies that KOPAB ENCH is not contaminated by the training dataset, and additionally reports performance on platforms withheld from synthesis (Appendix H.1) and the overlap of dependency edges between the two corpora (Appendix H.2). 4.1

Setup

We fine-tune Qwen3.5-4B and Qwen3.5-9B with GRPO on the 1,781 training data of §3 (detailed in Appendix G), employing a binary reward r ∈ {0, 1} over the RESPONSE and ENVIRONMENT dimensions. Training with verl (Sheng et al., 2024) on 8×H100 GPUs, full details in Appendix I. We sample four independent rollouts per task and report both pass@1 and pass@4. To assess the statistical reliability of the pass@1 estimates, we additionally re-evaluate all models over eight independent random seeds and report 95% confidence intervals in Appendix C. 4.2

Main results

KOPA-B ENCH. As shown in Table 1, both finetuned models achieve substantial gains over their base counterparts. Our 4B model raises pass@1 from 0.18 to 0.31 (+13pp) and the 9B model from

Training objective

pass@1

pass@4

Action

Qwen3.5-4B (base)

0.1758

0.3103

0.2140

SFT (Ours) GRPO (Ours)

0.2724 0.3094

0.3586 0.4690

0.3048 0.3462

Table 2: Training objective ablation, evaluated on KOPA-B ENCH with four rollouts per task; metrics are defined as in Table 1. SFT and GRPO consume the identical 1,781 EDGE tasks and share the same base checkpoint.

0.33 to 0.43 (+10pp), approaching Qwen3.5-27B (0.45) at one-third of its parameters. The 95% confidence intervals estimated over eight independent seeds do not overlap between our models and their base counterparts, so the gains are statistically significant (Appendix C). They also extend beyond the platforms seen during synthesis: on the 31 tasks grounded in platforms withheld from synthesis entirely, the 4B model improves by +22.6pp in pass@4, exceeding its +15.9pp pass@4 gain over the full benchmark (Appendix H.1). Other benchmark. Training on the EDGE corpus does not come at the cost of out-of-distribution performance. On the BFCL benchmark, both finetuned models on aggregate improve over their base counterparts. Multi-turn improves the most (+4.04pp for 4B, +5.87pp for 9B), consistently exceeding the single-turn gains. Despite our training distribution being predominantly single-turn, the multi-turn generalization likely stems from the structural alignment between multi-step toolcall trajectories and multi-turn conversational structures. 4.3

Ablation Study

Effect of the training objective. The gains in Table 1 could in principle come from the GRPO objective rather than from the EDGE corpus itself. To isolate the objective’s contribution, we supervisefine-tune the same base checkpoint on the same 1,781 tasks. We generate trajectories with Qwen3.5397B-A17B under the environment used for GRPO rollouts (identical tools, system prompt, and runtime context) and train on the verified ones. As shown in Table 2, SFT alone already raises pass@1 by +9.7pp and Action by +9.0pp over the base model, so most of the improvement is attributable to the corpus rather than to the objective. GRPO then adds a further consistent gain across all three metrics, most visibly in pass@4

(0.3586 → 0.4690, +11.0pp). Since both methods consume the identical task set, this gap reflects how much signal each extracts per task rather than data quantity. SFT imitates a single verified reference trajectory and is upper-bounded by the teacher, whereas GRPO samples multiple rollouts per prompt and learns from their relative rewards. Where one query admits several valid trajectories and intermediate failures are common (Appendix B.4), this exploration extracts more signal than single-reference imitation. Effect of trajectory diversification. To assess whether the structural diversity introduced by our synthesis pipeline drives the observed gains, we split the corpus into purely sequential, purely parallel, and M IXED trajectories—the last combining both structures within a single trajectory and forming the largest group (36.7%, Table 12)— and train on three subsets: pure-sequential, parallel + M IXED, and full. Note that parallel + M IXED removes only purely sequential trajectories, since the retained M IXED ones still carry sequential subchains. As shown in Table 3, pure-sequential lags behind (0.2327 pass@1) because it never observes parallel composition, whereas parallel + M IXED (0.3080) is exposed to both and approaches full (0.3094). The two separate at pass@4, where full scores 0.4690 and parallel + M IXED 0.4000. full trails parallel + M IXED on ACTION (0.3462 vs. 0.4070), but the gap is one of task composition, not trajectory quality. ACTION is an efficiencyweighted score averaged only over the tasks a model solves (Appendix B.4), and the tasks that full alone solves are longer-horizon ones, on which even correct trajectories run long and score lower on efficiency. Since adding purely sequential data causes no pass@1 regression while improving pass@4, we adopt the full mixture for our final model. The two families are complementary: sequential trajectories teach precise output-to-input handoffs across junctions, while parallel ones infuse comparative and conditional patterns that a single sequential chain cannot express. Effect of data filtering. We investigate the effect of filtering the synthesized training data. LLM synthesis introduces anomalies such as malformed structures and stale labels from time-varying APIs (Appendix F). Removing these tasks raises pass@1 from 0.242 to 0.309 (+6.7pp), with consistent gains across all metrics (Figure 3).

Training dataset

pass@1

pass@4

Action

Qwen3.5-4B (base)

0.1758

0.3103

0.2140

pure-sequential parallel + M IXED full

0.2327 0.3080 0.3094

0.3724 0.4000 0.4690

0.3648 0.4070 0.3462

Table 3: Trajectory composition ablation, evaluated on KOPA-B ENCH with four rollouts per task; metrics are defined as in Table 1. M IXED trajectories are retained in both parallel + M IXED and full; pure-sequential contains only purely sequential trajectories. full uses the complete filtered dataset and corresponds to our final model, Qwen3.5-4B (Ours).

Effect of execution-grounded dynamic graph. Phase A first builds a skeleton graph from LLM feasibility scores, then prunes edges whose posteriors fall too low (§3.2). We ask whether this pruning truly separates executable from non-executable edges, rather than discarding them at random. Figure 4 compares the execution success rate of three edge sets, namely the LLM-judged skeleton, our converged G⋆ , and the pruned set removed during the loop. The skeleton serves as a controlled proxy for how prior synthesis pipelines establish links between calls (Yin et al., 2025; Liu et al., 2025; Prabhakar et al., 2025). It fixes each dependency from signature matching and a single LLM judgment, and it never checks that dependency against a live endpoint. It shares only this property with the prior pipelines by construction and is not an endto-end reimplementation of any of them, which lets the comparison isolate the effect of execution grounding while holding the candidate edge set fixed. The results show this separation clearly. Only 50.2% of the skeleton edges actually execute, whereas G⋆ reaches 62.7% (+12.5pp). Since both start from the same LLM-proposed edges, this gain comes entirely from execution, because the loop removes edges the LLM judged plausible but that fail in practice. The pruned edges, in turn, execute only 14.8% of the time, far below G⋆ . If pruning were random, the two sets would execute at similar rates. Instead, the loop removes exactly the edges that live APIs reject, which any pipeline that fixes dependencies without executing them cannot detect. Figure 4b confirms that the loop prunes dead edges rather than promising ones. Among the pruned edges, 70.5% never succeed in any trial, against 27.7% of those retained in G⋆ . The residual zero-success edges in G⋆ are concentrated among

Figure 3: Effect of trajectory-level filtering on the dataset. Filtering improves every metric.

Figure 4: Execution-grounded refinement improves the graph and prunes the right edges. Edges are grouped as the pre-refinement Skeleton, the surviving graph G⋆ , and the Pruned set; light and dark bars count all tested edges (te ≥ 1) and high-evidence edges (te ≥ 10); the te ≥ 10 bars are not comparable across groups, since failing edges are pruned before accumulating trials.

edges tried only a few times, so they reflect underexploration rather than demonstrated failure.

5

Conclusion

In this work, we address multi-step tool-calling over live Korean public APIs by introducing KOPA-B ENCH, which reveals performance limitations in existing LLMs. To bridge this gap, we present EDGE, an execution-grounded framework that synthesizes verified multi-step trajectories. Our findings demonstrate that training small open-source models with GRPO on the resulting dataset yields substantial gains, nearly matching much larger architectures and generalizing to BFCL. Thus, KOPA-B ENCH and EDGE offer a scalable foundation for deploying reliable LLM agents in public applications.

Limitations Reliance on live endpoints. Because the benchmark and the synthesized dataset are grounded in live APIs, both are subject to endpoint drift:

schemas, availability, and returned records can change over time. We mitigate this by filtering tasks tied to real-time data sources (Appendix F) and by evaluating environment state through hash comparison, but exact reproduction still depends on the stability of public services we do not control. Beyond Korean public APIs. Both KOPAB ENCH and the EDGE pipeline are built on Korean public-sector APIs across six domains. This isolates the multi-step, high-cardinality setting we target, but leaves open how far our findings transfer to other languages, to commercial or private APIs, and to administrative systems in other countries. We leave extending EDGE to these broader settings to future work. Scope of comparison. Our ablations compare EDGE against controlled variants of itself, including a schema-only skeleton that approximates settings in which tool dependencies are specified but not executed. This design isolates the contribution of execution grounding while holding the surrounding pipeline fixed. Running an existing synthesis system end to end on our tool inventory and training on its output remains outside the scope of the present study, so system-level comparisons with prior synthesis pipelines are left open.

Acknowledgements This work was conducted as part of the Sovereign AI Foundation Model Project (GPU Track), organized by the Ministry of Science and ICT (MSIT), South Korea (PJT-26-010016). We thank LG AI Research for valuable discussions and feedback on this work.

References Shipra Agrawal and Navin Goyal. 2013. Thompson sampling for contextual bandits with linear payoffs. In Proceedings of the 30th International Conference on Machine Learning. Aelim Ahn, Sooyeon Lee, Hyosun Wang, Chiwan Park, Daeryong Kim, Jihyeon Roh, Kichang Yang, Wonjun Jang, Hwang Woosung, Min Seok Kim, and Jihoon Kang. 2026. OrchestrationBench: LLM-driven agentic planning and tool use in multi-domain scenarios. In The Fourteenth International Conference on Learning Representations. Anthropic. 2026. Claude Sonnet 4.6 system card. Technical report, Anthropic.

Victor Barres, Honghua Dong, Soham Ray, Xujie Si, and Karthik Narasimhan. 2025. τ 2 -bench: Evaluating conversational agents in a dual-control environment. arXiv:2506.07982. Mingyang Chen, Haoze Sun, Tianpeng Li, Fan Yang, Hao Liang, KeerLu, Bin Cui, Wentao Zhang, Zenan Zhou, and Weipeng Chen. 2025. Facilitating multiturn function calling for LLMs via compositional instruction tuning. In The Thirteenth International Conference on Learning Representations. Eunbi Choi, Kibong Choi, Sehyun Chun, Seokhee Hong, Junwon Hwang, Hyojin Jeon, Ahra Jo, Hyunjik Jo, Yeonsik Jo, Joonkee Kim, Seonghwan Kim, Soyeon Kim, Sunkyoung Kim, Yireun Kim, Yongil Kim, Changhun Lee, Haeju Lee, Jinsik Lee, Kyungmin Lee, and 39 others. 2026. EXAONE 4.5 technical report. Gemma Team. 2026. Gemma 4 technical report. Technical report, Google DeepMind. Zhicheng Guo, Sijie Cheng, Hao Wang, Shihao Liang, Yujia Qin, Peng Li, Zhiyuan Liu, Maosong Sun, and Yang Liu. 2024. StableToolBench: Towards stable large-scale benchmarking on tool learning of large language models. In Findings of the Association for Computational Linguistics: ACL 2024. Wenqi Huang, Charley Lee, Leonard Tng, and Serena Ge. 2026. DeepSWE: Measuring frontier coding agents on original, long-horizon engineering tasks. arXiv:2607.07946. Shinbok Lee, Gaeun Seo, Daniel Lee, Byeongil Ko, Sunghee Jung, and Myeongcheol Shin. 2024. FunctionChat-Bench: Comprehensive evaluation of language models’ generative capabilities in Korean tool-use dialogs. arXiv:2411.14054. Minghao Li, Yingxiu Zhao, Bowen Yu, Feifan Song, Hangyu Li, Haiyang Yu, Zhoujun Li, Fei Huang, and Yongbin Li. 2023. API-Bank: A comprehensive benchmark for tool-augmented LLMs. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Alexander H. Liu, Kartik Khandelwal, Sandeep Subramanian, Victor Jouault, Abhinav Rastogi, Adrien Sadé, Alan Jeffares, Albert Jiang, Alexandre Cahill, Alexandre Gavaudan, Alexandre Sablayrolles, Amélie Héliou, Amos You, Andy Ehrenberg, Andy Lo, Anton Eliseev, Antonia Calvi, Avinash Sooriyarachchi, Baptiste Bout, and 101 others. 2026. Ministral 3. arXiv:2601.08584. Weiwen Liu, Xu Huang, Xingshan Zeng, xinlong hao, Shuai Yu, Dexun Li, Shuai Wang, Weinan Gan, Zhengying Liu, Yuanqing Yu, Zezhong WANG, Yuxian Wang, Wu Ning, Yutai Hou, Bin Wang, Chuhan Wu, Wang Xinzhi, Yong Liu, Yasheng Wang, and 8 others. 2025. ToolACE: Winning the points of LLM function calling. In The Thirteenth International Conference on Learning Representations.

Zuxin Liu, Thai Quoc Hoang, Jianguo Zhang, Ming Zhu, Tian Lan, Shirley Kokane, Juntao Tan, Weiran Yao, Zhiwei Liu, Yihao Feng, Rithesh R N, Liangwei Yang, Silvio Savarese, Juan Carlos Niebles, Huan Wang, Shelby Heinecke, and Caiming Xiong. 2024. APIGen: Automated pipeline for generating verifiable and diverse function-calling datasets. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track. Seiji Maekawa, Jackson Hassell, Pouya Pezeshkpour, Tom Mitchell, and Estevam Hruschka. 2026. Towards reliable benchmarking: A contamination free, controllable evaluation framework for multi-step LLM function calling. In The Fourteenth International Conference on Learning Representations. Mike A Merrill, Alexander Glenn Shaw, Nicholas Carlini, Boxuan Li, Harsh Raj, Ivan Bercovich, Lin Shi, Jeong Yeon Shin, Thomas Walshe, E. Kelly Buchanan, Junhong Shen, Guanghao Ye, Haowei Lin, Jason Poulos, Maoyu Wang, Marianna Nezhurina, Di Lu, Orfeas Menis Mastromichalakis, Zhiwei Xu, and 65 others. 2026. Terminal-Bench: Benchmarking agents on hard, realistic tasks in command line interfaces. In The Fourteenth International Conference on Learning Representations. Mahmoud Mohammadi, Yipeng Li, Jane Lo, and Wendy Yip. 2025. Evaluation and benchmarking of LLM agents: A survey. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2. NVIDIA. 2025. Nemotron-SFT-agentic-v2. https://huggingface.co/datasets/nvidia/ Nemotron-SFT-Agentic-v2. OpenAI. 2025. GPT-5.1 instant and GPT-5.1 thinking system card addendum. Technical report, OpenAI. Shishir G Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng-Jie Ji, Vishnu Suresh, Ion Stoica, and Joseph E. Gonzalez. 2025. The berkeley function calling leaderboard (BFCL): From tool use to agentic evaluation of large language models. In Forty-second International Conference on Machine Learning. Akshara Prabhakar, Zuxin Liu, Ming Zhu, Jianguo Zhang, Tulika Manoj Awalgaonkar, Shiyu Wang, Zhiwei Liu, Haolin Chen, Thai Quoc Hoang, Juan Carlos Niebles, Shelby Heinecke, Weiran Yao, Huan Wang, Silvio Savarese, and Caiming Xiong. 2025. APIGenMT: Agentic pipeline for multi-turn data generation via simulated agent-human interplay. In The Thirtyninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track. Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Lauren Hong, Runchu Tian, Ruobing Xie, Jie Zhou, Mark Gerstein, Dahai Li, Zhiyuan Liu, and Maosong Sun. 2024. ToolLLM: Facilitating large language models to master 16000+ real-world APIs. In The Twelfth International Conference on Learning Representations.

Qwen Team. 2026. Qwen3.5: Towards native multimodal agents. Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu. 2024. HybridFlow: A flexible and efficient RLHF framework. arXiv:2409.19256. Aaditya Singh, Adam Fry, Adam Perelman, Adam Tart, Adi Ganesh, Ahmed El-Kishky, Aidan McLaughlin, Aiden Low, AJ Ostrow, Akhila Ananthram, Akshay Nathan, Alan Luo, Alec Helyar, Aleksander Madry, Aleksandr Efremov, Aleksandra Spyra, Alex BakerWhitcomb, Alex Beutel, Alex Karpenko, and 467 others. 2026. OpenAI GPT-5 system card. Yihong Tang, Kehai Chen, Liang Yue, Jinxin Fan, Caishen Zhou, Xiaoguang Li, Yuyang Zhang, Mingming Zhao, Shixiong Kai, Kaiyang Guo, Xingshan Zeng, Wenjing Cun, Lifeng Shang, and Min Zhang. 2025. Empowering real-world: A survey on the technology, practice, and evaluation of LLM-driven industry agents. arXiv:2510.17491. Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik R Narasimhan. 2025. τ -bench: A benchmark for tool-agent-user interaction in real-world domains. In The Thirteenth International Conference on Learning Representations. Fan Yin, Zifeng Wang, I-Hung Hsu, Jun Yan, Ke Jiang, Yanfei Chen, Jindong Gu, Long Le, Kai-Wei Chang, Chen-Yu Lee, Hamid Palangi, and Tomas Pfister. 2025. Magnet: Multi-turn tool-use data synthesis and distillation via graph translation. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers).

A

Related Work

Function-calling benchmarks. Function-calling evaluation has been driven by foundational benchmarks such as API-Bank (Li et al., 2023) and BFCL. To broaden scope, ToolBench (Qin et al., 2024) scales to thousands of real-world REST APIs, StableToolBench (Guo et al., 2024) adds simulated environments to address the instability of live APIs, and τ 2 -Bench (Barres et al., 2025) scores agents against the resulting environment state through a user simulator. For Korean, FunctionChat-Bench (Lee et al., 2024) evaluates general-purpose tools and OrchestrationBench (Ahn et al., 2026) evaluates tool coordination that requires structured planning. None, however, targets public-sector APIs; we introduce a benchmark over live Korean public APIs where solving a query requires chaining interdependent calls over many record responses.

Function-calling data synthesis. Since real function-calling data are scarce, a common approach is to synthesize them with LLMs. To keep synthesized calls valid, ToolACE (Liu et al., 2025) adds a verification stage, and AWM (Maekawa et al., 2026) generates calls inside simulated API environments that return consistent responses. For multi-step tasks Magnet (Yin et al., 2025), BUTTON (Chen et al., 2025), and APIGen-MT (Prabhakar et al., 2025) build such call chains by linking functions through their signatures or LLMproposed plans. In all of them, links between calls are fixed without checking them against the live APIs, and each call is assumed to return a single result, leaving unaddressed the unstable endpoints and multi-record responses of real public APIs.

B

Evaluation Details

B.1

Task Example

We illustrate the benchmark with a representative task drawn from the finance domain (Figure 5). The task requires the model to compute how far a stock’s latest closing price sits above the refixing floor of a convertible bond, given the current date. The user query is “Referring to the convertible bonds issued by Lightron Fiber-optic Devices in January 2026, calculate by what percentage yesterday’s closing price differs from the refixing floor price.” Figure 5 shows the full task definition and the evaluation workflow. To solve this task, the model must resolve the company’s DART corporate code, retrieve its January 2026 convertible-bond issuance record to obtain the refixing floor price, look up the prior day’s closing price, and finally compute the percentage gap. The expected trajectory is: 1. find_dart_corp_code(company_name="Lightron") -> corp code "00367482" 2. get_dart_convertible_bond_issuance_decision( corp_code="00367482", start="20260101", end="20260130") -> refixing floor 532 3. get_stock_price_trend(ticker="069540", strtDd="20260129", endDd="20260129") -> closing price 1680 4. evaluate_expression("100 * (1680 - 532) / 532") -> 215.78...

The system state includes context injected into the agent’s system prompt at runtime. PredefinedSystemState entries supply static values (e.g., current_date: 2026-01-30), while ActionSystemState entries are resolved dynamically by executing the specified function at the start

Figure 5: A representative task from the Finance domain of KOPA-B ENCH. Each task pairs a natural-language query with system state injected at runtime. The model calls the provided tools to solve the query and output a final answer, which is evaluated against the annotated ground truth (golden actions, expected response, and resulting environment state).

of each episode (e.g., retrieving the API key via get_dart_api_key()). B.2

RESPONSE Evaluation

Response evaluation verifies whether a model’s final answer matches the ground truth. We support three answer types, each appending a formatspecifying suffix to the system prompt. String type. The model emits its answer as ANSWER:{answer}. We extract it via pattern matching, normalize it by removing whitespace and special characters, and compare it against all acceptable values. Unmatched cases fall back to LLM-asa-Judge. Number type. The model encloses its answer in \boxed{} as a sympy-parseable number or ex-

pression (e.g., \boxed{42}). We normalize the extracted content through comma removal and float conversion, evaluating non-numeric strings symbolically with sympy, then compare to the ground truth under approximate equality. Failed cases fall back to LLM-as-a-Judge. Judge type. For open-ended tasks such as those in the Law domain, the model answers in the same format, and we rely solely on LLM-as-a-Judge without pattern-based extraction. B.3

ENVIRONMENT Evaluation

We use golden to denote the annotated groundtruth value throughout the remainder of the paper. ENVIRONMENT evaluation verifies that the system state resulting from the model’s tool execution matches the golden state. Each task runs in a freshly launched MCP server instance, isolating its state from other tasks so that evaluation is unaffected by residual state. At the end of each episode, we retrieve the final state and compare it against the golden state annotation via SHA-256 hash comparison, following Yao et al. (2025). B.4

ACTION Evaluation

ACTION evaluation measures both the correctness and efficiency of the model’s tool-call trajectory. Motivation. Prior function-calling benchmarks such as BFCL evaluate models primarily through action correctness, verifying whether predicted tool-calls match the golden actions. Correctness alone is insufficient for two reasons. First, a golden action’s presence does not guarantee that only correct actions were taken. Given the golden action fillTank(amount=20), the trajectory [fillTank(amount=20), fillTank(amount=10)] contains it yet reaches an incorrect final state. Second, one outcome is often reachable through multiple valid sequences. [fillTank(amount=10), fillTank(amount=10)] produces the same state as fillTank(amount=20) but is penalized under strict matching. We therefore use ACTION not as a primary success criterion but as a quality measure gated by outcome verification (§2.4), decomposed into correctness and efficiency. Tool-call correctness. For each ground-truth action, we search the entire predicted trajectory for a matching call, without requiring that calls appear in the annotated order. A match requires that

Model

pass@1

CI width

Qwen3.5-27B gemma-4-26B-A4B-it Qwen3.5-9B EXAONE-4.5-33B gemma-4-E4B-it Qwen3.5-4B Ministral-3-8B-Instruct-2512

0.4681 ± 0.0146 0.3672 ± 0.0151 0.3526 ± 0.0249 0.2224 ± 0.0122 0.1845 ± 0.0110 0.1603 ± 0.0204 0.1418 ± 0.0133

0.0292 0.0301 0.0498 0.0244 0.0220 0.0407 0.0265

Qwen3.5-4B (Ours) Qwen3.5-9B (Ours)

0.2931 ± 0.0169 0.4138 ± 0.0098

0.0338 0.0195

Table 4: Pass@1 on KOPA-B ENCH, reported as mean ± 95% confidence-interval half-width, computed over eight independent random seeds using the Student’s tdistribution (df = 7); the last column gives the full CI width. These eight-seed estimates complement the fourrun averages reported in Table 1.

the function name is identical, all required parameters are present, no unexpected parameters are included, parameter types conform to the schema under BFCL-style AST type checking, and parameter values match the ground truth for the arguments listed in value_compare_args. The correctness score Corr is the fraction of golden actions successfully matched. Tool-call efficiency. Efficiency combines a duplicate-call ratio (CR) and minimal-path coverage (MPC): CR = 1 −

Ndup , Ntotal

MPC =

Nopt , Ntotal

Eff =

CR + MPC 2

(3) where Ndup is the number of redundant duplicate calls, Ntotal the total number of calls, and Nopt the annotated optimal call count. Each task specifies a max_allowed_calls budget; if Ntotal exceeds it, Eff is set to 0. The final ACTION score is the mean of the two, (Corr + Eff)/2.

C

Statistical Reliability of Main Results

Given the modest size of KOPA-B ENCH (145 tasks), point estimates alone could be misleading. To assess whether the improvements reported in Table 1 are statistically meaningful, we re-evaluate every model over n=8 independent random seeds and report 95% confidence intervals for pass@1. Note that Table 1 in the main text reports the original estimates averaged over four runs, whereas this appendix reports the revised eight-seed estimates; the two are consistent, and all conclusions in §4 hold under both. Methodology. Prior agentic benchmarks report confidence intervals over multiple runs via a normal

Domain

Platforms

Traffic

Korea Expressway Corporation, Seoul Open Data Plaza DART, KRX, Korea Deposit Insurance Corporation Educational Data Platform, HRD Korea Korean Law Info Center Open Assembly Gyeonggi Data Dream

Finance Education Law Politics District Admin. Total

share a step, the mean tool-call count (4.92) exceeds the mean step count (3.32). • Headroom — the efficiency headroom, defined as max_allowed_calls/Nopt , where Nopt is the number of golden actions. It quantifies how tight the efficiency constraint is for each task: a value of 1.0 reduces ACTION efficiency to zero after the first redundant call, whereas larger values provide more tolerance before the max_allowed_calls ceiling is reached. • Avail. tools — the mean number of tools exposed to the agent (available_tools) per task. A larger pool increases the tool selection difficulty. • Parallel % — the fraction of tasks containing at least one set of parallel calls, i.e., two or more golden actions sharing the same step index.

#Tasks 22 29 30 22 25 17 145

Table 5: Selected platforms grouped by domain.

(z) approximation, i.e., estimate ± 1.96 · SE, e.g., Terminal-Bench (Merrill et al., 2026) and DeepSWE (Huang et al., 2026). Since our number of seeds is small (n=8), we instead use the Student’s t-distribution, s x̄ ± t0.975, df=7 · √ , n

(4)

where x̄ and s are the mean and standard deviation of pass@1 across seeds, yielding more conservative intervals than the normal approximation. Results. As shown in Table 4, the intervals are narrow (widths of roughly 0.02–0.05 in pass@1), indicating that the results are highly stable across seeds despite the limited number of examples. Crucially, the confidence intervals of our fine-tuned models do not overlap with those of their base counterparts: [0.2762, 0.3100] vs. [0.1400, 0.1807] for the 4B pair, and [0.4040, 0.4235] vs. [0.3277, 0.3775] for the 9B pair. This confirms that the improvements reported in §4 are statistically robust rather than an artifact of the small benchmark size.

D

Benchmark Statistics

The benchmark spans 10 platforms across six domains, listed in Table 5. Per-domain statistics characterizing the structure and difficulty of the 145 evaluation tasks are reported in Table 6. All quantities are computed directly from the golden action annotations and task metadata. Metric definitions. • Steps — the number of sequential reasoning steps required to solve the task. Two golden actions sharing a step index are issued in parallel and count as a single step. • Tool calls — the total number of golden actions in the task. Because parallel actions

E

Execution-Grounded Dynamic Graph: Details and Analysis

This appendix expands on Phase A (§3.2). We describe the posterior update scheme (§E.1), the pruning procedure (§E.2), and analyses of graph convergence (§E.3,§E.4). All hyperparameters are listed in Table 9. E.1

Update Scheme

This appendix details the outcome taxonomy and the posterior-update rule summarized in §3.2, together with the stabilizing mechanisms. Each edge execution is classified into one of the outcomes in Table 7, grouped into structural faults, which implicate the binding, and environmental faults, which reflect infrastructure noise independent of the binding. Successes increment the edge’s Beta αe ; failures increment βe , with structural faults penalized far more heavily than environmental ones. The asymmetry between structural and environmental penalties reflects the public-API environment: rate limits, timeouts, and server errors strike edges at random, regardless of whether the binding is correct, so they should only weakly lower an edge’s posterior. E.2

Pruning Details

Pruning rule. An edge e with posterior θe ∼ Beta(αe , βe ) is pruned at iteration t if and only if it has accumulated at least nmin trials and its posterior mass above the viability threshold τ has

Domain

#Tasks

Steps

Tool calls

Headroom

Avail. tools

Parallel %

Traffic Finance Education Law Politics District Administration

22 29 30 22 25 17

3.64 3.69 2.93 3.00 3.72 2.76

5.14 5.65 4.30 3.55 5.48 5.41

2.21 2.24 2.79 3.63 2.10 1.55

7.5 6.6 7.7 6.7 8.7 8.0

55% 83% 57% 27% 52% 82%

Total / mean

145

3.32

4.92

2.45

7.5

59%

Table 6: Per-domain benchmark statistics. Means are weighted by the number of tasks per domain.

Outcome

Group

Semantic success Generic-only success

— —

Update αe += 1.0 αe += 0.3

Schema/type mismatch structural Missing-field error structural

βe += 1.5 βe += 1.0

Rate limit Timeout Server error Authorization error Endpoint unavailable

environmental environmental environmental environmental environmental

βe += 0.1 βe += 0.3 βe += 0.3 βe += 0.5 βe += 0.8

Uncategorized failure

—

βe += 0.5

Table 7: Outcome taxonomy and the corresponding Betaposterior update. Successes increment αe , more for a semantic success (1.0) than for a generic-only success (0.3); structural failures increment βe by 1.0–1.5, environmental failures by only 0.1–0.8. A success is semantic if at least one non-generic input parameter (i.e., not an API key, pagination, or format field) carried a value, and generic-only otherwise.

fallen below the confidence level ε: prune(e) ⇐⇒ ne ≥ nmin ∧ Pr[θe > τ ] < ε, (5) where Pr[θe > τ ] = 1 − Iτ (αe , βe ) is the complement of the regularized incomplete beta function, and ne is the edge’s trial count. Settings. The Beta prior of each admitted skeleton edge is set from the LLM-judged feasibility score se ∈ [0, 1] of its binding, as Beta(1 + κse , 1 + κ(1 − se )), so that κ controls how strongly the prior commits to the initial feasibility estimate. An edge enters the skeleton only if se ≥ smin . During each iteration the sampler draws paths by Thompson sampling, taking an εexp greedy random edge with probability εexp ; edges are then updated (Table 7) and pruned by Eq. (5). E.3

Graph Convergence Analysis

Execution success rises as the loop prunes unreliable edges. The comparisons above contrast edge sets at convergence; Figure 6 shows the same

Figure 6: Execution success rises as unreliable edges are pruned. Trajectory pass rate and step success rate of sampled paths over the 100 Phase A iterations; faint lines are per-iteration values and bold lines a moving average. As the loop prunes structurally failing edges and samples toward reliable ones, both rise steadily, indicating that refinement concentrates the graph on executable dependencies.

effect emerging during the run. As iterations proceed, structurally failing edges accumulate posterior mass below the pruning threshold and are removed, while Thompson sampling increasingly routes execution toward edges with demonstrated success. The trajectory pass rate and step success rate of sampled paths rise steadily as a result, by +31pp and +28pp over the 100 iterations, consistent with the cross-sectional gap above between the retained graph and the pruned set. The loop thus progressively concentrates the graph on executable dependencies rather than reshaping it arbitrarily. E.4 Prior Calibration: LLM Judgment versus Execution Evidence Prior graph-based methods commit to edge reliability from the LLM’s judgment alone. Our loop instead grounds it in execution, and this section shows that both the posteriors and the pruning decisions follow observed outcomes, pruning even highly rated edges when execution contradicts the LLM. The skeleton anchors each edge’s Beta prior to its LLM feasibility score se (§3.2), and the loop

F

Figure 7: Pruning is governed by execution, not by the LLM prior. Each point is an admitted edge, plotting its feasibility prior se against the converged posterior mean θ̄e . Retained and pruned edges separate along the posterior axis near τ = 0.5 rather than along the prior axis, so an edge can carry a high prior yet be pruned once execution contradicts it.

revises this into a posterior with mean θ̄e . The revision is worth its cost only if θ̄e departs from se , since a posterior that merely tracked the score would make the loop redundant. Figure 7 plots θ̄e against se for every admitted edge probed at least once, with the diagonal marking θ̄e = se and the pruning threshold τ . Execution evidence overrides the prior. Edges spread broadly along the posterior axis rather than concentrating on the diagonal, so θ̄e is driven by observed outcomes rather than se . The deliberately weak prior κ = 2.0 lets a few executions move an edge well away from its initial score, and the correction runs both ways, as high-prior edges (se ≥ 0.6) often fall below the diagonal while some low-prior edges (se ≤ 0.3) rise above it. Pruning follows the posterior, not the prior. Because the decision reads θ̄e rather than se , retained and pruned edges interleave across the full range of se . The boundary runs horizontally near τ , so the two sets separate along the posterior axis, the signature of an execution-grounded decision. The loop thus removes precisely the edges the LLM endorsed but the live APIs reject, which a schemaonly pipeline would retain. The cold-start guard (nmin = 2, Eq. 5) keeps this safe, as none of the 10,174 pruned edges had zero trials.

Data Filtering Details

To optimize efficiency, we filter the synthesized corpus with a three-stage cascading pipeline with increasing computational costs. Table 8 presents representative examples of execution flaws, textual anomalies, and logical leaks identified and pruned during synthesis. Stage 1 (Rule-based filters). Deterministic rules remove four types: non-reproducible instances (e.g., real-time queries without a fixed timestamp), erroneous data whose ground truth is an error signal, redundant near-duplicates, and malformed instances violating format or schema constraints. Stage 2 (LLM-based static filters). An LLM judge then removes ill-posed instances: nongradable queries whose correctness cannot be verified automatically, answer-leaking queries that disclose the target answer, and decorative chains whose intermediate calls do not constrain the final answer. The latter two require agreement between two independent LLM judges. Stage 3 (Independent re-solving). Neither stage detects errors in the ground truth itself. Whereas existing pipelines treat it as a fixed reference (Liu et al., 2024), we re-solve every task independently; a re-solution judged correct actively replaces the original stored answer. Throughout this pipeline, only open-source model outputs become training labels and proprietary models serve solely for verification.

G

Training Data Analysis

Training data statistics. Each task is typed by the Phase B taxonomy of §3.3, and Table 12 reports the resulting training data distribution. The key effect of the cardinality-based junction typing is visible among the single-junction sequential trajectories. The high-cardinality fields of Korean public APIs (§3.3) make one-to-many junctions the norm, so that FAN (2≤ni ≤ϕ) and D RV (ni >ϕ) together account for 64.9% of these trajectories (398 of 613), far outnumbering the oneto-one S EQ junctions (ni =1, 35.1%) to which a naive one-to-one template would be confined. The remaining tasks are template-based non-sequential types (S EM, C MP, C OND, 28.9% of the corpus) and Mixed tasks (36.7%) that combine more than one of these patterns within a single trajectory. Beyond type composition, the corpus is structurally demanding. Table 10 reports its call structure over the 1,781 tasks. A task issues 4.13 tool

Stage

Filter Class

Representative Anomaly Example (User Query / Trajectory Flaw)

Stage 1

Real-time source

Trajectory components invoke APIs that fetch live highway gating status without a fixed reference timestamp, causing the ground-truth targets in the training dataset to dynamically change per execution (violating data determinism). Example: “실시간 영업소 진입조절 현황에 뜬 각 노선별로 휴게시설 경로 정보가 몇 건씩 있는지 확인해 줄 수 있어?” (Gloss: “Can you check how many rest area route information items are available for each route listed in the real-time tollbooth entry control status?”)

Stage 1

Linguistic anomaly

Query contaminated with foreign tokens or mixed scripts during synthesis, violating standard format constraints. Example: “전월 대비人口 증감 수치를 알려줘." or "먼저 확인한 뒤,そこに 포함된 각 영업소별로”. (Gloss: “Tell me the change in 人口 compared to the previous month." / "After checking first, そこに for each tollbooth included there.”)

Stage 2

Answer leakage

The query phrasing explicitly discloses the target intermediate parameter, rendering the associated tool call redundant. Example: “2023년 김포시에 있는 야구 가능한 공공체육시설 주소에서 시군명을 뽑 아서, 그 시군별로 경기도 내 골프장이 몇 건씩 있는지 알려줄래?” (Gloss: “Extract the municipality name from the address of baseball-available public sports facilities in Gimpo City in 2023, and tell me how many golf courses there are in Gyeonggi-do for that municipality?”)

Table 8: Taxonomy and representative examples of filtered data anomalies, categorized by filtering stage, filter class, and associated query/trajectory flaws.

Symbol Description

Value

Phase A: skeleton graph construction K Intra-domain candidate 15 neighbors / source tool Kx Cross-domain candidate 10 neighbors / source tool 0.3 smin Min. feasibility score to admit an edge κ Prior strength 2.0 Skeleton-construction model Qwen3.5-122B Phase A: execution-grounded graph update T Graph-update iterations M Paths sampled per iteration εexp ε-greedy exploration rate τprune Viability threshold (Eq. 5) εprune Pruning confidence level (Eq. 5) nmin Min. trials before pruning

100 500 0.1 0.5 0.3 2

Phase B: trajectory synthesis ϕ Fan-out budget (§3.3)

5

Table 9: EDGE hyperparameters and the values used in all reported runs.

calls on average and up to 36, and 57.6% of tasks require at least four sequential hops, so a model must sustain long chains of dependent calls rather than answer in a single step. The structural depth, together with the one-to-many chaining detailed in Table 11, is what the model must learn to handle. Comparison with existing tool-use datasets. What distinguishes our corpus from prior tool-use datasets is the prevalence of one-to-many chaining,

Metric

Value

Calls per task (mean / med. / min / max) 4.13 / 4 / 2 / 36 Steps: ≤3 / 4–5 / ≥6 42.3 / 50.9 / 6.7 % Calls: sequential / parallel 63.8 / 36.2 %

Table 10: Training-data statistics over the 1,781 synthesized tasks: the distribution of calls per task, the share of tasks by step count, and the split of calls into sequential and parallel.

a direct consequence of the high-cardinality fields of Korean public APIs. Table 11 compares our corpus against representative tool-use datasets: 74.1% of calls in our corpus return at least two records and 81.2% of chained calls are one-to-many, with a median chained cardinality of 27 (up to 224,958), whereas the prior datasets stay below 5% one-tomany chaining with single-digit cardinality. This is precisely the regime that the cardinality-based junction typing of §3.3 is designed to handle, and that a one-to-one synthesis template cannot reach.

H

Data Contamination Audit

Because KOPA-B ENCH and the EDGE training corpus are both built on the same live public APIs, they necessarily draw on a shared pool of tools. We therefore audit how much the 145 evaluation tasks overlap with the 1,781 training tasks, separating genuine leakage from the shared pool that

Dataset

Environment

ToolBench-v1 APIGen-MT Nemotron (NVIDIA, 2025) ToolACE Ours

one-to-many chaining (%)

RapidAPI τ -bench τ 2 -bench synthetic KOPA-B ENCH

1.9 10.0 1.6 4.5 81.2

chained cardinality median

p90

max

1 1 1 1 27

16 3 1 5 157

127 16 15 10 224,958

Table 11: Comparison of chaining-related statistics across datasets. one-to-many chaining is the share of multi-step trajectories that contain at least one junction consuming a multi-record tool output. chained cardinality reports the distribution of records consumed at such junctions.

Type

Count

%

M IXED D RV S EM S EQ C MP C OND FAN

654 350 303 215 161 50 48

36.7 19.7 17.0 12.1 9.0 2.8 2.7

Total

1781

100.0

Table 12: Distribution of task types in training set Level Verbatim query leakage Tool-universe overlap (functions) Class-level held-out platforms Tasks using only unseen functions Mean per-task function coverage

Value 0.0% (0/145) 62.2% (135/217) 3/12 12.4% (18/145) 62.9%

Table 13: Train–eval contamination audit between KOPA-B ENCH (145 tasks) and the EDGE training corpus (1,781 tasks).

a function-calling benchmark and its training corpus are expected to share. We define tool-universe overlap as the fraction of distinct evaluation gold functions that also appear as a gold function anywhere in training. This measures the shared API pool, which is intended by design and is distinct from query or task leakage. We report the audit at five levels, summarized in Table 13. No evaluation query appears verbatim in training (0 of 145), so there is no direct query leakage. The tooluniverse overlap is 62.2% (135 of 217 gold functions), which reflects the shared API pool rather than contamination. Verbatim query leakage at 0% is the strongest available evidence that the benchmark is not contaminated by the training corpus. H.1

Held-Out Platform Evaluation

Tool-universe overlap alone does not show whether the reported gains depend on the shared pool. We

evaluate on the subset of KOPA-B ENCH grounded in platforms that never enter synthesis: Seoul Open Data Plaza in the Traffic domain, and DART and KRX in the Finance domain. These three platforms of Table 5 appear only in KOPA-B ENCH, and none of their API tools appears in any training trajectory, so the held-out unit is the platform rather than the individual function. In total, 31 benchmark tasks are grounded in them, 10 on Seoul Open Data Plaza and 21 on the two finance platforms. The only tool these tasks share with training is evaluate_expression, the general-purpose evaluator of Appendix B.1, which takes an arithmetic or symbolic expression as a string and returns its value. Many tasks close with an arithmetic step and so call it, but it queries no external service and carries no platform-specific API knowledge. Once it is set aside as a general-purpose utility, every public-API tool involved in the 31 tasks is absent from training. Table 14 reports pass@4 on this subset, as defined in Table 1. Fine-tuning improves performance on the held-out platforms rather than degrading it, from 0.2903 to 0.5161 (9 to 16 of the 31 tasks), a gain of +22.6pp that exceeds the +15.9pp the same model gains over the full benchmark (0.3103 → 0.4690). The improvement is therefore not confined to the platforms seen during synthesis, as it would be if the model had acquired tool-chain templates or platform-specific call procedures from training. H.2

Dependency Edge Overlap

The audits above compare tool inventories and whole tasks. A model could still benefit from having seen the same dependency structure during training even when the individual tools differ, so we audit overlap at the level of edges between calls, at two resolutions. A function-pair edge (u, v) records that an output field of tool u supplies an

Domain Held-out platform Traffic Finance Total

n

Base

Ours

Seoul Open Data Plaza 10 0.2000 0.6000 DART•KRX 21 0.3333 0.4762

Resolution

Eval Train Jaccard Unseen

Function-pair Adjacency

178 592

687 2,911

0.000 0.004

100% 95.6%

31 0.2903 0.5161

Table 14: Performance on the 31 KOPA-B ENCH tasks grounded in platforms held out from synthesis, for Qwen3.5-4B before and after training on the EDGE corpus. n is the number of tasks per platform; values are pass@4 as defined in Table 1.

input parameter of tool v, which is the dependency relation that EDGE synthesizes and that the model must recover at evaluation time. An adjacency edge links two consecutive calls in a trajectory regardless of whether a dependency is present, serving as a loose upper bound on structural overlap. We count every edge occurrence across all 145 evaluation tasks without deduplication, so a structure repeated across tasks contributes proportionally to how often the model encounters it. The Jaccard similarity is computed over sets of unique edges, |Eeval ∩ Etrain |/|Eeval ∪ Etrain |, and Unseen denotes the fraction of edge occurrences in evaluation whose edge never appears in training. Edges that merely acquire an API key are excluded: in the open-API setting of KOPA-B ENCH, key acquisition precedes every call and is thus a fixed property of the environment rather than a task-specific structure. At the dependency resolution, this exclusion is automatic, since an API key is a generic parameter filled through the D EFAULT branch of Eq. (1) rather than U PSTREAM (§3.2), and therefore never forms a dependency edge. Table 15 shows that the dependency edges do not overlap at all. The Jaccard similarity is 0.0, and all 178 evaluation dependency-edge occurrences are unseen in training, despite the 62.2% tool-universe overlap. That is, the benchmark and the training corpus draw on a shared pool of tools but connect them in different ways. Even under the loose adjacency bound, 95.6% of occurrences remain unseen, so this conclusion does not depend on how strictly an edge is defined. Together with §H.1, the audits point in the same direction. The shared tool pool is intended by design, yet the platforms behind 31 evaluation tasks never enter synthesis, and the dependency structure the model must recover at evaluation time is entirely new.

Table 15: Overlap of edges between calls in KOPAB ENCH and the EDGE training corpus. Eval and Train are edge occurrence counts without deduplication, while Jaccard is computed over the corresponding sets of unique edges. Function-pair is the dependency relation; adjacency links any two consecutive calls and serves as an upper bound. API key acquisition edges are excluded.

I

Training Details

We fine-tune Qwen3.5-4B (dense, 28-layer, bfloat16) with Group Relative Policy Optimization (GRPO) on the multi-step tool-calling corpus of §3, using verl with a Megatron-LM policy backend and a vLLM rollout backend on a single 8×NVIDIA H100 80 GB node. The GRPO objective uses a train batch of 8, aggregates the loss by token-mean, and applies a k3 KL penalty of 0.03 to the loss with no KL term in the reward. We use a context length of 65K tokens. For the multi-turn agent, we cap each episode at 15 assistant turns and 16K tool-response tokens, with episode and tool timeouts of 600 s and 60 s.

J

Prompts

This section provides the prompt templates used by the EDGE synthesis pipeline (§3): dependency extraction and argument generation in Phase A (§J.1), query generation in Phase B (§J.2), and the LLMbased static filter (§J.3). J.1

Phase A: Graph Construction

Phase A invokes two prompts. Dependency Extraction (Table 16) extracts dependencies when building the skeleton graph, returning the feasibility score se , the binding set Be , and the default candidates De of Eq. (1) in a single structured response; feasibility scoring is thus not a separate call but the feasibility_score field of this prompt. Source-Argument Generation (Table 17) generates valid argument values for the source tool when an edge is probed against the live API, with API key and pagination parameters excluded and injected separately. Both prompts are run with Qwen3.5122B.

J.2

Phase B: Trajectory and Query Synthesis

Phase B generates a Korean query for each completed trajectory whose answer requires the full tool-call chain. All Korean prompts are shown in English translation. All calls share a system prompt (Table 18) that enforces a conversational tone, prohibits procedural and API references, restricts value citations to seed call arguments or process-step filter values, and mandates a five-item self-check before output. The user prompt varies by trajectory type (§3.3). Sequential trajectories (S EQ, FAN, D RV and compositions) use one template (Table 19) with a pattern-specific assembly rule. Nonsequential trajectories (S EM, C MP, C OND) use two stages, a per-step sub-query prompt (Table 20) and a composition prompt (Table 21) that keeps every step logically necessary. A separate answer prompt (Table 22) handles these, with per-target enumeration for S EM and C MP and two-sentence branch selection for C OND. Phase B synthesis uses both Qwen3.5-122B and Qwen3.5-397B-A17B. J.3

Data Filtering

The static filtering pipeline deploys dedicated LLM classifiers to detect and eliminate three categories of ill-posed data: (a) non-gradable queries, (b) answer-leaking queries, and (c) decorative chains. We present the complete system prompt utilized for detecting answer-leaking queries in Table 23 as the representative example of our filtering rubrics. The validation stage uses Qwen3.5-397BA17B and GPT-5 (Singh et al., 2026) as the referee. J.4

Evaluation Prompts

This section provides the prompts used during evaluation (§4). The agent under test receives the system prompt of Table 24. When deterministic answer extraction fails, RESPONSE grading falls back to an LLM judge, using a shared prompt for the string and judge response types (Table 25) and a separate prompt with numeric-equivalence rules for the number type (Table 26).

Prompt for Dependency Extraction You are given two API functions: a **Source** function and a **Target** function. Your task is to: 1. Evaluate the structural feasibility of using Source's output as input to Target. 2. Extract explicit **binding sets**: which output fields of Source map to which input parameters of Target. Source API Function: {source_api_str} Target API Function: {target_api_str} **Instructions**: - For each Target input parameter (EXCLUDING common request params like KEY, TYPE, START_INDEX, END_INDEX, dataFormat, numOfRows, pageNo, pIndex, pSize, service, resultType, OC, api_key, service_key, serviceKey), determine if any Source output field can provide its value. - If a transformation is needed (e.g., date format change, type casting), describe it briefly in the "transform" field. - Assign a confidence score (0.0-1.0) to each binding. - Also provide an overall feasibility_score (0.0-1.0) using these guidelines: - 0.9-1.0: Direct, clear output-to-input dependency - 0.7-0.8: Strong dependency, minor transformation needed - 0.4-0.6: Moderate dependency, uncertain match - 0.2-0.3: Weak, tangential relationship - 0.0-0.1: No meaningful dependency **CROSS-DOMAIN BINDING - CRITICAL**: Field names often differ across domains even when they represent the SAME concept. You MUST look at the **description/meaning** of fields, NOT just their names. If two fields represent the same real-world concept, they SHOULD be bound even if names differ. Common cross-domain mappings to look for: - Region/location: SIGUN_NM, REGION_NM, rgn, rgnSeNm, ctpvNm, [region name] -> all mean "region name" - Region code: SIGUN_CD, REGION_CD, rgnCd -> all mean "region code" - Year: SUM_YY, fyr, crtrYr, baseyy, YEAR, srchyear -> all mean "year" - Name/title: BILL_NAME, CNTNTS_TITLE, schlNm, opnId -> entity identifiers - Date: USE_YMD, BGN_DE, END_DE -> date values Example cross-domain bindings: education.output["ctpvNm"] -> hrdk.input["rgn"] (both = region name, confidence: 0.7) national_assembly.output["BILL_NO"] -> law.input["MST"] (both = bill identifier, confidence: 0.6) dataseoul.output["USE_YMD"] -> expressway.input["exDate"] (both = date, confidence: 0.7) Do NOT reject a binding just because field names look different. DO reject a binding only if the semantic meaning is clearly incompatible. **Directional Constraint**: Evaluate ONLY Source -> Target direction. **Suggested Defaults for Unbound Parameters**: For Target input parameters that are NOT bound to any Source output field (excluding KEY/TYPE/pagination), suggest 3 realistic default value candidates that would make the API call succeed. This is critical for parameters that would otherwise require user input (e.g., region names, dates, IDs). Use your knowledge of the API's domain to suggest DIVERSE values that are likely to return non-empty results. For example, if the parameter is a region name, suggest 3 different regions. Return as JSON: { "feasibility_score": 0.75, "reason": "Brief explanation of the feasibility assessment", "bindings": [ { "source_field": "field_from_source_output", "target_param": "param_in_target_input", "transform": null, "confidence": 0.9 } ], "suggested_defaults": { "unbound_param_1": ["candidate_1", "candidate_2", "candidate_3"], "unbound_param_2": ["candidate_1", "candidate_2", "candidate_3"] } } If no meaningful bindings exist, return an empty bindings list and a low feasibility_score.

Table 16: Prompt for dependency extraction in Phase A. A single call returns the feasibility score, binding set, and default candidates of Eq. (1).

Prompt for Source-argument generation You are generating valid arguments for a Korean public API tool call. Tool Name: {tool_name} Tool Description: {tool_description} Input Parameters Schema: {input_schema} Rules: - Generate realistic, valid argument values for each parameter. - DO NOT generate values for API key parameters (KEY, api_key, service_key, serviceKey) - these will be injected automatically. - DO NOT generate values for pagination/format parameters (numOfRows, pageNo, dataFormat, type, start_index, end_index) these are handled separately. - For year parameters, use "2023" unless the description suggests otherwise. - For region/location parameters, use a common Korean region (e.g., Seoul, Busan, Gyeonggi). - For school-related parameters, leave as empty string "" to get all results. - For optional parameters with no clear default, use empty string "". - Return ONLY the parameters that should be included in the API call. Return as JSON: { "arguments": {"param1": "value1", "param2": "value2"} }

Table 17: Prompt for source-argument generation in Phase A, used when probing an edge against the live API.

Query-Generation System Prompt You are a query generator that writes Korean questions and precise answers in the **tone of a real user asking a chatbot/LLM**. ## User perspective The asker: - Has domain knowledge (e.g., expressways, rest areas, routes, committees, bills) - Does NOT know about APIs, functions, data structures, or field names - Does NOT care how the data is retrieved -- only what they want to know - Asks in a **natural conversational tone typical of everyday chatbot/LLM usage**, not in a report or paper style ## Tone (conversational) -- applied to the query only The question must take the form **a real user would send to a chatbot/LLM**. Formal or report-like tone must be avoided. **Vary sentence-ending forms** -- do not repeat one form: - "...How many items?" / "...How many cases?" / "...What is it like?" - "Please tell me..." / "Tell me..." / "Could you tell me?" - "I'm curious about..." / "I'd like to know..." / "I wonder..." **Natural segmentation is encouraged**: do not cram all conditions into one sentence; if it can be split into short sentences, splitting is OK. **Conversational markers (optional, do not overuse)**: "a little", "just", "by any chance", "by the way", etc. (Four natural-vs-unnatural example pairs omitted -- contrast between report tone and chatbot tone.) **Hard rules**: - No English API field names - No procedural/mechanism vocabulary ("retrieve", "API", "endpoint", etc.) - No exposure of meta-processing (tie-breaking, sorting steps, etc.) ## Handling tool descriptions Input tool descriptions may contain procedural/mechanism vocabulary. This is implementation detail and must NEVER appear in the question. - Procedural vocabulary: "retrieve", "in order to retrieve", "search", "verify", "call", "fetch", "obtained from", "shown in", "extracted from" - Mechanism vocabulary: "OpenAPI", "API", "function", "endpoint", "response" - Data-processing vocabulary: "the retrieved result", "the returned data", "the relevant list", "result code", "response data" - English API field names (code-style): SIGUN_NM, BLL_NM, UNIT_CD, CONF_ID, etc. -- convert to the Korean description shown in parentheses - Vague placeholders: "a specific X", "the relevant X", "some X" -- if a concrete value is in the args, it MUST be used ## Handling parenthetical clarifications in tool descriptions (A) Short formal title / source label -- keep - "Fund Annual Current Balance (Settlement) (Fiscal Scale)" -- standard Korean statistical naming convention - Typically <= 10 characters, no commas (B) Long supplementary clarification -- strip, keep only the core entity - "... status information (engineers, technicians, etc., applicant qualifications, application procedure, ...)" -> use only "... status information" in the query - Typically > 10 characters or contains commas Rule: **If the parenthetical contains even one comma, it is almost always type (B) supplementary -- strip the whole parenthetical**. ## Avoid repeating organizational prefix in tool names When two tools in the chain share the same source/organization prefix, use the prefix **only once** in the query. ## No exposure of meta-processing or search mechanism The query is **what the user asks**. System processing details (tie-breaking, sorting, selection steps) must **NEVER appear** in the query. ## Query value citation rule (mandatory) **Concrete values** appearing in the query (years, codes, region names, thresholds, quarters, etc.) must be one of: 1. A value in the seed step's call args 2. The filter_value of a process step (DRV threshold) 3. All other values -- forbidden in the query (metadata field values from result_rows, guesses based on world knowledge, etc.) Information that does not qualify must be expressed in generalized form: "for each X", "by region", "matching a certain condition", etc. ## Handling long memo / text values Long text in seed args (> 30 characters; long combined memo/description/address) must NOT be copied verbatim into the query. Either generalize or extract a short essential portion. ## Naturalize answer fields (answer_fields) When answer fields are raw English / special expressions, use generalized or unified expressions in the query. ## Self-check before emitting output (mandatory) After drafting the query, verify the following five items before output: 1. Zero implementation-detail traces: no procedural vocabulary, no API field names, no mechanism vocabulary, no dataprocessing vocabulary 2. Legal value citation: every concrete value (year / code / region /threshold) exists in the trace's seed args or in a process step's filter_value 3. Parenthetical cleanup: comma-separated list-style parentheticals from tool descriptions were not copied verbatim 4. No prefix repetition: the same source/organization prefix does not appear twice in one query 5. No meta exposure: system rules such as tie-breaking, sorting, or processing steps do not appear in the query If any of the above is violated, **revise before output**. ## Output format JSON: {"query": "natural Korean question from the user's perspective",

"answer": "concrete answer grounded in execution results", "answer_fields": ["field1", "field2"]} - answer_fields: list of field names in the result records from which the answer values were taken (empty array when only a count is answered).

Table 18: System prompt for the Phase B query generator, which writes a Korean question in the tone of a real user without exposing any implementation detail.

Chain Query–Answer Assembly Based on the following chain components and execution results, write a query+answer pair. ## Chain composition ### 1. Data source (do NOT cite source vocabulary -- only entities and fields) - {seed_desc} {seed_args_section} ### 2. Step connection {bridge_section} ### 3. Downstream conditions {downstream_section} ### 4. Question type {answer_section} ## Execution results {execution_results} ## Assembly rule {pattern_assembly_rule} ## Chain-specific additional rules (In addition to the general rules in the system prompt, extra constraints specific to this chain's composition.) - **Actual value of the connected field**: in the query express only meaning (e.g., "for each city name"); in the answer include the actual value (e.g., "Yangju-si: 12 cases, Seongnam-si: 5 cases") - **Bridge value is a chain shortcut**: if the user knew that value they could skip the source step -> NEVER expose it in the query ## Input value naturalization (apply whenever an argument value is mentioned in the query) - Term/session codes -> natural language: "21" -> "21st", "414" -> "414th session" - Date codes -> natural language: "2023-04" -> "April 2023", "20230401" -> "April 1, 2023" - Year -> natural language: "2023" -> "year 2023" - Quarter -> natural language: "1" -> "Q1", "4" -> "Q4" ## Output (JSON) { "query": "natural Korean question", "answer": "concrete answer grounded in execution results", "answer_fields": ["field1", "field2"] } Answer rules: - Count question -> exactly reflect the count from execution results (e.g., "Total of 18 cases") - Field-value question -> extract the actual value from execution results - No results -> "There is no data matching this condition" - No placeholders such as "can be verified", "a list is provided", etc. - **When there are two or more targets, separate each target on its own bullet/numbered line**: example: "- Yangju-si: 1504 cases" / "- Seongnam-si: 1504 cases"

Table 19: Prompt that assembles a full query–answer pair from the chain components and execution results in Phase B.

Sub-Query Generation You generate a natural Korean sub-query for a SINGLE tool call step. ## Input - Tool: {tool_name} - Description: {tool_description} - Arguments (MUST be included in the sub-query in natural Korean): {arguments_str} - Result Summary: {result_summary_str} {step_context} ## Steps (follow in order) ### Step 1: Identify terminology Use the tool description as the domain term directly (it is already short). e.g., "expressway real-time traffic information" -> domain term = "expressway real-time traffic" ### Step 2: Convert arguments to Korean {arg_handling_rule} ### Step 3: Determine question type {question_type_rule} ### Step 4: Write the sub-query Combine the domain term + converted arguments + question type into one natural Korean question. ## FORBIDDEN (any of these -> invalid output) - API field names as-is: ROUTE_CD, STN_NM, ORG_CD, YEAR, era_co, etc. - Procedural language: "retrieve", "search", "call", "verified" - Vague terms: "recent", "some", "main", "several", "various" - Vague catch-all: "detailed info", "personal info etc." -> pick at most 4 specific fields - Result count numbers: "49 lawmakers", "191 minutes" -> the user does not know the count - Generalizing domain terms: "expressway traffic" -> "traffic info" (X), "traffic" (X) - Placeholder substitution: "specific city", "the relevant region", "specific route", "specific organization", "specific year" -- if a concrete value is in the arguments, it MUST be used. If the code cannot be naturalized into Korean, keep the code value (e.g., "city code 41150") ## Output JSON: { "sub_query": "natural Korean sub-question", "core_intent": "core intent (2-5 words)" }

Table 20: Prompt for generating a natural-language sub-query for a single tool-call step in Phase B.

Sub-Query Composition You compose multiple sub-queries into ONE natural Korean question. ## Input - Sub-queries: {sub_queries_str} - Junction Info: {junction_info_str} ## Steps (follow in order) ### Step 1: Apply pattern-specific composition rule {pattern_rule} ### Step 2: Determine answer type {answer_type_rule} ### Step 3: Write the composed query Merge sub-queries into one question following the pattern rule and answer type. Preserve exact domain terminology from subqueries (e.g., "expressway traffic" stays as-is, do NOT generalize to "traffic info"). ## FORBIDDEN (any of these -> invalid output) - Procedural language: "retrieved", "searched", "verified", "called", "fetched", "first", "then" - Data format / schema mentions: "JSON", "json", "format", "schema", "field" -- the user does not need to know the backend data format - Vague terms: "recent", "some", "main", "latest" - API field names: PRDC_YM_NM, ROUTE_CD, ORG_CD, STN_NM, etc. - Result count numbers: "out of 49 lawmakers", "191 minutes" -- user does not know counts - Vague catch-all: "detailed info", "personal info etc." -> max 4 specific fields - Enumeration: "list all", "enumerate everything" -> ask count instead ("total how many?") - Generalizing domain terms: "real-time traffic" -> "traffic info" (X) - Intermediate result leakage (connected fields only): if step N produces a value that becomes step N+1's input argument ( connected field), do NOT use that value in the query -- it enables shortcutting. Result fields that are NOT passed to the next step may appear in the query. - Placeholder substitution: "specific city", "the relevant region", "specific route", "specific organization", "specific year" -- if the sub-query has a concrete value it MUST be reflected. NEVER use "specific X" - **Invented names / synthesized identifiers**: if the sub-query contains code values (NAAS_CD=14M56632, CURR_COMMITTEE_ID =9700407, etc.), keep them **as-is**. Do NOT invent person names (e.g., Hong Gildong, Kim Cheolsu), committee names (e.g ., National Assembly Library, etc.). If you do not know which name corresponds to the code, write it as "code + description" e.g., "the lawmaker with code 14M56632"). ## INPUT VALUE NATURALIZATION (apply whenever an input value is used in the query) - Term/session codes -> natural language: "21" -> "21st", "414" -> "414th session" - Date codes -> natural language: "2023-04" -> "April 2023", "20230401" -> "April 1, 2023" - Year -> natural language: "2023" -> "year 2023" - Quarter -> natural language: "1" -> "Q1", "4" -> "Q4" - Code values may stay as-is, but pair with a Korean description when meaningful: "committee code AO" (OK), "city code 41390" (OK) ## MINIMALITY RULE (every step must be necessary) Each step must feed data to the next step -- the composed query must make EVERY step logically necessary. - The query must describe WHY the first step is needed (what it provides to the next step). - If removing step 1 would still allow answering the query -> the query is BAD. - BAD: "Based on the term of the 21st National Budget Committee's minutes, how many legislative activities are there in total ?" (term=21 is already in the query -> step 1 is skippable) - GOOD: "For the term in which the Budget Committee held the most meetings, how many legislative activities are there in total ?" (step 1 is needed to find which term) ## Output JSON: { "query": "natural Korean composed question", "step_justifications": [ {"step": 1, "reason": "why this step is needed"}, {"step": 2, "reason": "why this step is needed"} ] }

Table 21: Prompt that composes multiple sub-queries into a single natural Korean question in Phase B, enforcing that every step is logically necessary.

Answer Generation You generate a concrete answer for the given query based on execution results. ## Query {query} ## Execution Results (full step trace, for context) {execution_results_str} {answer_material_block} ## Rules - Answer MUST contain concrete values (numbers, names, dates) from the actual results. - **COUNT questions**: the number in the answer MUST be **taken verbatim from the `count` value of the terminal_calls material **. - Do NOT count records/rows yourself. Even when rows is a sample, the real total is `count`. - e.g., if count=5, the answer is "5 cases" (NOT 3 cases or 0 cases). - **Distinguishing parallel calls (CMP / SEM / FAN)**: when there are multiple entries in `terminal_calls`, **match each call' s input_args with its count/rows by the input values** in the answer. - e.g., [{"input_args": {"SIGUN_CD": "41280"}, "count": 5}, {"input_args": {"SIGUN_CD": "41190"}, "count": 5}] -> answer "City code 41280: 5 cases, 41190: 5 cases". - None of them must be omitted. - **VALUE questions**: use only the actual values present in rows. Do NOT invent names/values that are not in rows. - **When there are two or more targets, separate each target on its own bullet/numbered line** (for easier verification): Example (CMP / SEM parallel calls): - City code 41280: 5 cases - City code 41190: 5 cases Example (multiple organizations): - National Assembly Library: 422 cases - National Assembly Budget Office: 132 cases Each bullet must lead with the target identifier. A plain sentence is allowed when there is only one target. - **COND-pattern-only structure (queries with conditional branching)**: 1. **You MUST check the `MATCHED BRANCH` marker in the execution results**. seed count=0 means the "none (else)" branch. 2. **Write as two separate sentences**: the first states only the condition outcome (ending with a period); the second states only the matched branch's result. - seed count > 0: "[condition] exists. [matched branch result]" - seed count == 0: "[condition] does not exist. [matched branch result]" 3. Use only the result values of the matched branch in the answer. Data for the unmatched branch is not in the trace, so do NOT include it in the answer. - Example (seed=0, else): "There is no lawmaker with code HE428991. The number of National Assembly committee status items matching the committee name '2002 World Cup International Sports Event Support Special' is 3 in total." - Example (seed>0, then): "There are game producers in Gyeonggi-do. There are 9 facilities available for soccer." - FORBIDDEN: joining the two clauses with causal connectives such as "because none exists" or "because it exists" - FORBIDDEN: listing values directly without the branch-selection sentence - FORBIDDEN: an answer that starts with "exists" when seed count=0 - FORBIDDEN placeholders: "a list is provided", "can be verified", "verifiable from the data" - If result is empty (count=0): "There is no data matching this condition" - Answer should directly and specifically answer the query. - Keep the answer concise but complete. ## Output JSON: { "answer_reasoning": "step-by-step explanation of how the answer was derived (which tool's which result was used; how counts/values were combined)", "answer": "concrete answer", "answer_fields": ["field1", "field2"] } - answer_reasoning: the reasoning trace that derived the answer. Describe which execution step's results were used as evidence. - answer_fields: list of **field names** in the result records from which the answer's actual values were taken. e.g., if you answered with a stock name, ["jmNm"]; with name and grade, ["jmNm", "grdNm"]. When only a count is answered, [] (empty array).

Table 22: Prompt that generates the final answer for a query from the execution-result trace in Phase B.

Answer Leakage Detection You are a classifier evaluating the data quality of queries used for testing Korean tool-using agents. Target Class (answer_leaked_conditional): Cases where a query contains a conditional/filter clause meant to determine the target object, but the value of that target object is already explicitly stated elsewhere in the query, rendering the conditional check vacuous (logically meaningless). Core Rubrics: - Does the query contain a conditional clause such as "Z of [Object] that has both X and Y" or "Z of [Object] satisfying X and Y"? - [Important-A] Is the actual value of the [Object] (e.g., municipality name, department name, route name) already specified elsewhere in the query text? Values embedded within actions (tool calls) do NOT constitute a leak---this can be a normal sequential operation where the agent discovers and uses the value from a prior step's result. - [Important-B] Does the entity type of the suspected leaking value match that of the [Object] to be determined by the conditional clause? If the conditional clause determines a municipality but the query explicitly mentions a different attribute (e.g., phone number, address), it is not leakage. The leaked value must match the target entity type of the conditional clause for the check to be vacuous. - [Important-C] If the intermediate clause's result is explicitly stated in RESPONSE_VALUE, that clause is a required answer component for evaluation---hence, not leakage (likely a parallel call or essential info). - Is the conditional check merely formal because the [Object] is already determined (solely by looking at the query) without going through the conditions? True Cases (remove=true): - "In Suwon City, what is the Z of the municipality that has both X and Y?" -> The "municipality with both X and Y" should be determined via filtering, but "Suwon City" is already explicitly mentioned in the query. - "Which facility in Icheon City has both a pottery workshop and a Confucian school with an area >= 4000?" -> The target of the conditional check is already explicitly stated as "Icheon City", making the condition vacuous. False Cases (remove=false): - Simple single lookup ("What is the population of Suwon City?") - Valid conditional/comparison query---The conditional/comparison result actually determines the target of the answer: * "Party names of the National Assembly that has fewer minutes between the 21st and 14th assemblies" -> It must be determined through the check which assembly is referred to. * "Population of the municipality with the largest number of libraries" -> The specific municipality must be determined through the conditional check. - Valid lookup query---Simple argument usage without conditional clauses: * "How many charter bus companies are there in Suwon City?" -> No conditional clause at all, just a lookup. - Valid sequential discovery---The conditional clause is in the query, the [Object] value is not in the query, and actions discover the [Object] from a prior tool result to use in the next step: * "Status of housing land development projects in municipalities with a permitted quarrying area <= 5688 among permitted quarrying companies?" -> No municipality name in the query. actions filter step 1 results to derive the municipality, then use it in step 2. This is a valid multi-step query where the condition actually determines the answer. * Do not judge it as a leak just by seeing the municipality name (e.g., "Icheon City") embedded in the actions. - When the suspected leaking value and the conditional clause target entity are of different types: * "At company X, what is the phone number of...?" (The company name is specified, and the phone number is asked) -> The company name is specified, but the attribute being asked is the phone number. Unless the conditional clause determines the company name, it is not leaked. - When the intermediate result is included in RESPONSE_VALUE---that clause is essential information for evaluation: * If the query says "Check A and what about B?" and response.value also includes the answer to A -> Clause A is not leakage (it is either a parallel call or part of the final answer). The reference RESPONSE_VALUE (GT) and ACTIONS are provided together. Adhere to the following principles: - The primary baseline for determining answer leakage is the query text---actions serve only as a secondary signal. - If a value embedded within actions is also explicitly specified in the query -> Strong signal of answer leakage. - If a value embedded within actions is absent from the query -> Presumed to be a sequential discovery, not answer leakage. - If the answer to an intermediate clause is included in RESPONSE_VALUE -> That clause is required for evaluation, not leakage. Output must strictly follow the JSON schema below (No markdown, comments, or extra text): { "remove": true | false, "explanation": "one-sentence reason" }

Table 23: Prompt utilized in Data Filtering to detect and filter out answer-leaking queries.

Prompt for the Agent under Test [BASE INSTRUCTION] You are a Tool-Executing Agent. Your goal is to answer the user's question accurately and efficiently using the available tools. [RUNTIME-CONTEXT INJECTION -- appended only when the task defines a system state] Use the information in <RUNTIME_CONTEXT> when it is directly needed to execute the tool calls. <RUNTIME_CONTEXT> {key_1}: {value_1} {key_2}: {value_2} ... </RUNTIME_CONTEXT> [RESPONSE-FORMAT SUFFIX -- selected by the task's response type] (String) At the very end, output your final answer in exactly this format: ANSWER: {} The answer must appear inside the curly braces and nowhere else. (Number) Put your final answer within \boxed{} as a sympy-parseable number or expression only (e.g., \boxed{42}, \boxed{sqrt(2)}). No text or explanations - just the numerical value or mathematical expression. (Judge) At the very end, output your final answer in exactly this format: ANSWER: {} Place your comprehensive and explicit answer inside the braces, fully addressing all aspects of the question.

Table 24: System prompt shown to the agent under test, assembled per task from a base instruction, an optional runtime-context block, and a response-format suffix selected by the task’s response type.

LLM-as-a-Judge Prompt (String / Judge) Your task is to verify whether a student's answer matches the correct string answer. You are NOT expected to derive or infer a new correct answer - your only job is to determine whether the student's final answer semantically matches the provided correct answer. Grading Guidelines: - The student answer must express the same meaning as the correct answer. - Minor grammatical or stylistic differences are allowed. - However, if grammar is so broken that the meaning becomes unclear, confusing, or significantly degraded, REJECT. - Synonyms and paraphrases are acceptable as long as they preserve meaning. - Additional irrelevant information is acceptable only if it does not contradict or obscure the core meaning. - If the student's answer contains additional information beyond the correct answer: ACCEPT if it does not distort the overall meaning and introduces no logical contradiction. - If the student's response contains mutually contradictory statements or conclusions, REJECT. - If the student's response consists of or contains meaningless repetition of tokens, words, or phrases, REJECT. [Student Question, Student Answer, and Correct Answer inserted here.] Go step by step through the grading criteria. Do not draw any conclusions until you have finished your reasoning. End your response with exactly one of the following two lines (no markdown, no bold, no extra punctuation): Decision: ACCEPT Decision: REJECT

Table 25: LLM-as-a-Judge prompt shared by the string and judge response types, used as a fallback when deterministic extraction and matching fail.

LLM-as-a-Judge Prompt (Number) Your task is to verify whether a student's answer to a specific question is correct, based on a provided correct answer. You are not expected to solve the question yourself - only to evaluate whether the student's final answer matches the correct one. Grading Guidelines: - Only the final answer should be evaluated. Disregard mistakes made during intermediate steps. - Ignore formatting issues such as LaTeX errors, but do not ignore mathematically significant notation errors. - Any discrepancy that alters the mathematical meaning of the answer should result in rejection. - If the student's answer is a computable expression (e.g., "1 + 2 + 3", "sqrt(16) + 1"), compute it first, then compare. Do not reject solely because it is not in simplified form. - Final answers may be unsimplified or in equivalent forms. - Missing units can be ignored. - When comparing decimals, use the shorter decimal as the precision standard; if the longer one rounds to match, ACCEPT (e.g., correct 3.14 vs. student 3.14159 -> ACCEPT). - For non-math questions, the answer need not mirror the correct answer in all details, as long as it reasonably and completely addresses the core aspects. - If the response contains mutually contradictory statements, REJECT. - If the response consists of or contains meaningless repetition of tokens, words, or phrases, REJECT. [Student Question, Student Answer, and Correct Answer inserted here.] Go step by step through the grading criteria. Do not draw any conclusions until you have finished your reasoning. End your response with exactly one of the following two lines (no markdown, no bold, no extra punctuation): Decision: ACCEPT Decision: REJECT

Table 26: LLM-as-a-Judge prompt for the number response type, extending the grading guidelines with numericequivalence rules.

Record · ID 660841 · SHA-256 23c1ecaa4a6f3c35
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.