End-to-end comparison with prior synthesis systems

Establish how existing function-calling data-synthesis systems perform in end-to-end comparisons with EDGE when run on the same 2,318-tool Korean public API inventory and used to produce training data for the same evaluation setting.

Background

The paper compares EDGE primarily with controlled variants of its own pipeline, including a schema-only skeleton that does not execute proposed dependencies against live endpoints. This design isolates the contribution of execution grounding but does not constitute a full system-level comparison with prior synthesis methods.

The unresolved comparison would require running an existing synthesis system end to end on the paper’s tool inventory, training models on the resulting data, and evaluating them under the same KOPA-Bench conditions. Such an experiment would clarify whether EDGE’s gains arise specifically from its execution-grounded dynamic graph and high-cardinality trajectory handling relative to complete alternative pipelines.

References

Running an existing synthesis system end to end on our tool inventory and training on its output remains outside the scope of the present study, so system-level comparisons with prior synthesis pipelines are left open.

Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe  (2609.05395 - Kim et al., 4 Sep 2026) in Section: Limitations, paragraph “Scope of comparison”