Determine whether lexical similarity rates are differentially biased across corpora

Determine whether the use of TF-IDF cosine similarity biases measured redundancy rates unequally across the real MCP, BFCL v4, and UltraTool corpora, despite measuring lexical rather than semantic equivalence.

Background

The paper measures tool redundancy using TF-IDF cosine similarity. This captures lexical overlap but may fail to identify semantically equivalent tools whose descriptions use different wording, thereby potentially lowering all reported redundancy rates. The authors explicitly state that they have not established whether this limitation affects the three corpora equally, leaving the comparability of the measured rates unresolved.

References

This biases all three rates downward and we have no reason to think it biases them unequally, but we have not shown that.

What a Random Draw from the MCP Registry Contains, and What Tool-Use Benchmarks Contain Instead  (2609.10962 - Afsar, 10 Sep 2026) in Section 7, Threats to validity, paragraph “Similarity is lexical”