Robust generalization of supervised fine-tuning for multilingual API grounding

Determine whether supervised fine-tuning improvements in argument-language consistency extend robustly to unseen APIs, diverse argument structures, and cross-lingual transfer in multilingual API calling.

Background

The paper identifies Argument Language Mismatch (ALM) as a failure in which a LLM selects the correct API but generates argument values in a language inconsistent with the user input or system requirements. Although supervised fine-tuning improves language consistency when training and test distributions are aligned, the authors explicitly identify uncertainty about whether this benefit transfers to settings involving unseen APIs, more varied argument structures, and languages not represented in training.

The experiments address some of these dimensions through a low-API-overlap generalization split and evaluation on unseen languages, but the broader question of robust transfer across diverse API and argument configurations remains explicitly unresolved.

References

However, it remains unclear whether such improvements extend robustly to more challenging settings, such as unseen APIs, diverse argument structures, or cross-lingual transfer.

When the API Speaks the Wrong Language: Revisiting Post-Training for Multilingual Tool Use  (2608.11715 - Chauhan et al., 12 Aug 2026) in Section 1, Introduction, page 1