Close the adaptive-routing performance gap
Develop practical adaptive-routing signals that close the approximately 14-percentage-point accuracy gap between the best tested routing method and the oracle under the evaluated multilingual natural language inference cascade, thereby making routing preferable to always-expensive inference under comparable compute budgets.
References
Whether the remaining gap relative to computationally expensive approaches such as cross-encoder reranking or LLM-based verification can be closed without additional inference remains an open question.
What the evidence supports is that the tested representation and confidence signals do not recover enough of that headroom to make routing worthwhile under the tested models, signals, and compute budgets---no practical method we evaluated exceeds simply running the expensive model on every input. The gap between what is achievable and what these signals achieve is roughly 14 accuracy points, and closing it is an open problem rather than a closed one.