Model-scale and data effects in consequence-aware ATC understanding

Determine how model scale and training-data effects influence performance under consequence-aware evaluation for safety-critical air traffic control language understanding, extending beyond the single 8B open-weight backbone used in the fine-tuning study.

Background

The paper evaluates several models in zero-shot and prompting settings but conducts fine-tuning experiments with only one 8B open-weight backbone, Qwen3-8B. Consequently, the reported gains from risk-aware fine-tuning do not establish how the method behaves across substantially different model sizes or training-data regimes.

The unresolved issue is explicitly identified as a limitation because understanding scale and data effects is necessary for assessing whether the observed consequence-aware performance improvements generalize beyond the particular fine-tuning configuration studied.

References

Our fine-tuning study uses one 8B open-weight backbone, so model scale and data effects remain open.

Beyond Semantic Accuracy: Consequence-Aware Evaluation for Safety-Critical Language Understanding  (2608.24621 - Chang et al., 25 Aug 2026) in Limitations section