Training Away the Ranking Cost of Generated Rationales

Determine whether the ranking cost introduced by generating rationales in behavioral language models can be eliminated through faithfulness objectives or evidence-grounded rewards, rather than avoided by using the scored readout at serving time.

Background

The paper finds that, across 13 model–domain cells, reading a decision probability directly from the model generally produces more accurate behavioral ranking than generating a rationale before obtaining the decision. The ranking degradation is associated with abandonment of dominant predictive features and convergence on stock phrasing.

The authors recommend retaining generated rationales for interpretability while sourcing ranking from the scored readout. They explicitly leave unresolved whether training interventions—specifically faithfulness objectives or evidence-grounded rewards—can directly remove the ranking penalty associated with rationale generation, which would avoid relying on a separate scored serving pathway.

References

The remaining question, which we leave open, is whether the ranking cost of generation can be trained away directly, through faithfulness objectives or evidence-grounded rewards, rather than routed around at serving time as we recommend below.

— Scored vs. Generated Readouts in Behavioral Language Models: An Empirical Study of Elicitation Format  (2609.09882 - Kraisingkorn et al., 9 Sep 2026) in Section Conclusion, immediately before Section Deployment Guidance