Identify the item feature driving model choices

Identify which feature of the size-tile items drives language-model choices when those choices are item-sensitive but do not vary with the fit cost of discriminating the target, using a manipulation not present in the current item set.

Background

The paper finds that LLMs respond reliably to which size-tile item they receive, but their achieved coordinate does not systematically change with fit_cost, the quantity that measures the description quality sacrificed when selecting the posterior-maximising option instead of the fit-maximising option. The authors rule out the models’ own option preferences and positional effects as explanations.

The unresolved issue is therefore what item property the models are tracking. The current item set does not contain the manipulation needed to distinguish among possible explanations, so the authors explicitly leave the question open rather than proposing a speculative account.

References

This leaves an open question that our data can pose but not answer. Models are item-sensitive on size, so their choices depend on which item they see, yet their coordinate does not move with fit_cost, which is the item property the design varies. Whatever they respond to is a feature of the item uncorrelated with the fit cost of discriminating. We can rule out two candidates. It is not the model's own option preference, which the marginal null removes, and it is not position, because the coordinate is computed on the canonical option and permutation consistency is above the marginal null. Identifying what it is would require a manipulation the current item set does not contain, so we state the question and leave it open rather than speculate.

Consistency Without Alignment: Item-Sensitive Language Models Indistinguishable From Random  (2609.00576 - Huynh, 1 Sep 2026) in Results, subsection “The divergence curve is shallow under the primary rule”

We do not know whether the consistency-without-alignment result is a property of this task in general or is concentrated on the one tile with enough resolution to detect either strategy at all; the coarser tiles may show the same dissociation without our being able to measure it.

Consistency Without Alignment: Item-Sensitive Language Models Indistinguishable From Random  (2609.00576 - Huynh, 1 Sep 2026) in Section “Discussion,” paragraph “One tile carries the primary analysis”