Generalize the instability ordering across model families and temperatures
Determine whether the ordering of response instability—highest for self-referential subjective-experience questions, intermediate for unresolvable philosophical questions, and lowest for verifiable questions—holds for other large language model families and at sampling temperatures other than 0.7.
References
Whether the same ordering holds for other model families, or at other sampling temperatures, is untested.
These findings open as many questions as they close. First, which task properties set the direction and magnitude of the distortion: the exploratory subjective-task contrast motivates, without establishing, a systematic test that crosses objectivity, demonstrability, and answer familiarity, since a task's position on the intellective-to-judgmental continuum (Steiner, 1972; Laughlin et al., 2006) may govern whether belief-anchored groups over- or under-converge.