Disentangle value differences from leakage propensity in evaluation scores
Determine whether between-model differences in bias on value leakage evaluations arise primarily from differences in underlying model values or from differences in propensity for value leakage, thereby clarifying whether lower measured bias reflects weaker preferences (for example, weaker pro-company or moral preferences) or greater resistance to leaking values into answers.
References
Second, it is unclear whether different scores on our evaluations stem from different values or from a different tendency to value leakage.
Restricted to the three leaky models, $\rho$ falls to 0.50, which at $n{=}3$ is indistinguishable from no relationship ($p=1.0$). The separation between leaky and resistant models therefore drives the correlation, while ordering inside the leaky group remains unresolved.