Generality of the confidence confound beyond activation steering

Determine whether the confidence confound identified in activation-steering debiasing also affects inference-time debiasing techniques beyond activation steering.

Background

The study shows that activation-steering debiasing directions learned from biased and anti-biased prompts are strongly aligned with model-confidence directions. Consequently, apparent reductions in bias can result from abstention or increased uncertainty rather than correction of stereotyped associations.

The authors leave unresolved whether this same confidence-related confound occurs in other inference-time debiasing methods, making its broader prevalence and implications for fairness evaluation an open research question.

References

Future work will study whether a debiasing direction which is independent of model confidence can be found, and whether this confound impacts inference-time debiasing techniques beyond activation steering.

— Latent space bias directions in LLMs capture confidence, not fairness  (2610.08559 - Buttigieg et al., 6 Oct 2026) in Section Discussion and conclusions