Dissecting Persona-Driven Reasoning in Language Models via Activation Patching (2507.20936v1)

Published 28 Jul 2025 in cs.LG, cs.AI, and cs.CL

Abstract: LLMs exhibit remarkable versatility in adopting diverse personas. In this study, we examine how assigning a persona influences a model's reasoning on an objective task. Using activation patching, we take a first step toward understanding how key components of the model encode persona-specific information. Our findings reveal that the early Multi-Layer Perceptron (MLP) layers attend not only to the syntactic structure of the input but also process its semantic content. These layers transform persona tokens into richer representations, which are then used by the middle Multi-Head Attention (MHA) layers to shape the model's output. Additionally, we identify specific attention heads that disproportionately attend to racial and color-based identities.

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Follow-up Questions

We haven't generated follow-up questions for this paper yet.

Generate Now

Dissecting Persona-Driven Reasoning in Language Models via Activation Patching (2507.20936v1)

Summary

Follow-up Questions

Related Papers

Authors (2)