TY - JOUR
T1 - A mechanistic interpretability perspective on personality in large language models
AU - Dan, Yuhao
AU - Yu, Lang
AU - Lin, Jiaju
AU - Chen, Qin
AU - Zhou, Jie
AU - Tian, Junfeng
AU - Bai, Qingchun
AU - He, Liang
N1 - Publisher Copyright:
© 2026 Elsevier Ltd. All rights are reserved, including those for text and data mining, AI training, and similar technologies.
PY - 2026/12
Y1 - 2026/12
N2 - Large language models (LLMs) are increasingly reported to exhibit stable personality-like behaviors under standardized psychometric evaluations such as the Big Five, yet how such behaviors arise from internal computation remains unclear. In this work, we take a mechanistic perspective and aim to characterize how Transformer computation paths give rise to personality-consistent behavior. To enable this analysis, we introduce TraitTrace, to our knowledge the first dataset specifically designed for mechanistic studies of personality in LLMs, comprising 1800 human-curated situational prompts for circuit discovery and validation. Our analysis shows that personality-consistent behaviors are supported by compact circuits that involve only a small fraction of model components, rather than being diffusely distributed across the network. Across Llama2-7B-Chat and Phi-2, the extracted personality circuits achieve near-full-model Hit@10 performance, with an average score of 0.99, while retaining, on average, only about 8% of nodes and fewer than 0.05% of edges for each trait and level. Structural analyses reveal that high and low levels of the same trait are implemented by circuits that share similar nodes but differ in routing. Additionally, we find that Neuroticism exhibits comparatively lower cross-trait overlap, consistent with prior psychometric observations that it is less correlated with other Big Five traits. Furthermore, causal intervention experiments reveal that a small number of early-layer nodes act as bottlenecks that disproportionately govern personality expression. Together, these results provide a circuit-level account of personality in LLMs, bridging behavioral assessment and mechanistic interpretability, and suggesting new directions for more interpretable and targeted personality control and alignment.
AB - Large language models (LLMs) are increasingly reported to exhibit stable personality-like behaviors under standardized psychometric evaluations such as the Big Five, yet how such behaviors arise from internal computation remains unclear. In this work, we take a mechanistic perspective and aim to characterize how Transformer computation paths give rise to personality-consistent behavior. To enable this analysis, we introduce TraitTrace, to our knowledge the first dataset specifically designed for mechanistic studies of personality in LLMs, comprising 1800 human-curated situational prompts for circuit discovery and validation. Our analysis shows that personality-consistent behaviors are supported by compact circuits that involve only a small fraction of model components, rather than being diffusely distributed across the network. Across Llama2-7B-Chat and Phi-2, the extracted personality circuits achieve near-full-model Hit@10 performance, with an average score of 0.99, while retaining, on average, only about 8% of nodes and fewer than 0.05% of edges for each trait and level. Structural analyses reveal that high and low levels of the same trait are implemented by circuits that share similar nodes but differ in routing. Additionally, we find that Neuroticism exhibits comparatively lower cross-trait overlap, consistent with prior psychometric observations that it is less correlated with other Big Five traits. Furthermore, causal intervention experiments reveal that a small number of early-layer nodes act as bottlenecks that disproportionately govern personality expression. Together, these results provide a circuit-level account of personality in LLMs, bridging behavioral assessment and mechanistic interpretability, and suggesting new directions for more interpretable and targeted personality control and alignment.
KW - Big five
KW - Dataset construction
KW - Large language models
KW - Mechanistic interpretability
KW - Personality
KW - Transformer circuits
UR - https://www.scopus.com/pages/publications/105041109492
U2 - 10.1016/j.ipm.2026.104962
DO - 10.1016/j.ipm.2026.104962
M3 - 文章
AN - SCOPUS:105041109492
SN - 0306-4573
VL - 63
JO - Information Processing and Management
JF - Information Processing and Management
IS - 8
M1 - 104962
ER -