跳到主要导航 跳到搜索 跳到主要内容

A mechanistic interpretability perspective on personality in large language models

  • East China Normal University
  • Pennsylvania State University
  • Xiaohongshu
  • Shanghai Open University

科研成果: 期刊稿件文章同行评审

摘要

Large language models (LLMs) are increasingly reported to exhibit stable personality-like behaviors under standardized psychometric evaluations such as the Big Five, yet how such behaviors arise from internal computation remains unclear. In this work, we take a mechanistic perspective and aim to characterize how Transformer computation paths give rise to personality-consistent behavior. To enable this analysis, we introduce TraitTrace, to our knowledge the first dataset specifically designed for mechanistic studies of personality in LLMs, comprising 1800 human-curated situational prompts for circuit discovery and validation. Our analysis shows that personality-consistent behaviors are supported by compact circuits that involve only a small fraction of model components, rather than being diffusely distributed across the network. Across Llama2-7B-Chat and Phi-2, the extracted personality circuits achieve near-full-model Hit@10 performance, with an average score of 0.99, while retaining, on average, only about 8% of nodes and fewer than 0.05% of edges for each trait and level. Structural analyses reveal that high and low levels of the same trait are implemented by circuits that share similar nodes but differ in routing. Additionally, we find that Neuroticism exhibits comparatively lower cross-trait overlap, consistent with prior psychometric observations that it is less correlated with other Big Five traits. Furthermore, causal intervention experiments reveal that a small number of early-layer nodes act as bottlenecks that disproportionately govern personality expression. Together, these results provide a circuit-level account of personality in LLMs, bridging behavioral assessment and mechanistic interpretability, and suggesting new directions for more interpretable and targeted personality control and alignment.

源语言英语
文章编号104962
期刊Information Processing and Management
63
8
DOI
出版状态已出版 - 12月 2026

学术指纹

探究 'A mechanistic interpretability perspective on personality in large language models' 的科研主题。它们共同构成独一无二的学术指纹。

引用此