Skip to main navigation Skip to search Skip to main content

A mechanistic interpretability perspective on personality in large language models

  • Yuhao Dan
  • , Lang Yu
  • , Jiaju Lin
  • , Qin Chen*
  • , Jie Zhou
  • , Junfeng Tian
  • , Qingchun Bai
  • , Liang He
  • *Corresponding author for this work
  • East China Normal University
  • Pennsylvania State University
  • Xiaohongshu
  • Shanghai Open University

Research output: Contribution to journalArticlepeer-review

Abstract

Large language models (LLMs) are increasingly reported to exhibit stable personality-like behaviors under standardized psychometric evaluations such as the Big Five, yet how such behaviors arise from internal computation remains unclear. In this work, we take a mechanistic perspective and aim to characterize how Transformer computation paths give rise to personality-consistent behavior. To enable this analysis, we introduce TraitTrace, to our knowledge the first dataset specifically designed for mechanistic studies of personality in LLMs, comprising 1800 human-curated situational prompts for circuit discovery and validation. Our analysis shows that personality-consistent behaviors are supported by compact circuits that involve only a small fraction of model components, rather than being diffusely distributed across the network. Across Llama2-7B-Chat and Phi-2, the extracted personality circuits achieve near-full-model Hit@10 performance, with an average score of 0.99, while retaining, on average, only about 8% of nodes and fewer than 0.05% of edges for each trait and level. Structural analyses reveal that high and low levels of the same trait are implemented by circuits that share similar nodes but differ in routing. Additionally, we find that Neuroticism exhibits comparatively lower cross-trait overlap, consistent with prior psychometric observations that it is less correlated with other Big Five traits. Furthermore, causal intervention experiments reveal that a small number of early-layer nodes act as bottlenecks that disproportionately govern personality expression. Together, these results provide a circuit-level account of personality in LLMs, bridging behavioral assessment and mechanistic interpretability, and suggesting new directions for more interpretable and targeted personality control and alignment.

Original languageEnglish
Article number104962
JournalInformation Processing and Management
Volume63
Issue number8
DOIs
StatePublished - Dec 2026

Keywords

  • Big five
  • Dataset construction
  • Large language models
  • Mechanistic interpretability
  • Personality
  • Transformer circuits

Fingerprint

Dive into the research topics of 'A mechanistic interpretability perspective on personality in large language models'. Together they form a unique fingerprint.

Cite this