跳到主要导航 跳到搜索 跳到主要内容

PertEval: Unveiling Real Knowledge Capacity of LLMs with Knowledge-Invariant Perturbations

  • Jiatong Li
  • , Renjun Hu
  • , Kunzhe Huang
  • , Yan Zhuang
  • , Qi Liu*
  • , Mengxiao Zhu
  • , Xing Shi
  • , Wei Lin
  • *此作品的通讯作者
  • University of Science and Technology of China
  • Alibaba Group Holding Ltd.

科研成果: 期刊稿件会议文章同行评审

摘要

Expert-designed close-ended benchmarks are indispensable in assessing the knowledge capacity of large language models (LLMs). Despite their widespread use, concerns have mounted regarding their reliability due to limited test scenarios and an unavoidable risk of data contamination. To rectify this, we present PertEval, a toolkit devised for in-depth probing of LLMs' knowledge capacity through knowledge-invariant perturbations. These perturbations employ human-like restatement techniques to generate on-the-fly test samples from static benchmarks, meticulously retaining knowledge-critical content while altering irrelevant details. Our toolkit further includes a suite of response consistency analyses that compare performance on raw vs. perturbed test sets to precisely assess LLMs' genuine knowledge capacity. Six representative LLMs are re-evaluated using PertEval. Results reveal significantly inflated performance of the LLMs on raw benchmarks, including an absolute 25.8% overestimation for GPT-4. Additionally, through a nuanced response pattern analysis, we discover that PertEval retains LLMs' uncertainty to specious knowledge, and reveals their potential rote memorization to correct options which leads to overestimated performance. We also find that the detailed response consistency analyses by PertEval could illuminate various weaknesses in existing LLMs' knowledge mastery and guide the development of refinement. Our findings provide insights for advancing more robust and genuinely knowledgeable LLMs. Our code is available at https://github.com/aigc-apps/PertEval.

源语言英语
期刊Advances in Neural Information Processing Systems
37
出版状态已出版 - 2024
已对外发布
活动38th Conference on Neural Information Processing Systems, NeurIPS 2024 - Vancouver, 加拿大
期限: 9 12月 202415 12月 2024

指纹

探究 'PertEval: Unveiling Real Knowledge Capacity of LLMs with Knowledge-Invariant Perturbations' 的科研主题。它们共同构成独一无二的指纹。

引用此