跳到主要导航 跳到搜索 跳到主要内容

A survey of datasets in medicine for large language models

  • Deshiwei Zhang
  • , Xiaojuan Xue
  • , Peng Gao
  • , Zhijuan Jin
  • , Menghan Hu*
  • , Yue Wu
  • , Xiayang Ying*
  • *此作品的通讯作者
  • Southeast University, Nanjing
  • East China Normal University
  • Tongji University
  • Shanghai Jiao Tong University

科研成果: 期刊稿件文献综述同行评审

摘要

With the advent of models such as ChatGPT and other models, large language models (LLMs) have demonstrated unprecedented capabilities in understanding and generating natural language, presenting novel opportunities and challenges within the medicine domain. While there have been many studies focusing on the employment of LLMs in medicine, comprehensive reviews of the datasets utilized in this field remain scarce. This survey seeks to address this gap by providing a comprehensive overview of the datasets in medicine fueling LLMs, highlighting their unique characteristics and the critical roles they play at different stages of LLMs’ development: pre-training, fine-tuning, and evaluation. Ultimately, this survey aims to underline the significance of datasets in realizing the full potential of LLMs to innovate and improve healthcare outcomes.

源语言英语
页(从-至)457-478
页数22
期刊Intelligence and Robotics
4
4
DOI
出版状态已出版 - 12月 2024

学术指纹

探究 'A survey of datasets in medicine for large language models' 的科研主题。它们共同构成独一无二的学术指纹。

引用此