摘要
With the advent of models such as ChatGPT and other models, large language models (LLMs) have demonstrated unprecedented capabilities in understanding and generating natural language, presenting novel opportunities and challenges within the medicine domain. While there have been many studies focusing on the employment of LLMs in medicine, comprehensive reviews of the datasets utilized in this field remain scarce. This survey seeks to address this gap by providing a comprehensive overview of the datasets in medicine fueling LLMs, highlighting their unique characteristics and the critical roles they play at different stages of LLMs’ development: pre-training, fine-tuning, and evaluation. Ultimately, this survey aims to underline the significance of datasets in realizing the full potential of LLMs to innovate and improve healthcare outcomes.
| 源语言 | 英语 |
|---|---|
| 页(从-至) | 457-478 |
| 页数 | 22 |
| 期刊 | Intelligence and Robotics |
| 卷 | 4 |
| 期 | 4 |
| DOI | |
| 出版状态 | 已出版 - 12月 2024 |
学术指纹
探究 'A survey of datasets in medicine for large language models' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver