摘要
Existing distantly supervised relation extractors usually rely on noisy data for both model training and evaluation, which may lead to garbage-in-garbage-out systems. To alleviate the problem, we study whether a small clean dataset could help improve the quality of distantly supervised models. We show that besides getting a more convincing evaluation of models, a small clean dataset also helps us to build more robust denoising models. Specifically, we propose a new criterion for clean instance selection based on influence functions. It collects sample-level evidence for recognizing good instances (which is more informative than loss-level evidence). We also propose a teacher-student mechanism for controlling purity of intermediate results when bootstrapping the clean set. The whole approach is model-agnostic and demonstrates strong performances on both denoising real (NYT) and synthetic noisy datasets.
| 源语言 | 英语 |
|---|---|
| 页(从-至) | 2528-2539 |
| 页数 | 12 |
| 期刊 | Proceedings - International Conference on Computational Linguistics, COLING |
| 卷 | 29 |
| 期 | 1 |
| 出版状态 | 已出版 - 2022 |
| 活动 | 29th International Conference on Computational Linguistics, COLING 2022 - Hybrid, Gyeongju, 韩国 期限: 12 10月 2022 → 17 10月 2022 |
学术指纹
探究 'Few Clean Instances Help Denoising Distant Supervision' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver