跳到主要导航 跳到搜索 跳到主要内容

Estimate unlabeled-data-distribution for semi-supervised PU learning

  • East China Normal University
  • Fudan University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Traditional supervised classifiers use only labeled data (features/label pairs) as the training set, while the unlabeled data is used as the testing set. In practice, it is often the case that the labeled data is hard to obtain and the unlabeled data contains the instances that belong to the predefined class beyond the labeled data categories. This problem has been widely studied in recent years and the semi-supervised learning is an efficient solution to learn from positive and unlabeled examples(or PU learning). Among all the semi-supervised PU learning methods, it's hard to choose just one approach to fit all unlabeled data distribution. This paper proposes an automatic KL-divergence based semi-supervised learning method by using unlabeled data distribution knowledge. Meanwhile, a new framework is designed to integrate different semi-supervised PU learning algorithms in order to take advantage of the former methods. The experimental results show that (1)data distribution information is very helpful for the semi-supervised PU learning method; (2)the proposed framework can achieve higher precision when compared with the-state-of-the-art method.

源语言英语
主期刊名Web Technologies and Applications - 14th Asia-Pacific Web Conference, APWeb 2012, Proceedings
22-33
页数12
DOI
出版状态已出版 - 2012
活动14th Asia Pacific Web Technology Conference, APWeb 2012 - Kunming, 中国
期限: 11 4月 201213 4月 2012

出版系列

姓名Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
7235 LNCS
ISSN(印刷版)0302-9743
ISSN(电子版)1611-3349

会议

会议14th Asia Pacific Web Technology Conference, APWeb 2012
国家/地区中国
Kunming
时期11/04/1213/04/12

指纹

探究 'Estimate unlabeled-data-distribution for semi-supervised PU learning' 的科研主题。它们共同构成独一无二的指纹。

引用此