Abstract
This paper describes three learning methods for document filtering that use unlabeled data. The proposed methods are based on a committee of the classifiers which are trained on a small set of labeled data and then augmented by a large number of unlabeled data. By taking advantage of unlabeled data, the effective number of labeled data needed is significantly reduced and the filtering accuracy is increased. The use of unlabeled data is important because obtaining labeled data is difficult and time-consuming, while unlabeled data are abundant. For all proposed methods, the experimental results show that the accuracy is improved up to 9.2% with only two-thirds as many labeled data as the method which does not use unlabeled data.
| Original language | English |
|---|---|
| Pages | 328-333 |
| Number of pages | 6 |
| Publication status | Published - 2001 |
| Event | 2001 IEEE International Symposium on Industrial Electronics Proceedings (ISIE 2001) - Pusan, Korea, Republic of Duration: 12 Jun 2001 → 16 Jun 2001 |
Conference
| Conference | 2001 IEEE International Symposium on Industrial Electronics Proceedings (ISIE 2001) |
|---|---|
| Country/Territory | Korea, Republic of |
| City | Pusan |
| Period | 12/06/01 → 16/06/01 |
Fingerprint
Dive into the research topics of 'Document filtering boosted by unlabeled data'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver