摘要
The essence of uncertain data management has been well adopted since data uncertainty widely exists in lots of applications, such as Web, sensor networks, etc. Most of the uncertain data models are based on the possible world semantics. Because the number of the possible worlds will blowup exponentially with the growth of the data set, it is much more challenging to handle uncertain data than deterministic data. In this paper, we take the first attempt to study the rarity, an important statistic that describes the proportion of items with the same frequency, upon uncertain data. We have proposed three novel solutions, including an exact method and an approximate method to compute the rarity of a given frequency respectively, and a method to find the frequency of the maximum rarity. Analysis in theorem and extensive experimental results demonstrate the effectiveness and efficiency of the proposed solutions.
| 源语言 | 英语 |
|---|---|
| 页(从-至) | 2028-2039 |
| 页数 | 12 |
| 期刊 | Science China Information Sciences |
| 卷 | 54 |
| 期 | 10 |
| DOI | |
| 出版状态 | 已出版 - 10月 2011 |
指纹
探究 'Computing rarity on uncertain data' 的科研主题。它们共同构成独一无二的指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver