TY - JOUR
T1 - XML filtering based-on probabilistic SLCA
AU - Zhang, Chen Jing
AU - Wang, Xiao Ling
AU - Zhou, Ao Ying
PY - 2014/9/1
Y1 - 2014/9/1
N2 - Uncertain data management is becoming an important research focus. Uncertain management of XML data which is the main store and exchange standard of web data is naturally becoming a hot point. One of the branches is keyword-based search over probabilistic XML. In recent work of keyword search over probabilistic XML, only the independent and the mutually-exclusive relationships among sibling nodes have been discussed. Because of the complexity of representation and computation, more general relationship among sibling nodes has got little attention up to now. This paper addresses the problem of keyword filtering over probabilistic XML data model PrXML{exp, ind, mux}. In the model, exp node is used to represent more general relationship among sibling nodes. tab is defined as keyword distribution probability table of one subtree. The dot product, Cartesian product, and addition operation of tab are also defined. Then the computation of different type of nodes' tab are given. Furthermore, an algorithm of how to obtain SLCAs and the probability of being a SLCA node is also given without generating possible worlds. Finally, the features and efficiency of our method are evaluated with extensive experimental results.
AB - Uncertain data management is becoming an important research focus. Uncertain management of XML data which is the main store and exchange standard of web data is naturally becoming a hot point. One of the branches is keyword-based search over probabilistic XML. In recent work of keyword search over probabilistic XML, only the independent and the mutually-exclusive relationships among sibling nodes have been discussed. Because of the complexity of representation and computation, more general relationship among sibling nodes has got little attention up to now. This paper addresses the problem of keyword filtering over probabilistic XML data model PrXML{exp, ind, mux}. In the model, exp node is used to represent more general relationship among sibling nodes. tab is defined as keyword distribution probability table of one subtree. The dot product, Cartesian product, and addition operation of tab are also defined. Then the computation of different type of nodes' tab are given. Furthermore, an algorithm of how to obtain SLCAs and the probability of being a SLCA node is also given without generating possible worlds. Finally, the features and efficiency of our method are evaluated with extensive experimental results.
KW - Keyword distribution probability table
KW - Keywords filtering
KW - Probabilistic XML
KW - Smallest lowest common ancestor
KW - Uncertain data
UR - https://www.scopus.com/pages/publications/84907575642
U2 - 10.3724/SP.J.1016.2014.01959
DO - 10.3724/SP.J.1016.2014.01959
M3 - 文章
AN - SCOPUS:84907575642
SN - 0254-4164
VL - 37
SP - 1959
EP - 1971
JO - Jisuanji Xuebao/Chinese Journal of Computers
JF - Jisuanji Xuebao/Chinese Journal of Computers
IS - 9
ER -