跳到主要导航 跳到搜索 跳到主要内容

Scalable parallel join for huge tables

  • Nianlong Weng
  • , Minqi Zhou
  • , Ming Chien Shan
  • , Aoying Zhou
  • Fudan University
  • East China Normal University
  • SAP Research

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

The parallel join processing which combines tuples from two or more relational tables together in a parallel manner is becoming more and more important and imperative to be solved, since tables may be huge, especially in this big data era. A few algorithms have already been proposed based on the prevailing mapreduce paradigm, while most of them impose both high communication costs and synchronization costs. In this paper, we propose a novel algorithm for scalable parallel join processing for the column-wise stored data analyzing. To cater for the prevailing deployed Hadoop system, we adopt the Hadoop Distributed File System (HDFS) as the file system across over a large set of machines. Tables are projected (i.e., vertical partition), segmented (i.e., horizontal partition), clustered and placed in a column-wise format over the distributed file system based on Gray Code. By effectively fetching the dedicated tuples from other tables on demand based on an optimized bloom filter strategy, each segment (i.e., partition) is capable in accomplishing the join processing individually with dramatically reduced communication cost, and consequently achieves the desired scalable parallelism. Tuples are transmitted in a demand driven manner across the network, rather than the hash-based movement in the mapreduce paradigm. Our extensive performance studies confirm the effectiveness and efficiency of our methods.

源语言英语
主期刊名Proceedings - 2013 IEEE International Congress on Big Data, BigData 2013
出版商IEEE Computer Society
157-164
页数8
ISBN(印刷版)9780768550060
DOI
出版状态已出版 - 2013
活动2013 IEEE International Congress on Big Data, BigData Congress 2013 - Santa Clara, CA, 美国
期限: 27 6月 20132 7月 2013

出版系列

姓名Proceedings - 2013 IEEE International Congress on Big Data, BigData 2013

会议

会议2013 IEEE International Congress on Big Data, BigData Congress 2013
国家/地区美国
Santa Clara, CA
时期27/06/132/07/13

指纹

探究 'Scalable parallel join for huge tables' 的科研主题。它们共同构成独一无二的指纹。

引用此