跳到主要导航 跳到搜索 跳到主要内容

Efficient Mining Multi-Mers in a Variety of Biological Sequences

  • Jingsong Zhang
  • , Jianmei Guo
  • , Ming Zhang
  • , Xiangtian Yu
  • , Xiaoqing Yu
  • , Weifeng Guo
  • , Tao Zeng*
  • , Luonan Chen*
  • *此作品的通讯作者
  • CAS - Center for Excellence in Molecular Cell Science
  • Alibaba Group Holding Ltd.
  • Naval Medical University
  • Shanghai Institute of Technology
  • Northwestern Polytechnical University Xian

科研成果: 期刊稿件文章同行评审

摘要

Counting the occurrence frequency of each kk-mer in a biological sequence is a preliminary yet important step in many bioinformatics applications. However, most kk-mer counting algorithms rely on a given kk to produce single-length kk-mers, which is inefficient for sequence analysis for different kk. Moreover, existing kk-mer counters focus more on DNA and RNA sequences and less on protein ones. In practice, the analysis of kk-mers in protein sequences can provide substantial biological insights in structure, function, and evolution. To this end, an efficient algorithm, called MulMer (Multiple-Mer mining), is proposed to mine kk-mers of various lengths termed multi-mers via inverted-index technique, which is orders of magnitude faster than the conventional forward-index methods. Moreover, to the best of our knowledge, MulMer is the first able to mine multi-mers in a variety of sequences, including DNA, RNA, and protein sequences.

源语言英语
文章编号8341507
页(从-至)949-958
页数10
期刊IEEE/ACM Transactions on Computational Biology and Bioinformatics
17
3
DOI
出版状态已出版 - 1 5月 2020
已对外发布

指纹

探究 'Efficient Mining Multi-Mers in a Variety of Biological Sequences' 的科研主题。它们共同构成独一无二的指纹。

引用此