跳到主要导航 跳到搜索 跳到主要内容

Analysis on n-gram statistics and linguistic features of whole genome protein sequences

  • Qi Wen Dong*
  • , Xiao Long Wang
  • , Lei Lin
  • *此作品的通讯作者
  • Harbin Inst. of Technol.

科研成果: 期刊稿件文章同行评审

摘要

To obtain the statistical sequence analysis on a large number of genomic and proteomic sequences available for different organisms, the n-grams of whole genome protein sequences from 20 organisms were extracted. Their linguistic features were analyzed by two tests: Zipf power law and Shannon entropy, developed for analysis of natural languages and symbolic sequences. The natural genome proteins and the artificial genome proteins were compared with each other and some statistical features of n-grams were discovered. The results show that: the n-grams of whole genome protein sequences approximately follow the Zipf law when n is larger than 4; the Shannon n-gram entropy of natural genome proteins is lower than that of artificial proteins; a simple unigram model can distinguish different organisms; there exist organism-specific usages of "phrases" in protein sequences. It is suggested that further detailed analysis on n-gram of whole genome protein sequences will result in a powerful model for mapping the relationship of protein sequence, structure and function.

源语言英语
页(从-至)694-698
页数5
期刊Journal of Harbin Institute of Technology (New Series)
15
5
出版状态已出版 - 10月 2008
已对外发布

学术指纹

探究 'Analysis on n-gram statistics and linguistic features of whole genome protein sequences' 的科研主题。它们共同构成独一无二的学术指纹。

引用此