首页 | 本学科首页   官方微博 | 高级检索  
   检索      


Identification of repeat structure in large genomes using repeat probability clouds
Authors:Gu Wanjun  Castoe Todd A  Hedges Dale J  Batzer Mark A  Pollock David D
Institution:aDepartment of Biochemistry and Molecular Genetics, University of Colorado School of Medicine, Aurora, CO 80045, USA;bDepartment of Epidemiology, Tulane University Health Sciences Center, New Orleans, LA 70112, USA;cDepartment of Biological Sciences, Biological Computation and Visualization Center, and Center for Bio-Modular Multi-Scale Systems, Louisiana State University, Baton Rouge, LA 70803, USA
Abstract:The identification of repeat structure in eukaryotic genomes can be time-consuming and difficult because of the large amount of information (not, vert, similar3 × 109 bp) that needs to be processed and compared. We introduce a new approach based on exact word counts to evaluate, de novo, the repeat structure present within large eukaryotic genomes. This approach avoids sequence alignment and similarity search, two of the most time-consuming components of traditional methods for repeat identification. Algorithms were implemented to efficiently calculate exact counts for any length oligonucleotide in large genomes. Based on these oligonucleotide counts, oligonucleotide excess probability clouds, or “P-clouds,” were constructed. P-clouds are composed of clusters of related oligonucleotides that occur, as a group, more often than expected by chance. After construction, P-clouds were mapped back onto the genome, and regions of high P-cloud density were identified as repetitive regions based on a sliding window approach. This efficient method is capable of analyzing the repeat content of the entire human genome on a single desktop computer in less than half a day, at least 10-fold faster than current approaches. The predicted repetitive regions strongly overlap with known repeat elements as well as other repetitive regions such as gene families, pseudogenes, and segmental duplicons. This method should be extremely useful as a tool for use in de novo identification of repeat structure in large newly sequenced genomes.
Keywords:Alignment  Complete genome annotation  Oligonucleotide counts  P-clouds  Repeat structure
本文献已被 ScienceDirect PubMed 等数据库收录!
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号