首页 | 本学科首页   官方微博 | 高级检索  
     


Word organization in coding DNA: A mathematical model
Authors:Indranil Mukhopadhyay  Anup Som  Satyabrata Sahoo
Affiliation:(1) Department of Human Genetics, University of Pittsburgh, 15261 Pittsburgh, PA, USA;(2) Center for Evolutionary Functional Genomics, The Biodesign Institute, Arizona State University, 85287-5301 Tempe, AZ, USA;(3) Department of Physics, Raidighi College, WB-743383 Raidighi, India
Abstract:This article deals with the relationship between vocabulary (total number of distinct oligomers or “words”) and text-length (total number of oligomers or “words”) for a coding DNA sequence (CDS). For natural human languages, Heaps established a mathematical formula known as Heaps' law, which relates vocabulary to text-length. Our analysis shows that Heaps' law fails to model this relationship for CDSs. Here we develop a mathematical model to establish the relationship between the number of type of words (vocabulary) and the number of words sampled (text-length) for CDSs, when non-overlapping nucleotide strings with the same length are treated as words. We use tangent-hyperbolic function, which captures the saturation property of vocabulary. Based on the parameters of the model, we formulate a mathematical equation, known as “equation of word organization”, whose parameters essentially indicate that nucleotide organization of coding sequences are different from one another. We also compare the word organization of CDSs with the random word distribution and conclude that a CDS is neither similar to a natural human language nor to a random one. Moreover, these sequences have their unique nucleotide organization and it is completely structured for specific biological functioning. IM and AS contributed equally to this work.
Keywords:Coding DNA  Vocabulary  Text-length  Heaps' law  Mathematical modeling
本文献已被 ScienceDirect SpringerLink 等数据库收录!
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号