首页 | 本学科首页   官方微博 | 高级检索  
     


Integrated modeling of protein-coding genes in the Manduca sexta genome using RNA-Seq data from the biochemical model insect
Affiliation:1. School of Life Sciences, Chongqing University, Chongqing, China;2. Postdoctoral Station of Biomedical Engineering, Chongqing University, Chongqing, China;3. State Key Laboratory of Silkworm Genome Biology, Southwest University, Chongqing, China
Abstract:The genome sequence of Manduca sexta was recently determined using 454 technology. Cufflinks and MAKER2 were used to establish gene models in the genome assembly based on the RNA-Seq data and other species' sequences. Aided by the extensive RNA-Seq data from 50 tissue samples at various life stages, annotators over the world (including the present authors) have manually confirmed and improved a small percentage of the models after spending months of effort. While such collaborative efforts are highly commendable, many of the predicted genes still have problems which may hamper future research on this insect species. As a biochemical model representing lepidopteran pests, M. sexta has been used extensively to study insect physiological processes for over five decades. In this work, we assembled Manduca datasets Cufflinks 3.0, Trinity 4.0, and Oases 4.0 to assist the manual annotation efforts and development of Official Gene Set (OGS) 2.0. To further improve annotation quality, we developed methods to evaluate gene models in the MAKER2, Cufflinks, Oases and Trinity assemblies and selected the best ones to constitute MCOT 1.0 after thorough crosschecking. MCOT 1.0 has 18,089 genes encoding 31,666 proteins: 32.8% match OGS 2.0 models perfectly or near perfectly, 11,747 differ considerably, and 29.5% are absent in OGS 2.0. Future automation of this process is anticipated to greatly reduce human efforts in generating comprehensive, reliable models of structural genes in other genome projects where extensive RNA-Seq data are available.
Keywords:Gene annotation  Tobacco hornworm  Automated gene modeling  Arthropod genomics  OGS"  },{"  #name"  :"  keyword"  ,"  $"  :{"  id"  :"  kwrd0040"  },"  $$"  :[{"  #name"  :"  text"  ,"  _"  :"  official gene set  ORF"  },{"  #name"  :"  keyword"  ,"  $"  :{"  id"  :"  kwrd0050"  },"  $$"  :[{"  #name"  :"  text"  ,"  _"  :"  open reading frame  L"  },{"  #name"  :"  keyword"  ,"  $"  :{"  id"  :"  kwrd0060"  },"  $$"  :[{"  #name"  :"  text"  ,"  _"  :"  length  ML"  },{"  #name"  :"  keyword"  ,"  $"  :{"  id"  :"  kwrd0070"  },"  $$"  :[{"  #name"  :"  text"  ,"  _"  :"  match length  QL"  },{"  #name"  :"  keyword"  ,"  $"  :{"  id"  :"  kwrd0080"  },"  $$"  :[{"  #name"  :"  text"  ,"  _"  :"  query length  SL"  },{"  #name"  :"  keyword"  ,"  $"  :{"  id"  :"  kwrd0090"  },"  $$"  :[{"  #name"  :"  text"  ,"  _"  :"  subject length  M"  },{"  #name"  :"  keyword"  ,"  $"  :{"  id"  :"  kwrd0100"  },"  $$"  :[{"  #name"  :"  text"  ,"  _"  :"  MAKER  C"  },{"  #name"  :"  keyword"  ,"  $"  :{"  id"  :"  kwrd0110"  },"  $$"  :[{"  #name"  :"  text"  ,"  _"  :"  Cufflinks  T"  },{"  #name"  :"  keyword"  ,"  $"  :{"  id"  :"  kwrd0120"  },"  $$"  :[{"  #name"  :"  text"  ,"  _"  :"  Trinity  O"  },{"  #name"  :"  keyword"  ,"  $"  :{"  id"  :"  kwrd0130"  },"  $$"  :[{"  #name"  :"  text"  ,"  _"  :"  Oases  U"  },{"  #name"  :"  keyword"  ,"  $"  :{"  id"  :"  kwrd0140"  },"  $$"  :[{"  #name"  :"  text"  ,"  _"  :"  UniProt Arthropoda  Y"  },{"  #name"  :"  keyword"  ,"  $"  :{"  id"  :"  kwrd0150"  },"  $$"  :[{"  #name"  :"  text"  ,"  _"  :"  C/T/O  S"  },{"  #name"  :"  keyword"  ,"  $"  :{"  id"  :"  kwrd0160"  },"  $$"  :[{"  #name"  :"  text"  ,"  _"  :"  similarity ratio of lengths  MLI"  },{"  #name"  :"  keyword"  ,"  $"  :{"  id"  :"  kwrd0170"  },"  $$"  :[{"  #name"  :"  text"  ,"  _"  :"  match length index  S1/S2"  },{"  #name"  :"  keyword"  ,"  $"  :{"  id"  :"  kwrd0180"  },"  $$"  :[{"  #name"  :"  text"  ,"  _"  :"  Selection 1 or 2  “P”"  },{"  #name"  :"  keyword"  ,"  $"  :{"  id"  :"  kwrd0190"  },"  $$"  :[{"  #name"  :"  text"  ,"  _"  :"  perfect  “N”"  },{"  #name"  :"  keyword"  ,"  $"  :{"  id"  :"  kwrd0200"  },"  $$"  :[{"  #name"  :"  text"  ,"  _"  :"  near perfect  “O”"  },{"  #name"  :"  keyword"  ,"  $"  :{"  id"  :"  kwrd0210"  },"  $$"  :[{"  #name"  :"  text"  ,"  _"  :"  okay  “B”"  },{"  #name"  :"  keyword"  ,"  $"  :{"  id"  :"  kwrd0220"  },"  $$"  :[{"  #name"  :"  text"  ,"  _"  :"  bad  “W”"  },{"  #name"  :"  keyword"  ,"  $"  :{"  id"  :"  kwrd0230"  },"  $$"  :[{"  #name"  :"  text"  ,"  _"  :"  worst
本文献已被 ScienceDirect 等数据库收录!
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号