期刊界 All Journals 搜尽天下杂志传播学术成果专业期刊搜索期刊信息化学术搜索

1.

Inverse folding of RNA pseudoknot structures

James ZM Gao Linda YM Li Christian M Reidys 《Algorithms for molecular biology : AMB》2010,5(1):27

Background

RNA exhibits a variety of structural configurations. Here we consider a structure to be tantamount to the noncrossing Watson-Crick and G-U-base pairings (secondary structure) and additional cross-serial base pairs. These interactions are called pseudoknots and are observed across the whole spectrum of RNA functionalities. In the context of studying natural RNA structures, searching for new ribozymes and designing artificial RNA, it is of interest to find RNA sequences folding into a specific structure and to analyze their induced neutral networks. Since the established inverse folding algorithms, RNAinverse, RNA-SSD as well as INFO-RNA are limited to RNA secondary structures, we present in this paper the inverse folding algorithm Inv which can deal with 3-noncrossing, canonical pseudoknot structures. 相似文献

2.

RNACompress: Grammar-based compression and informational complexity measurement of RNA secondary structure 总被引：1，自引：0，他引：1

Qi Liu Yu Yang Chun Chen Jiajun Bu Yin Zhang Xiuzi Ye 《BMC bioinformatics》2008,9(1):176

Background

With the rapid emergence of RNA databases and newly identified non-coding RNAs, an efficient compression algorithm for RNA sequence and structural information is needed for the storage and analysis of such data. Although several algorithms for compressing DNA sequences have been proposed, none of them are suitable for the compression of RNA sequences with their secondary structures simultaneously. This kind of compression not only facilitates the maintenance of RNA data, but also supplies a novel way to measure the informational complexity of RNA structural data, raising the possibility of studying the relationship between the functional activities of RNA structures and their complexities, as well as various structural properties of RNA based on compression. 相似文献

3.

4SALE – A tool for synchronous RNA sequence and secondary structure alignment and editing

Philipp N Seibel Tobias Müller Thomas Dandekar Jörg Schultz Matthias Wolf 《BMC bioinformatics》2006,7(1):498-7

Background

In sequence analysis the multiple alignment builds the fundament of all proceeding analyses. Errors in an alignment could strongly influence all succeeding analyses and therefore could lead to wrong predictions. Hand-crafted and hand-improved alignments are necessary and meanwhile good common practice. For RNA sequences often the primary sequence as well as a secondary structure consensus is well known, e.g., the cloverleaf structure of the t-RNA. Recently, some alignment editors are proposed that are able to include and model both kinds of information. However, with the advent of a large amount of reliable RNA sequences together with their solved secondary structures (available from e.g. the ITS2 Database), we are faced with the problem to handle sequences and their associated secondary structures synchronously. 相似文献

4.

Enhancement of accuracy and efficiency for RNA secondary structure prediction by sequence segmentation and MapReduce

Boyu Zhang Daniel T Yehdego Kyle L Johnson Ming-Ying Leung Michela Taufer 《BMC structural biology》2013,13(Z1):S3

Background

Ribonucleic acid (RNA) molecules play important roles in many biological processes including gene expression and regulation. Their secondary structures are crucial for the RNA functionality, and the prediction of the secondary structures is widely studied. Our previous research shows that cutting long sequences into shorter chunks, predicting secondary structures of the chunks independently using thermodynamic methods, and reconstructing the entire secondary structure from the predicted chunk structures can yield better accuracy than predicting the secondary structure using the RNA sequence as a whole. The chunking, prediction, and reconstruction processes can use different methods and parameters, some of which produce more accurate predictions than others. In this paper, we study the prediction accuracy and efficiency of three different chunking methods using seven popular secondary structure prediction programs that apply to two datasets of RNA with known secondary structures, which include both pseudoknotted and non-pseudoknotted sequences, as well as a family of viral genome RNAs whose structures have not been predicted before. Our modularized MapReduce framework based on Hadoop allows us to study the problem in a parallel and robust environment.

Results

On average, the maximum accuracy retention values are larger than one for our chunking methods and the seven prediction programs over 50 non-pseudoknotted sequences, meaning that the secondary structure predicted using chunking is more similar to the real structure than the secondary structure predicted by using the whole sequence. We observe similar results for the 23 pseudoknotted sequences, except for the NUPACK program using the centered chunking method. The performance analysis for 14 long RNA sequences from the Nodaviridae virus family outlines how the coarse-grained mapping of chunking and predictions in the MapReduce framework exhibits shorter turnaround times for short RNA sequences. However, as the lengths of the RNA sequences increase, the fine-grained mapping can surpass the coarse-grained mapping in performance.

Conclusions

By using our MapReduce framework together with statistical analysis on the accuracy retention results, we observe how the inversion-based chunking methods can outperform predictions using the whole sequence. Our chunk-based approach also enables us to predict secondary structures for very long RNA sequences, which is not feasible with traditional methods alone.

相似文献

5.

Correlation between nucleotide composition and folding energy of coding sequences with special attention to wobble bases

Jan C Biro 《Theoretical biology & medical modelling》2008,5(1):14

Background

The secondary structure and complexity of mRNA influences its accessibility to regulatory molecules (proteins, micro-RNAs), its stability and its level of expression. The mobile elements of the RNA sequence, the wobble bases, are expected to regulate the formation of structures encompassing coding sequences. 相似文献

6.

Improving consensus structure by eliminating averaging artifacts

Dukka B KC 《BMC structural biology》2009,9(1):12-9

Background

Common structural biology methods (i.e., NMR and molecular dynamics) often produce ensembles of molecular structures. Consequently, averaging of 3D coordinates of molecular structures (proteins and RNA) is a frequent approach to obtain a consensus structure that is representative of the ensemble. However, when the structures are averaged, artifacts can result in unrealistic local geometries, including unphysical bond lengths and angles. 相似文献

7.

RNA STRAND: The RNA Secondary Structure and Statistical Analysis Database

Mirela Andronescu Vera Bereg Holger H Hoos Anne Condon 《BMC bioinformatics》2008,9(1):340

Background

The ability to access, search and analyse secondary structures of a large set of known RNA molecules is very important for deriving improved RNA energy models, for evaluating computational predictions of RNA secondary structures and for a better understanding of RNA folding. Currently there is no database that can easily provide these capabilities for almost all RNA molecules with known secondary structures. 相似文献

8.

Accelerated probabilistic inference of RNA structure evolution

Ian?Holmes Email author 《BMC bioinformatics》2005,6(1):73

Background

Pairwise stochastic context-free grammars (Pair SCFGs) are powerful tools for evolutionary analysis of RNA, including simultaneous RNA sequence alignment and secondary structure prediction, but the associated algorithms are intensive in both CPU and memory usage. The same problem is faced by other RNA alignment-and-folding algorithms based on Sankoff's 1985 algorithm. It is therefore desirable to constrain such algorithms, by pre-processing the sequences and using this first pass to limit the range of structures and/or alignments that can be considered. 相似文献

9.

Directed acyclic graph kernels for structural RNA analysis

Kengo Sato Toutai Mituyama Kiyoshi Asai Yasubumi Sakakibara 《BMC bioinformatics》2008,9(1):318

Background

Recent discoveries of a large variety of important roles for non-coding RNAs (ncRNAs) have been reported by numerous researchers. In order to analyze ncRNAs by kernel methods including support vector machines, we propose stem kernels as an extension of string kernels for measuring the similarities between two RNA sequences from the viewpoint of secondary structures. However, applying stem kernels directly to large data sets of ncRNAs is impractical due to their computational complexity. 相似文献

10.

Integrated multi-level quality control for proteomic profiling studies using mass spectrometry

David A Cairns David N Perkins Anthea J Stanley Douglas Thompson Jennifer H Barrett Peter J Selby Rosamonde E Banks 《BMC bioinformatics》2008,9(1):1-19

Background

Evolutionary conservation of RNA secondary structure is a typical feature of many functional non-coding RNAs. Since almost all of the available methods used for prediction and annotation of non-coding RNA genes rely on this evolutionary signature, accurate measures for structural conservation are essential.

Results

We systematically assessed the ability of various measures to detect conserved RNA structures in multiple sequence alignments. We tested three existing and eight novel strategies that are based on metrics of folding energies, metrics of single optimal structure predictions, and metrics of structure ensembles. We find that the folding energy based SCI score used in the RNAz program and a simple base-pair distance metric are by far the most accurate. The use of more complex metrics like for example tree editing does not improve performance. A variant of the SCI performed particularly well on highly conserved alignments and is thus a viable alternative when only little evolutionary information is available. Surprisingly, ensemble based methods that, in principle, could benefit from the additional information contained in sub-optimal structures, perform particularly poorly. As a general trend, we observed that methods that include a consensus structure prediction outperformed equivalent methods that only consider pairwise comparisons.

Conclusion

Structural conservation can be measured accurately with relatively simple and intuitive metrics. They have the potential to form the basis of future RNA gene finders, that face new challenges like finding lineage specific structures or detecting mis-aligned sequences. 相似文献

11.

A method for rapid similarity analysis of RNA secondary structures

Na Liu Tianming Wang 《BMC bioinformatics》2006,7(1):493

Background

Owing to the rapid expansion of RNA structure databases in recent years, efficient methods for structure comparison are in demand for function prediction and evolutionary analysis. Usually, the similarity of RNA secondary structures is evaluated based on tree models and dynamic programming algorithms. We present here a new method for the similarity analysis of RNA secondary structures. 相似文献

12.

Empirical comparison of cross-platform normalization methods for gene expression data

Jason Rudy Faramarz Valafar 《BMC bioinformatics》2011,12(1):1-22

Background

The prediction of secondary structure, i.e. the set of canonical base pairs between nucleotides, is a first step in developing an understanding of the function of an RNA sequence. The most accurate computational methods predict conserved structures for a set of homologous RNA sequences. These methods usually suffer from high computational complexity. In this paper, TurboFold, a novel and efficient method for secondary structure prediction for multiple RNA sequences, is presented.

Results

TurboFold takes, as input, a set of homologous RNA sequences and outputs estimates of the base pairing probabilities for each sequence. The base pairing probabilities for a sequence are estimated by combining intrinsic information, derived from the sequence itself via the nearest neighbor thermodynamic model, with extrinsic information, derived from the other sequences in the input set. For a given sequence, the extrinsic information is computed by using pairwise-sequence-alignment-based probabilities for co-incidence with each of the other sequences, along with estimated base pairing probabilities, from the previous iteration, for the other sequences. The extrinsic information is introduced as free energy modifications for base pairing in a partition function computation based on the nearest neighbor thermodynamic model. This process yields updated estimates of base pairing probability. The updated base pairing probabilities in turn are used to recompute extrinsic information, resulting in the overall iterative estimation procedure that defines TurboFold. TurboFold is benchmarked on a number of ncRNA datasets and compared against alternative secondary structure prediction methods. The iterative procedure in TurboFold is shown to improve estimates of base pairing probability with each iteration, though only small gains are obtained beyond three iterations. Secondary structures composed of base pairs with estimated probabilities higher than a significance threshold are shown to be more accurate for TurboFold than for alternative methods that estimate base pairing probabilities. TurboFold-MEA, which uses base pairing probabilities from TurboFold in a maximum expected accuracy algorithm for secondary structure prediction, has accuracy comparable to the best performing secondary structure prediction methods. The computational and memory requirements for TurboFold are modest and, in terms of sequence length and number of sequences, scale much more favorably than joint alignment and folding algorithms.

Conclusions

TurboFold is an iterative probabilistic method for predicting secondary structures for multiple RNA sequences that efficiently and accurately combines the information from the comparative analysis between sequences with the thermodynamic folding model. Unlike most other multi-sequence structure prediction methods, TurboFold does not enforce strict commonality of structures and is therefore useful for predicting structures for homologous sequences that have diverged significantly. TurboFold can be downloaded as part of the RNAstructure package at http://rna.urmc.rochester.edu. 相似文献

13.

A method for aligning RNA secondary structures and its application to RNA motif detection

Jianghui?Liu Jason?TL?Wang Jun?Hu Bin?Tian Email author 《BMC bioinformatics》2005,6(1):89

Background

Alignment of RNA secondary structures is important in studying functional RNA motifs. In recent years, much progress has been made in RNA motif finding and structure alignment. However, existing tools either require a large number of prealigned structures or suffer from high time complexities. This makes it difficult for the tools to process RNAs whose prealigned structures are unavailable or process very large RNA structure databases. 相似文献

14.

Thioredoxin 1 as a serum marker for breast cancer and its use in combination with CEA or CA15-3 for improving the sensitivity of breast cancer diagnoses

Byung-Joon Park Mee-Kyung Cha Il-Han Kim 《BMC research notes》2014,7(1):1-12

Background

Influenza B and C are single-stranded RNA viruses that cause yearly epidemics and infections. Knowledge of RNA secondary structure generated by influenza B and C will be helpful in further understanding the role of RNA structure in the progression of influenza infection.

Findings

All available protein-coding sequences for influenza B and C were analyzed for regions with high potential for functional RNA secondary structure. On the basis of conserved RNA secondary structure with predicted high thermodynamic stability, putative structures were identified that contain splice sites in segment 8 of influenza B and segments 6 and 7 of influenza C. The sequence in segment 6 also contains three unused AUG start codon sites that are sequestered within a hairpin structure.

Conclusions

When added to previous studies on influenza A, the results suggest that influenza splicing may share common structural strategies for regulation of splicing. In particular, influenza 3′ splice sites are predicted to form secondary structures that can switch conformation to regulate splicing. Thus, these RNA structures present attractive targets for therapeutics aimed at targeting one or the other conformation. 相似文献

15.

Degeneracy and genetic assimilation in RNA evolution

Reza Rezazadegan Christian Reidys 《BMC bioinformatics》2018,19(1):543

Background

The neutral theory of Motoo Kimura stipulates that evolution is mostly driven by neutral mutations. However adaptive pressure eventually leads to changes in phenotype that involve non-neutral mutations. The relation between neutrality and adaptation has been studied in the context of RNA before and here we further study transitional mutations in the context of degenerate (plastic) RNA sequences and genetic assimilation. We propose quasineutral mutations, i.e. mutations which preserve an element of the phenotype set, as minimal mutations and study their properties. We also propose a general probabilistic interpretation of genetic assimilation and specialize it to the Boltzmann ensemble of RNA sequences.

Results

We show that degenerate sequences i.e. sequences with more than one structure at the MFE level have the highest evolvability among all sequences and are central to evolutionary innovation. Degenerate sequences also tend to cluster together in the sequence space. The selective pressure in an evolutionary simulation causes the population to move towards regions with more degenerate sequences, i.e. regions at the intersection of different neutral networks, and this causes the number of such sequences to increase well beyond the average percentage of degenerate sequences in the sequence space. We also observe that evolution by quasineutral mutations tends to conserve the number of base pairs in structures and thereby maintains structural integrity even in the presence of pressure to the contrary.

Conclusions

We conclude that degenerate RNA sequences play a major role in evolutionary adaptation.

相似文献

16.

Identification of consensus RNA secondary structures using suffix arrays

Mohammad Anwar Truong Nguyen Marcel Turcotte 《BMC bioinformatics》2006,7(1):244-15

Background

The identification of a consensus RNA motif often consists in finding a conserved secondary structure with minimum free energy in an ensemble of aligned sequences. However, an alignment is often difficult to obtain without prior structural information. Thus the need for tools to automate this process. 相似文献

17.

Comparative genomics reveals 104 candidate structured RNAs from bacteria,archaea, and their metagenomes

Zasha Weinberg Joy X Wang Jarrod Bogue Jingying Yang Keith Corbino Ryan H Moy Ronald R Breaker 《Genome biology》2010,11(3):R31

Background

Structured noncoding RNAs perform many functions that are essential for protein synthesis, RNA processing, and gene regulation. Structured RNAs can be detected by comparative genomics, in which homologous sequences are identified and inspected for mutations that conserve RNA secondary structure. 相似文献

18.

A novel representation of RNA secondary structure based on element-contact graphs

Wenjie Shu Xiaochen Bo Zhiqiang Zheng Shengqi Wang 《BMC bioinformatics》2008,9(1):188

Background

Depending on their specific structures, noncoding RNAs (ncRNAs) play important roles in many biological processes. Interest in developing new topological indices based on RNA graphs has been revived in recent years, as such indices can be used to compare, identify and classify RNAs. Although the topological indices presented before characterize the main topological features of RNA secondary structures, information on RNA structural details is ignored to some degree. Therefore, it is necessity to identify topological features with low degeneracy based on complete and fine-grained RNA graphical representations. 相似文献

19.

The Subviral RNA Database: a toolbox for viroids, the hepatitis delta virus and satellite RNAs research

Lynda Rocheleau Martin Pelchat 《BMC microbiology》2006,6(1):24-5

Background

Viroids, satellite RNAs, satellites viruses and the human hepatitis delta virus form the 'brotherhood' of the smallest known infectious RNA agents, known as the subviral RNAs. For most of these species, it is generally accepted that characteristics such as cell movement, replication, host specificity and pathogenicity are encoded in their RNA sequences and their resulting RNA structures. Although many sequences are indexed in publicly available databases, these sequence annotation databases do not provide the advanced searches and data manipulation capability for identifying and characterizing subviral RNA motifs. 相似文献

20.

An image processing approach to computing distances between RNA secondary structures dot plots

Tor Ivry Shahar Michal Assaf Avihoo Guillermo Sapiro Danny Barash 《Algorithms for molecular biology : AMB》2009,4(1):4

Background

Computing the distance between two RNA secondary structures can contribute in understanding the functional relationship between them. When used repeatedly, such a procedure may lead to finding a query RNA structure of interest in a database of structures. Several methods are available for computing distances between RNAs represented as strings or graphs, but none utilize the RNA representation with dot plots. Since dot plots are essentially digital images, there is a clear motivation to devise an algorithm for computing the distance between dot plots based on image processing methods. 相似文献