首页 | 本学科首页   官方微博 | 高级检索  
相似文献
 共查询到20条相似文献,搜索用时 15 毫秒
1.
Protein chemical shifts encode detailed structural information that is difficult and computationally costly to describe at a fundamental level. Statistical and machine learning approaches have been used to infer correlations between chemical shifts and secondary structure from experimental chemical shifts. These methods range from simple statistics such as the chemical shift index to complex methods using neural networks. Notwithstanding their higher accuracy, more complex approaches tend to obscure the relationship between secondary structure and chemical shift and often involve many parameters that need to be trained. We present hidden Markov models (HMMs) with Gaussian emission probabilities to model the dependence between protein chemical shifts and secondary structure. The continuous emission probabilities are modeled as conditional probabilities for a given amino acid and secondary structure type. Using these distributions as outputs of first‐ and second‐order HMMs, we achieve a prediction accuracy of 82.3%, which is competitive with existing methods for predicting secondary structure from protein chemical shifts. Incorporation of sequence‐based secondary structure prediction into our HMM improves the prediction accuracy to 84.0%. Our findings suggest that an HMM with correlated Gaussian distributions conditioned on the secondary structure provides an adequate generative model of chemical shifts. Proteins 2013; © 2012 Wiley Periodicals, Inc.  相似文献   

2.
Information about the secondary and tertiary structure of a protein sequence can greatly assist biologists in the generation and testing of hypotheses, as well as design of experiments. The PROTINFO server enables users to submit a protein sequence and request a prediction of the three-dimensional (tertiary) structure based on comparative modeling, fold generation and de novo methods developed by the authors. In addition, users can submit NMR chemical shift data and request protein secondary structure assignment that is based on using neural networks to combine the chemical shifts with secondary structure predictions. The server is available at http://protinfo.compbio.washington.edu.  相似文献   

3.
A method for predicting type I and II β-turns using nuclear magnetic resonance (NMR) chemical shifts is proposed. Isolated β-turn chemical-shift data were collected from 1,798 protein chains. One-dimensional statistical analyses on chemical-shift data of three classes β-turn (type I, II, and VIII) showed different distributions at four positions, (i) to (i + 3). Considering the central two residues of type I β-turns, the mean values of Cο, Cα, HN, and NH chemical shifts were generally (i + 1) > (i + 2). The mean values of Cβ and Hα chemical shifts were (i + 1) < (i + 2). The distributions of the central two residues in type II and VIII β-turns were also distinguishable by trends of chemical shift values. Two-dimensional cluster analyses on chemical-shift data show positional distributions more clearly. Based on these propensities of chemical shift classified as a function of position, rules were derived using scoring matrices for four consecutive residues to predict type I and II β-turns. The proposed method achieves an overall prediction accuracy of 83.2 and 84.2 % with the Matthews correlation coefficient values of 0.317 and 0.632 for type I and II β-turns, indicating that its higher accuracy for type II turn prediction. The results show that it is feasible to use NMR chemical shifts to predict the β-turn types in proteins. The proposed method can be incorporated into other chemical-shift based protein secondary structure prediction methods.  相似文献   

4.
Abstract

The transient secondary structure and dynamics of an intrinsically unstructured linker domain from the 70 kDa subunit of human replication protein A was investigated using solution state NMR. Stable secondary structure, inferred from large secondary chemical shifts, was observed for a segment of the intrinsically unstructured linker domain when it is attached to an N-terminal protein interaction domain. Results from NMR relaxation experiments showed the rotational diffusion for this segment of the intrinsically unstructured linker domain to be correlated with the N-terminal protein interaction domain. When the N-terminal domain is removed, the stable secondary structure is lost and faster rotational diffusion is observed. The large secondary chemical shifts were used to calculate phi and psidihedral angles and these dihedral angles were used to build a backbone structural model. Restrained molecular dynamics were performed on this new structure using the chemical shift based dihedral angles and a single NOE distance as restraints. In the resulting family of structures a large, solvent exposed loop was observed for the segment of the intrinsically unstructured linker domain that had large secondary chemical shifts.  相似文献   

5.
Chemical shifts of amino acids in proteins are the most sensitive and easily obtainable NMR parameters that reflect the primary, secondary, and tertiary structures of the protein. In recent years, chemical shifts have been used to identify secondary structure in peptides and proteins, and it has been confirmed that 1Hα, 13Cα, 13Cβ, and 13C′ NMR chemical shifts for all 20 amino acids are sensitive to their secondary structure. Currently, most of the methods are purely based on one-dimensional statistical analyses of various chemical shifts for each residue to identify protein secondary structure. However, it is possible to achieve an increased accuracy from the two-dimensional analyses of these chemical shifts. The 2DCSi approach performs two-dimension cluster analyses of 1Hα, 1HN, 13Cα, 13Cβ, 13C′, and 15NH chemical shifts to identify protein secondary structure and the redox state of cysteine residue. For the analysis of paired chemical shifts of 6 data sets, each of the 20 amino acids has its own 15 two-dimension cluster scattering diagrams. Accordingly, the probabilities for identifying helix and extended structure were calculated by using our scoring matrix. Compared with existing the chemical shift-based methods, it appears to improve the prediction accuracy of secondary structure identification, particularly in the extended structure. In addition, the probability of the given residue to be helix or extended structure is displayed, allows the users to make decisions by themselves. Electronic Supplementary Material The online version of this article (doi:) contains supplementary material, which is available to authorized users. Grant sponsor: National Science Council of ROC; Grant numbers: NSC-94-2323-B006- 001, NSC-93-2212-E-006.  相似文献   

6.
7.
PsiCSI is a highly accurate and automated method of assigning secondary structure from NMR data, which is a useful intermediate step in the determination of tertiary structures. The method combines information from chemical shifts and protein sequence using three layers of neural networks. Training and testing was performed on a suite of 92 proteins (9437 residues) with known secondary and tertiary structure. Using a stringent cross-validation procedure in which the target and homologous proteins were removed from the databases used for training the neural networks, an average 89% Q3 accuracy (per residue) was observed. This is an increase of 6.2% and 5.5% (representing 36% and 33% fewer errors) over methods that use chemical shifts (CSI) or sequence information (Psipred) alone. In addition, PsiCSI improves upon the translation of chemical shift information to secondary structure (Q3 = 87.4%) and is able to use sequence information as an effective substitute for sparse NMR data (Q3 = 86.9% without (13)C shifts and Q3 = 86.8% with only H(alpha) shifts available). Finally, errors made by PsiCSI almost exclusively involve the interchange of helix or strand with coil and not helix with strand (<2.5 occurrences per 10000 residues). The automation, increased accuracy, absence of gross errors, and robustness with regards to sparse data make PsiCSI ideal for high-throughput applications, and should improve the effectiveness of hybrid NMR/de novo structure determination methods. A Web server is available for users to submit data and have the assignment returned.  相似文献   

8.
The combination of the wide availability of protein backbone and side-chain NMR chemical shifts with advances in understanding of their relationship to protein structure makes these parameters useful for the assessment of structural-dynamic protein models. A new chemical shift predictor (PPM) is introduced, which is solely based on physical?Cchemical contributions to the chemical shifts for both the protein backbone and methyl-bearing amino-acid side chains. To explicitly account for the effects of protein dynamics on chemical shifts, PPM was directly refined against 100?ns long molecular dynamics (MD) simulations of 35 proteins with known experimental NMR chemical shifts. It is found that the prediction of methyl-proton chemical shifts by PPM from MD ensembles is improved over other methods, while backbone C??, C??, C??, N, and HN chemical shifts are predicted at an accuracy comparable to the latest generation of chemical shift prediction programs. PPM is particularly suitable for the rapid evaluation of large protein conformational ensembles on their consistency with experimental NMR data and the possible improvement of protein force fields from chemical shifts.  相似文献   

9.
The transient secondary structure and dynamics of an intrinsically unstructured linker domain from the 70 kDa subunit of human replication protein A was investigated using solution state NMR. Stable secondary structure, inferred from large secondary chemical shifts, was observed for a segment of the intrinsically unstructured linker domain when it is attached to an N-terminal protein interaction domain. Results from NMR relaxation experiments showed the rotational diffusion for this segment of the intrinsically unstructured linker domain to be correlated with the N-terminal protein interaction domain. When the N-terminal domain is removed, the stable secondary structure is lost and faster rotational diffusion is observed. The large secondary chemical shifts were used to calculate phi and psi dihedral angles and these dihedral angles were used to build a backbone structural model. Restrained molecular dynamics were performed on this new structure using the chemical shift based dihedral angles and a single NOE distance as restraints. In the resulting family of structures a large, solvent exposed loop was observed for the segment of the intrinsically unstructured linker domain that had large secondary chemical shifts.  相似文献   

10.
Local structure prediction can facilitate ab initio structure prediction, protein threading, and remote homology detection. However, the accuracy of existing methods is limited. In this paper, we propose a knowledge-based prediction method that assigns a measure called the local match rate to each position of an amino acid sequence to estimate the confidence of our method. Empirically, the accuracy of the method correlates positively with the local match rate; therefore, we employ it to predict the local structures of positions with a high local match rate. For positions with a low local match rate, we propose a neural network prediction method. To better utilize the knowledge-based and neural network methods, we design a hybrid prediction method, HYPLOSP (HYbrid method to Protein LOcal Structure Prediction) that combines both methods. To evaluate the performance of the proposed methods, we first perform cross-validation experiments by applying our knowledge-based method, a neural network method, and HYPLOSP to a large dataset of 3,925 protein chains. We test our methods extensively on three different structural alphabets and evaluate their performance by two widely used criteria, Maximum Deviation of backbone torsion Angle (MDA) and Q(N), which is similar to Q(3) in secondary structure prediction. We then compare HYPLOSP with three previous studies using a dataset of 56 new protein chains. HYPLOSP shows promising results in terms of MDA and Q(N) accuracy and demonstrates its alphabet-independent capability.  相似文献   

11.
Successful prediction of the beta-hairpin motif will be helpful for understanding the of the fold recognition. Some algorithms have been proposed for the prediction of beta-hairpin motifs. However, the parameters used by these methods were primarily based on the amino acid sequences. Here, we proposed a novel model for predicting beta-hairpin structure based on the chemical shift. Firstly, we analyzed the statistical distribution of chemical shifts of six nuclei in not beta-hairpin and beta-hairpin motifs. Secondly, we used these chemical shifts as features combined with three algorithms to predict beta-hairpin structure. Finally, we achieved the best prediction, namely sensitivity of 92%, the specificity of 94% with 0.85 of Mathew’s correlation coefficient using quadratic discriminant analysis algorithm, which is clearly superior to the same method for the prediction of beta-hairpin structure from 20 amino acid compositions in the three-fold cross-validation. Our finding showed that the chemical shift is an effective parameter for beta-hairpin prediction, suggesting the quadratic discriminant analysis is a powerful algorithm for the prediction of beta-hairpin.  相似文献   

12.
NMR chemical shifts provide important local structural information for proteins and are key in recently described protein structure generation protocols. We describe a new chemical shift prediction program, SPARTA+, which is based on artificial neural networking. The neural network is trained on a large carefully pruned database, containing 580 proteins for which high-resolution X-ray structures and nearly complete backbone and 13Cβ chemical shifts are available. The neural network is trained to establish quantitative relations between chemical shifts and protein structures, including backbone and side-chain conformation, H-bonding, electric fields and ring-current effects. The trained neural network yields rapid chemical shift prediction for backbone and 13Cβ atoms, with standard deviations of 2.45, 1.09, 0.94, 1.14, 0.25 and 0.49 ppm for δ15N, δ13C’, δ13Cα, δ13Cβ, δ1Hα and δ1HN, respectively, between the SPARTA+ predicted and experimental shifts for a set of eleven validation proteins. These results represent a modest but consistent improvement (2–10%) over the best programs available to date, and appear to be approaching the limit at which empirical approaches can predict chemical shifts.  相似文献   

13.
NMR chemical shift predictions based on empirical methods are nowadays indispensable tools during resonance assignment and 3D structure calculation of proteins. However, owing to the very limited statistical data basis, such methods are still in their infancy in the field of nucleic acids, especially when non-canonical structures and nucleic acid complexes are considered. Here, we present an ab initio approach for predicting proton chemical shifts of arbitrary nucleic acid structures based on state-of-the-art fragment-based quantum chemical calculations. We tested our prediction method on a diverse set of nucleic acid structures including double-stranded DNA, hairpins, DNA/protein complexes and chemically-modified DNA. Overall, our quantum chemical calculations yield highly/very accurate predictions with mean absolute deviations of 0.3–0.6 ppm and correlation coefficients (r2) usually above 0.9. This will allow for identifying misassignments and validating 3D structures. Furthermore, our calculations reveal that chemical shifts of protons involved in hydrogen bonding are predicted significantly less accurately. This is in part caused by insufficient inclusion of solvation effects. However, it also points toward shortcomings of current force fields used for structure determination of nucleic acids. Our quantum chemical calculations could therefore provide input for force field optimization.  相似文献   

14.
Chao Fang  Yi Shang  Dong Xu 《Proteins》2018,86(5):592-598
Protein secondary structure prediction can provide important information for protein 3D structure prediction and protein functions. Deep learning offers a new opportunity to significantly improve prediction accuracy. In this article, a new deep neural network architecture, named the Deep inception‐inside‐inception (Deep3I) network, is proposed for protein secondary structure prediction and implemented as a software tool MUFOLD‐SS. The input to MUFOLD‐SS is a carefully designed feature matrix corresponding to the primary amino acid sequence of a protein, which consists of a rich set of information derived from individual amino acid, as well as the context of the protein sequence. Specifically, the feature matrix is a composition of physio‐chemical properties of amino acids, PSI‐BLAST profile, and HHBlits profile. MUFOLD‐SS is composed of a sequence of nested inception modules and maps the input matrix to either eight states or three states of secondary structures. The architecture of MUFOLD‐SS enables effective processing of local and global interactions between amino acids in making accurate prediction. In extensive experiments on multiple datasets, MUFOLD‐SS outperformed the best existing methods and other deep neural networks significantly. MUFold‐SS can be downloaded from http://dslsrv8.cs.missouri.edu/~cf797/MUFoldSS/download.html .  相似文献   

15.
Secondary structure plays an important role in determining the function of noncoding RNAs. Hence, identifying RNA secondary structures is of great value to research. Computational prediction is a mainstream approach for predicting RNA secondary structure. Unfortunately, even though new methods have been proposed over the past 40 years, the performance of computational prediction methods has stagnated in the last decade. Recently, with the increasing availability of RNA structure data, new methods based on machine learning (ML) technologies, especially deep learning, have alleviated the issue. In this review, we provide a comprehensive overview of RNA secondary structure prediction methods based on ML technologies and a tabularized summary of the most important methods in this field. The current pending challenges in the field of RNA secondary structure prediction and future trends are also discussed.  相似文献   

16.
目的 长链非编码RNA在遗传、代谢和基因表达调控等方面发挥着重要作用。然而,传统的实验方法解析RNA的三级结构耗时长、费用高且操作要求高。此外,通过计算方法来预测RNA的三级结构在近十年来无突破性进展。因此,需要提出新的预测算法来准确的预测RNA的三级结构。所以,本文发展可以用于提高RNA三级结构预测准确性的碱基关联图预测方法。方法 为了利用RNA理化特征信息,本文应用多层全卷积神经网络和循环神经网络的深度学习算法来预测RNA碱基间的接触概率,并通过注意力机制处理RNA序列中碱基间相互依赖的特征。结果 通过多层神经网络与注意力机制结合,本文方法能够有效得到RNA特征值中局部和全局的信息,提高了模型的鲁棒性和泛化能力。检验计算表明,所提出模型对序列长度L的4种标准(L/10、L/5、L/2、L)碱基关联图的预测准确率分别达到0.84、0.82、0.82和0.75。结论 基于注意力机制的深度学习预测算法能够提高RNA碱基关联图预测的准确率,从而帮助RNA三级结构的预测。  相似文献   

17.
A 4D approach for protein 1H chemical shift prediction was explored. The 4th dimension is the molecular flexibility, mapped using molecular dynamics simulations. The chemical shifts were predicted with a principal component model based on atom coordinates from a database of 40 protein structures. When compared to the corresponding non-dynamic (3D) model, the 4th dimension improved prediction by 6–7%. The prediction method achieved RMS errors of 0.29 and 0.50 ppm for Hα and HN shifts, respectively. However, for individual proteins the RMS errors were 0.17–0.34 and 0.34–0.65 ppm for the Hα and HN shifts, respectively. X-ray structures gave better predictions than the corresponding NMR structures, indicating that chemical shifts contain invaluable information about local structures. The 1H chemical shift prediction tool 4DSPOT is available from .  相似文献   

18.
张斌  尹京苑  薛丹 《生物信息学》2011,9(3):224-228,234
蛋白质二级结构对于研究其功能具有重要作用。采用主成分分析方法对氨基酸的基本物化属性及其二级结构倾向性进行降维降噪处理,使用径向基神经网络对蛋白质二级结构进行预测。主成分分析使得之前 20 ×12 矩阵变为 20 ×4 矩阵,极大地减少了神经网络输入端的维数。在仿真过程中,当窗口大小为 21,扩展函数为 7 时,预测精确度达到了 71. 81%。实验结果表明 RBF 神经网络可以有效的用于蛋白质二级结构的预测。  相似文献   

19.
We present an empirical method for identification of distinct structural motifs in proteins on the basis of experimentally determined backbone and 13Cβ chemical shifts. Elements identified include the N-terminal and C-terminal helix capping motifs and five types of β-turns: I, II, I′, II′ and VIII. Using a database of proteins of known structure, the NMR chemical shifts, together with the PDB-extracted amino acid preference of the helix capping and β-turn motifs are used as input data for training an artificial neural network algorithm, which outputs the statistical probability of finding each motif at any given position in the protein. The trained neural networks, contained in the MICS (motif identification from chemical shifts) program, also provide a confidence level for each of their predictions, and values ranging from ca 0.7–0.9 for the Matthews correlation coefficient of its predictions far exceed those attainable by sequence analysis. MICS is anticipated to be useful both in the conventional NMR structure determination process and for enhancing on-going efforts to determine protein structures solely on the basis of chemical shift information, where it can aid in identifying protein database fragments suitable for use in building such structures.  相似文献   

20.
We present a new method for predicting the secondary structure of globular proteins based on non-linear neural network models. Network models learn from existing protein structures how to predict the secondary structure of local sequences of amino acids. The average success rate of our method on a testing set of proteins non-homologous with the corresponding training set was 64.3% on three types of secondary structure (alpha-helix, beta-sheet, and coil), with correlation coefficients of C alpha = 0.41, C beta = 0.31 and Ccoil = 0.41. These quality indices are all higher than those of previous methods. The prediction accuracy for the first 25 residues of the N-terminal sequence was significantly better. We conclude from computational experiments on real and artificial structures that no method based solely on local information in the protein sequence is likely to produce significantly better results for non-homologous proteins. The performance of our method of homologous proteins is much better than for non-homologous proteins, but is not as good as simply assuming that homologous sequences have identical structures.  相似文献   

设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号