首页 | 本学科首页   官方微博 | 高级检索  
相似文献
 共查询到20条相似文献,搜索用时 953 毫秒
1.
The ab initio folding problem can be divided into two sequential tasks of approximately equal computational complexity: the generation of native-like backbone folds and the positioning of side chains upon these backbones. The prediction of side-chain conformation in this context is challenging, because at best only the near-native global fold of the protein is known. To test the effect of displacements in the protein backbones on side-chain prediction for folds generated ab initio, sets of near-native backbones (≤ 4 Å Cα RMS error) for four small proteins were generated by two methods. The steric environment surrounding each residue was probed by placing the side chains in the native conformation on each of these decoys, followed by torsion-space optimization to remove steric clashes on a rigid backbone. We observe that on average 40% of the χ1 angles were displaced by 40° or more, effectively setting the limits in accuracy for side-chain modeling under these conditions. Three different algorithms were subsequently used for prediction of side-chain conformation. The average prediction accuracy for the three methods was remarkably similar: 49% to 51% of the χ1 angles were predicted correctly overall (33% to 36% of the χ1+2 angles). Interestingly, when the inter-side-chain interactions were disregarded, the mean accuracy increased. A consensus approach is described, in which side-chain conformations are defined based on the most frequently predicted χ angles for a given method upon each set of near-native backbones. We find that consensus modeling, which de facto includes backbone flexibility, improves side-chain prediction: χ1 accuracy improved to 51–54% (36–42% of χ1+2). Implications of a consensus method for ab initio protein structure prediction are discussed. Proteins 33:204–217, 1998. © 1998 Wiley-Liss, Inc.  相似文献   

2.
In recent years, it has been repeatedly demonstrated that the coordinates of the main-chain atoms alone are sufficient to determine the side-chain conformations of buried residues of compact proteins. Given a perfect backbone, the side-chain packing method can predict the side-chain conformations to an accuracy as high as 1.2 Å RMS deviation (RMSD) with greater than 80% of the χ angles correct. However, similarly rigorous studies have not been conducted to determine how well these apply, if at all, to the more important problem of homology modeling per se. Specifically, if the available backbone is imperfect, as expected for practical application of homology modeling, can packing constraints alone achieve sufficiently accurate predictions to be useful? Here, by systematically applying such methods to the pairwise modeling of two repressor and two cro proteins from the closely related bacteriophages 434 and P22, we find that when the backbone RMSD is 0.8 Å, the prediction on buried side chain is accurate with an RMS error of 1.8 Å and approximately 70% of the χ angles correctly predicted. When the backbone RMSD is larger, in the range of 1.6–1.8 Å, the prediction quality is still significantly better than random, with RMS error at 2.2 Å on the buried side chains and 60% accuracy on χ angles. Together these results suggest the following rules-of-thumb for homology modeling of buried side chains. When the sequence identity between the modeled sequence and the template sequence is >50% (or, equivalently, the expected backbone RMSD is <1 Å), side-chain packing methods work well. When sequence identity is between 30–50%, reflecting a backbone RMS error of 1–2 Å, it is still valid to use side-chain packing methods to predict the buried residues, albeit with care. When sequence identity is below 30% (or backbone RMS error greater than 2 Å), the backbone constraint alone is unlikely to produce useful models. Other methods, such as those involving the use of database fragments to reconstruct a template backbone, may be necessary as a complementary guide for modeling.  相似文献   

3.
Rohl CA  Strauss CE  Chivian D  Baker D 《Proteins》2004,55(3):656-677
A major limitation of current comparative modeling methods is the accuracy with which regions that are structurally divergent from homologues of known structure can be modeled. Because structural differences between homologous proteins are responsible for variations in protein function and specificity, the ability to model these differences has important functional consequences. Although existing methods can provide reasonably accurate models of short loop regions, modeling longer structurally divergent regions is an unsolved problem. Here we describe a method based on the de novo structure prediction algorithm, Rosetta, for predicting conformations of structurally divergent regions in comparative models. Initial conformations for short segments are selected from the protein structure database, whereas longer segments are built up by using three- and nine-residue fragments drawn from the database and combined by using the Rosetta algorithm. A gap closure term in the potential in combination with modified Newton's method for gradient descent minimization is used to ensure continuity of the peptide backbone. Conformations of variable regions are refined in the context of a fixed template structure using Monte Carlo minimization together with rapid repacking of side-chains to iteratively optimize backbone torsion angles and side-chain rotamers. For short loops, mean accuracies of 0.69, 1.45, and 3.62 A are obtained for 4, 8, and 12 residue loops, respectively. In addition, the method can provide reasonable models of conformations of longer protein segments: predicted conformations of 3A root-mean-square deviation or better were obtained for 5 of 10 examples of segments ranging from 13 to 34 residues. In combination with a sequence alignment algorithm, this method generates complete, ungapped models of protein structures, including regions both similar to and divergent from a homologous structure. This combined method was used to make predictions for 28 protein domains in the Critical Assessment of Protein Structure 4 (CASP 4) and 59 domains in CASP 5, where the method ranked highly among comparative modeling and fold recognition methods. Model accuracy in these blind predictions is dominated by alignment quality, but in the context of accurate alignments, long protein segments can be accurately modeled. Notably, the method correctly predicted the local structure of a 39-residue insertion into a TIM barrel in CASP 5 target T0186.  相似文献   

4.
L Holm  C Sander 《Proteins》1992,14(2):213-223
An unknown protein structure can be predicted with fair accuracy once an evolutionary connection at the sequence level has been made to a protein of known 3-D structure. In model building by homology, one typically starts with a backbone framework, rebuilds new loop regions, and replaces nonconserved side chains. Here, we use an extremely efficient Monte Carlo algorithm in rotamer space with simulated annealing and simple potential energy functions to optimize the packing of side chains on given backbone models. Optimized models are generated within minutes on a workstation, with reasonable accuracy (average of 81% side chain chi 1 dihedral angles correct in the cores of proteins determined at better than 2.5 A resolution). As expected, the quality of the models decreases with decreasing accuracy of backbone coordinates. If the back-bone was taken from a homologous rather than the same protein, about 70% side chain chi 1 angles were modeled correctly in the core in a case of strong homology and about 60% in a case of medium homology. The algorithm can be used in automated, fast, and reproducible model building by homology.  相似文献   

5.
6.
Metropolis Monte Carlo (MMC) loop refinement has been performed on the three extracellular loops (ECLs) of rhodopsin and opsin-based homology models of the thyroid-stimulating hormone receptor transmembrane domain, a class A type G protein-coupled receptor. The Monte Carlo sampling technique, employing torsion angles of amino acid side chains and local moves for the six consecutive backbone torsion angles, has previously reproduced the conformation of several loops with known crystal structures with accuracy consistently less than 2?Å. A grid-based potential map, which includes van der Waals, electrostatics, hydrophobic as well as hydrogen-bond potentials for bulk protein environment and the solvation effect, has been used to significantly reduce the computational cost of energy evaluation. A modified sigmoidal distance-dependent dielectric function has been implemented in conjunction with the desolvation and hydrogen-bonding terms. A long high-temperature simulation with 2?kcal/mol repulsion potential resulted in extensive sampling of the conformational space. The slow annealing leading to the low-energy structures predicted secondary structure by the MMC technique. Molecular docking with the reported agonist reproduced the binding site within 1.5?Å. Virtual screening performed on the three lowest structures showed that the ligand-binding mode in the inter-helical region is dependent on the ECL conformations.  相似文献   

7.
Modeling of loops in protein structures   总被引:27,自引:0,他引:27       下载免费PDF全文
Comparative protein structure prediction is limited mostly by the errors in alignment and loop modeling. We describe here a new automated modeling technique that significantly improves the accuracy of loop predictions in protein structures. The positions of all nonhydrogen atoms of the loop are optimized in a fixed environment with respect to a pseudo energy function. The energy is a sum of many spatial restraints that include the bond length, bond angle, and improper dihedral angle terms from the CHARMM-22 force field, statistical preferences for the main-chain and side-chain dihedral angles, and statistical preferences for nonbonded atomic contacts that depend on the two atom types, their distance through space, and separation in sequence. The energy function is optimized with the method of conjugate gradients combined with molecular dynamics and simulated annealing. Typically, the predicted loop conformation corresponds to the lowest energy conformation among 500 independent optimizations. Predictions were made for 40 loops of known structure at each length from 1 to 14 residues. The accuracy of loop predictions is evaluated as a function of thoroughness of conformational sampling, loop length, and structural properties of native loops. When accuracy is measured by local superposition of the model on the native loop, 100, 90, and 30% of 4-, 8-, and 12-residue loop predictions, respectively, had <2 A RMSD error for the mainchain N, C(alpha), C, and O atoms; the average accuracies were 0.59 +/- 0.05, 1.16 +/- 0.10, and 2.61 +/- 0.16 A, respectively. To simulate real comparative modeling problems, the method was also evaluated by predicting loops of known structure in only approximately correct environments with errors typical of comparative modeling without misalignment. When the RMSD distortion of the main-chain stem atoms is 2.5 A, the average loop prediction error increased by 180, 25, and 3% for 4-, 8-, and 12-residue loops, respectively. The accuracy of the lowest energy prediction for a given loop can be estimated from the structural variability among a number of low energy predictions. The relative value of the present method is gauged by (1) comparing it with one of the most successful previously described methods, and (2) describing its accuracy in recent blind predictions of protein structure. Finally, it is shown that the average accuracy of prediction is limited primarily by the accuracy of the energy function rather than by the extent of conformational sampling.  相似文献   

8.
rap-1A, an anti-oncogene-encoded protein, is aras-p21-like protein whose sequence is over 80% homologous to p21 and which interacts with the same intracellular target proteins and is activated by the same mechanisms as p21, e.g., by binding GTP in place of GDP. Both interact with effector proteins in the same region, involving residues 32–47. However, activated rap-1A blocks the mitogenic signal transducing effects of p21. Optimal sequence alignment of p21 and rap-1A shows two insertions of rap-1A atras positions 120 and 138. We have constructed the three-dimensional structure of rap-1A bound to GTP by using the energy-minimized three-dimensional structure ofras-p21 as the basis for the modeling using a stepwise procedure in which identical and homologous amino acid residues in rap-1A are assumed to adopt the same conformation as the corresponding residues in p21. Side-chain conformations for homologous and nonhomologous residues are generated in conformations that are as close as possible to those of the corresponding side chains in p21. The entire structure has been subjected to a nested series of energy minimizations. The final predicted structure has an overall backbone deviation of 0.7 å from that ofras-p21. The effector binding domains from residues 32–47 are identical in both proteins (except for different side chains of different residues at position 45). A major difference occurs in the insertion region at residue 120. This region is in the middle of another effector loop of the p21 protein involving residues 115–126. Differences in sequence and structure in this region may contribute to the differences in cellular functions of these two proteins.  相似文献   

9.
10.
MOTIVATION: Side-chain positioning is a central component of homology modeling and protein design. In a common formulation of the problem, the backbone is fixed, side-chain conformations come from a rotamer library, and a pairwise energy function is optimized. It is NP-complete to find even a reasonable approximate solution to this problem. We seek to put this hardness result into practical context. RESULTS: We present an integer linear programming (ILP) formulation of side-chain positioning that allows us to tackle large problem sizes. We relax the integrality constraint to give a polynomial-time linear programming (LP) heuristic. We apply LP to position side chains on native and homologous backbones and to choose side chains for protein design. Surprisingly, when positioning side chains on native and homologous backbones, optimal solutions using a simple, biologically relevant energy function can usually be found using LP. On the other hand, the design problem often cannot be solved using LP directly; however, optimal solutions for large instances can still be found using the computationally more expensive ILP procedure. While different energy functions also affect the difficulty of the problem, the LP/ILP approach is able to find optimal solutions. Our analysis is the first large-scale demonstration that LP-based approaches are highly effective in finding optimal (and successive near-optimal) solutions for the side-chain positioning problem.  相似文献   

11.
Improved side-chain modeling for protein-protein docking   总被引:1,自引:0,他引:1  
Success in high-resolution protein-protein docking requires accurate modeling of side-chain conformations at the interface. Most current methods either leave side chains fixed in the conformations observed in the unbound protein structures or allow the side chains to sample a set of discrete rotamer conformations. Here we describe a rapid and efficient method for sampling off-rotamer side-chain conformations by torsion space minimization during protein-protein docking starting from discrete rotamer libraries supplemented with side-chain conformations taken from the unbound structures, and show that the new method improves side-chain modeling and increases the energetic discrimination between good and bad models. Analysis of the distribution of side-chain interaction energies within and between the two protein partners shows that the new method leads to more native-like distributions of interaction energies and that the neglect of side-chain entropy produces a small but measurable increase in the number of residues whose interaction energy cannot compensate for the entropic cost of side-chain freezing at the interface. The power of the method is highlighted by a number of predictions of unprecedented accuracy in the recent CAPRI (Critical Assessment of PRedicted Interactions) blind test of protein-protein docking methods.  相似文献   

12.
Side-chain modeling with an optimized scoring function   总被引:1,自引:0,他引:1       下载免费PDF全文
Modeling side-chain conformations on a fixed protein backbone has a wide application in structure prediction and molecular design. Each effort in this field requires decisions about a rotamer set, scoring function, and search strategy. We have developed a new and simple scoring function, which operates on side-chain rotamers and consists of the following energy terms: contact surface, volume overlap, backbone dependency, electrostatic interactions, and desolvation energy. The weights of these energy terms were optimized to achieve the minimal average root mean square (rms) deviation between the lowest energy rotamer and real side-chain conformation on a training set of high-resolution protein structures. In the course of optimization, for every residue, its side chain was replaced by varying rotamers, whereas conformations for all other residues were kept as they appeared in the crystal structure. We obtained prediction accuracy of 90.4% for chi(1), 78.3% for chi(1 + 2), and 1.18 A overall rms deviation. Furthermore, the derived scoring function combined with a Monte Carlo search algorithm was used to place all side chains onto a protein backbone simultaneously. The average prediction accuracy was 87.9% for chi(1), 73.2% for chi(1 + 2), and 1.34 A rms deviation for 30 protein structures. Our approach was compared with available side-chain construction methods and showed improvement over the best among them: 4.4% for chi(1), 4.7% for chi(1 + 2), and 0.21 A for rms deviation. We hypothesize that the scoring function instead of the search strategy is the main obstacle in side-chain modeling. Additionally, we show that a more detailed rotamer library is expected to increase chi(1 + 2) prediction accuracy but may have little effect on chi(1) prediction accuracy.  相似文献   

13.
Modeling protein loops using a phi i + 1, psi i dimer database.   总被引:1,自引:1,他引:0       下载免费PDF全文
We present an automated method for modeling backbones of protein loops. The method samples a database of phi i + 1 and psi i angles constructed from a nonredundant version of the Protein Data Bank (PDB). The dihedral angles phi i + 1 and psi i completely define the backbone conformation of a dimer when standard bond lengths, bond angles, and a trans planar peptide configuration are used. For the 400 possible dimers resulting from 20 natural amino acids, a list of allowed phi i + 1, psi i pairs for each dimer is created by pooling all such pairs from the loop segments of each protein in the nonredundant version of the PDB. Starting from the N-terminus of the loop sequence, conformations are generated by assigning randomly selected pairs of phi i + 1, psi i for each dimer from the respective pool using standard bond lengths, bond angles, and a trans peptide configuration. We use this database to simulate protein loops of lengths varying from 5 to 11 amino acids in five proteins of known three-dimensional structures. Typically, 10,000-50,000 models are simulated for each protein loop and are evaluated for stereochemical consistency. Depending on the length and sequence of a given loop, 50-80% of the models generated have no stereochemical strain in the backbone atoms. We demonstrate that, when simulated loops are extended to include flanking residues from homologous segments, only very few loops from an ensemble of sterically allowed conformations orient the flanking segments consistent with the protein topology. The presence of near-native backbone conformations for loops from five different proteins suggests the completeness of the dimeric database for use in modeling loops of homologous proteins. Here, we take advantage of this observation to design a method that filters near-native loop conformations from an ensemble of sterically allowed conformations. We demonstrate that our method eliminates the need for a loop-closure algorithm and hence allows for the use of topological constraints of the homologous proteins or disulfide constraints to filter near-native loop conformations.  相似文献   

14.
rap-1A, an anti-oncogene-encoded protein, is aras-p21-like protein whose sequence is over 80% homologous to p21 and which interacts with the same intracellular target proteins and is activated by the same mechanisms as p21, e.g., by binding GTP in place of GDP. Both interact with effector proteins in the same region, involving residues 32–47. However, activated rap-1A blocks the mitogenic signal transducing effects of p21. Optimal sequence alignment of p21 and rap-1A shows two insertions of rap-1A atras positions 120 and 138. We have constructed the three-dimensional structure of rap-1A bound to GTP by using the energy-minimized three-dimensional structure ofras-p21 as the basis for the modeling using a stepwise procedure in which identical and homologous amino acid residues in rap-1A are assumed to adopt the same conformation as the corresponding residues in p21. Side-chain conformations for homologous and nonhomologous residues are generated in conformations that are as close as possible to those of the corresponding side chains in p21. The entire structure has been subjected to a nested series of energy minimizations. The final predicted structure has an overall backbone deviation of 0.7 å from that ofras-p21. The effector binding domains from residues 32–47 are identical in both proteins (except for different side chains of different residues at position 45). A major difference occurs in the insertion region at residue 120. This region is in the middle of another effector loop of the p21 protein involving residues 115–126. Differences in sequence and structure in this region may contribute to the differences in cellular functions of these two proteins.  相似文献   

15.
The three-dimensional structures of leucine-rich repeat (LRR)-containing proteins from five different families were previously predicted based on the crystal structure of the ribonuclease inhibitor, using an approach that combined homology-based modeling, structure-based sequence alignment of LRRs, and several rational assumptions. The structural models have been produced based on very limited sequence similarity, which, in general, cannot yield trustworthy predictions. Recently, the protein structures from three of these five families have been determined. In this report we estimate the quality of the modeling approach by comparing the models with the experimentally determined structures. The comparison suggests that the general architecture, curvature, "interior/exterior" orientations of side chains, and backbone conformation of the LRR structures can be predicted correctly. On the other hand, the analysis revealed that, in some cases, it is difficult to predict correctly the twist of the overall super-helical structure. Taking into consideration the conclusions from these comparisons, we identified a new family of bacterial LRR proteins and present its structural model. The reliability of the LRR protein modeling suggests that it would be informative to apply similar modeling approaches to other classes of solenoid proteins.  相似文献   

16.
Achieving atomic-level accuracy in comparative protein models is limited by our ability to refine the initial, homolog-derived model closer to the native state. Despite considerable effort, progress in developing a generalized refinement method has been limited. In contrast, methods have been described that can accurately reconstruct loop conformations in native protein structures. We hypothesize that loop refinement in homology models is much more difficult than loop reconstruction in crystal structures, in part, because side-chain, backbone, and other structural inaccuracies surrounding the loop create a challenging sampling problem; the loop cannot be refined without simultaneously refining adjacent portions. In this work, we single out one sampling issue in an artificial but useful test set and examine how loop refinement accuracy is affected by errors in surrounding side-chains. In 80 high-resolution crystal structures, we first perturbed 6-12 residue loops away from the crystal conformation, and placed all protein side chains in non-native but low energy conformations. Even these relatively small perturbations in the surroundings made the loop prediction problem much more challenging. Using a previously published loop prediction method, median backbone (N-Calpha-C-O) RMSD's for groups of 6, 8, 10, and 12 residue loops are 0.3/0.6/0.4/0.6 A, respectively, on native structures and increase to 1.1/2.2/1.5/2.3 A on the perturbed cases. We then augmented our previous loop prediction method to simultaneously optimize the rotamer states of side chains surrounding the loop. Our results show that this augmented loop prediction method can recover the native state in many perturbed structures where the previous method failed; the median RMSD's for the 6, 8, 10, and 12 residue perturbed loops improve to 0.4/0.8/1.1/1.2 A. Finally, we highlight three comparative models from blind tests, in which our new method predicted loops closer to the native conformation than first modeled using the homolog template, a task generally understood to be difficult. Although many challenges remain in refining full comparative models to high accuracy, this work offers a methodical step toward that goal.  相似文献   

17.
Rebuilding flavodoxin from C alpha coordinates: a test study   总被引:4,自引:0,他引:4  
L S Reid  J M Thornton 《Proteins》1989,5(2):170-182
The tertiary structure of flavodoxin has been model built from only the X-ray crystallographic alpha-carbon coordinates. Main-chain atoms were generated from a dictionary of backbone structures. Side-chain conformations were initially set according to observed statistical distributions, clashes were resolved with reference to other knowledge-based parameters, and finally, energy minimization was applied. The RMSD of the model was 1.7 A across all atoms to the native structure. Regular secondary structural elements were modeled more accurately than other regions. About 40% of the chi 1 torsional angles were modeled correctly. Packing of side chains in the core was energetically stable but diverged significantly from the native structure in some regions. The modeling of protein structures is increasing in popularity but relatively few checks have been applied to determine the accuracy of the approach. In this work a variety of parameters have been examined. It was found that close contacts, and hydrogen-bonding patterns could identify poorly packed residues. These tests, however, did not indicate which residues had a conformation different from the native structure or how to move such residues to bring them into agreement. To assist in the modeling of interacting side chains a database of known interactions has been prepared.  相似文献   

18.
Wu S  Zhang Y 《PloS one》2008,3(10):e3400
We developed a composite machine-learning based algorithm, called ANGLOR, to predict real-value protein backbone torsion angles from amino acid sequences. The input features of ANGLOR include sequence profiles, predicted secondary structure and solvent accessibility. In a large-scale benchmarking test, the mean absolute error (MAE) of the phi/psi prediction is 28 degrees/46 degrees , which is approximately 10% lower than that generated by software in literature. The prediction is statistically different from a random predictor (or a purely secondary-structure-based predictor) with p-value <1.0 x 10(-300) (or <1.0 x 10(-148)) by Wilcoxon signed rank test. For some residues (ILE, LEU, PRO and VAL) and especially the residues in helix and buried regions, the MAE of phi angles is much smaller (10-20 degrees ) than that in other environments. Thus, although the average accuracy of the ANGLOR prediction is still low, the portion of the accurately predicted dihedral angles may be useful in assisting protein fold recognition and ab initio 3D structure modeling.  相似文献   

19.
20.
A theoretical and computational approach to ab initio structure prediction for polypeptides in water is described and applied to selected amino acid sequences for testing and preliminary validation. The method builds systematically on the extensive efforts applied to parameterization of molecular dynamics (MD) force fields, employs an empirically well-validated continuum dielectric model for solvation, and an eminently parallelizable approach to conformational search. The effective free energy of polypeptide chains is estimated from AMBER united atom potential functions, with internal degrees of freedom for both backbone and amino acid side chains explicitly treated. The hydration free energy of each structure is determined using the Generalized Born/Solvent Accessibility (GBSA) method, modified and reparameterized to include atom types consistent with the AMBER force field. The conformational search procedure employs a multiple copy, Monte Carlo simulated annealing (MCSA) protocol in full torsion angle space, applied iteratively on sets of structures of progressively lower free energy until a prediction of a structure with lowest effective free energy is obtained. Calibration tests for the effective energy function and search algorithm are performed on the alanine dipeptide, selected protein crystal structures, and united atom decoys on barnase, crambin, and six examples from the Rosetta set. Specific demonstration cases of the method are provided for the 8-mer sequence of Ala residues, a 12-residue peptide with longer side chains QLLKKLLQQLKQ, a de novo designed 16 residue peptide of sequence (AAQAA)3Y, a 15-residue sequence with a beta sheet motif, GEWTWDATKTFTVTE, and a 36 residue small protein, Villin headpiece. The Ala 8-mer readily formed an alpha-helix. An alpha-helix structure was predicted for the 16-mer, consistent with observed results from IR and CD spectroscopy and with the pattern in psi/straight phi angles of known protein structures. The predicted structure for the 12-mer, composed of a mix of helix and less regular elements of secondary structure, lies 2.65 A RMS from the observed crystal structure. Structure prediction for the 8-mer beta-motif resulted in form 4.50 A RMS from the crystal geometry. For Villin, the predicted native form is very close to the crystal structure, RMS values of 3.5 A (including sidechains), and 1.01 A (main chain only). The methodology permits a detailed analysis of the molecular forces which dominate various segments of the predicted folding trajectory. Analysis of the results in terms of internal torsional, electrostatic and van der Waals and the electrostatic and non-electrostatic contributions to hydration, including the hydrophobic effect, is presented.  相似文献   

设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号