首页 | 本学科首页   官方微博 | 高级检索  
相似文献
 共查询到20条相似文献,搜索用时 15 毫秒
1.
Protein sequence world is considerably larger than structure world. In consequence, numerous non-related sequences may adopt similar 3D folds and different kinds of amino acids may thus be found in similar 3D structures. By grouping together the 20 amino acids into a smaller number of representative residues with similar features, sequence world simplification may be achieved. This clustering hence defines a reduced amino acid alphabet (reduced AAA). Numerous works have shown that protein 3D structures are composed of a limited number of building blocks, defining a structural alphabet. We previously identified such an alphabet composed of 16 representative structural motifs (5-residues length) called Protein Blocks (PBs). This alphabet permits to translate the structure (3D) in sequence of PBs (1D). Based on these two concepts, reduced AAA and PBs, we analyzed the distributions of the different kinds of amino acids and their equivalences in the structural context. Different reduced sets were considered. Recurrent amino acid associations were found in all the local structures while other were specific of some local structures (PBs) (e.g Cysteine, Histidine, Threonine and Serine for the alpha-helix Ncap). Some similar associations are found in other reduced AAAs, e.g Ile with Val, or hydrophobic aromatic residues Trp with Phe and Tyr. We put into evidence interesting alternative associations. This highlights the dependence on the information considered (sequence or structure). This approach, equivalent to a substitution matrix, could be useful for designing protein sequence with different features (for instance adaptation to environment) while preserving mainly the 3D fold.  相似文献   

2.
3.
A statistical analysis of the PDB structures has led us to define a new set of small 3D structural prototypes called Protein Blocks (PBs). This structural alphabet includes 16 PBs, each one is defined by the (phi, psi) dihedral angles of 5 consecutive residues. The amino acid distributions observed in sequence windows encompassing these PBs are used to predict by a Bayesian approach the local 3D structure of proteins from the sole knowledge of their sequences. LocPred is a software which allows the users to submit a protein sequence and performs a prediction in terms of PBs. The prediction results are given both textually and graphically.  相似文献   

4.
The structural annotation of proteins with no detectable homologs of known 3D structure identified using sequence‐search methods is a major challenge today. We propose an original method that computes the conditional probabilities for the amino‐acid sequence of a protein to fit to known protein 3D structures using a structural alphabet, known as “Protein Blocks” (PBs). PBs constitute a library of 16 local structural prototypes that approximate every part of protein backbone structures. It is used to encode 3D protein structures into 1D PB sequences and to capture sequence to structure relationships. Our method relies on amino acid occurrence matrices, one for each PB, to score global and local threading of query amino acid sequences to protein folds encoded into PB sequences. It does not use any information from residue contacts or sequence‐search methods or explicit incorporation of hydrophobic effect. The performance of the method was assessed with independent test datasets derived from SCOP 1.75A. With a Z‐score cutoff that achieved 95% specificity (i.e., less than 5% false positives), global and local threading showed sensitivity of 64.1% and 34.2%, respectively. We further tested its performance on 57 difficult CASP10 targets that had no known homologs in PDB: 38 compatible templates were identified by our approach and 66% of these hits yielded correctly predicted structures. This method scales‐up well and offers promising perspectives for structural annotations at genomic level. It has been implemented in the form of a web‐server that is freely available at http://www.bo‐protscience.fr/forsa .  相似文献   

5.
We present a comprehensive evaluation of a new structure mining method called PB-ALIGN. It is based on the encoding of protein structure as 1D sequence of a combination of 16 short structural motifs or protein blocks (PBs). PBs are short motifs capable of representing most of the local structural features of a protein backbone. Using derived PB substitution matrix and simple dynamic programming algorithm, PB sequences are aligned the same way amino acid sequences to yield structure alignment. PBs are short motifs capable of representing most of the local structural features of a protein backbone. Alignment of these local features as sequence of symbols enables fast detection of structural similarities between two proteins. Ability of the method to characterize and align regions beyond regular secondary structures, for example, N and C caps of helix and loops connecting regular structures, puts it a step ahead of existing methods, which strongly rely on secondary structure elements. PB-ALIGN achieved efficiency of 85% in extracting true fold from a large database of 7259 SCOP domains and was successful in 82% cases to identify true super-family members. On comparison to 13 existing structure comparison/mining methods, PB-ALIGN emerged as the best on general ability test dataset and was at par with methods like YAKUSA and CE on nontrivial test dataset. Furthermore, the proposed method performed well when compared to flexible structure alignment method like FATCAT and outperforms in processing speed (less than 45 s per database scan). This work also establishes a reliable cut-off value for the demarcation of similar folds. It finally shows that global alignment scores of unrelated structures using PBs follow an extreme value distribution. PB-ALIGN is freely available on web server called Protein Block Expert (PBE) at http://bioinformatics.univ-reunion.fr/PBE/.  相似文献   

6.
Three-dimensional protein structures can be described with a library of 3D fragments that define a structural alphabet. We have previously proposed such an alphabet, composed of 16 patterns of five consecutive amino acids, called Protein Blocks (PBs). These PBs have been used to describe protein backbones and to predict local structures from protein sequences. The Q16 prediction rate reaches 40.7% with an optimization procedure. This article examines two aspects of PBs. First, we determine the effect of the enlargement of databanks on their definition. The results show that the geometrical features of the different PBs are preserved (local RMSD value equal to 0.41 A on average) and sequence-structure specificities reinforced when databanks are enlarged. Second, we improve the methods for optimizing PB predictions from sequences, revisiting the optimization procedure and exploring different local prediction strategies. Use of a statistical optimization procedure for the sequence-local structure relation improves prediction accuracy by 8% (Q16 = 48.7%). Better recognition of repetitive structures occurs without losing the prediction efficiency of the other local folds. Adding secondary structure prediction improved the accuracy of Q16 by only 1%. An entropy index (Neq), strongly related to the RMSD value of the difference between predicted PBs and true local structures, is proposed to estimate prediction quality. The Neq is linearly correlated with the Q16 prediction rate distributions, computed for a large set of proteins. An "expected" prediction rate QE16 is deduced with a mean error of 5%.  相似文献   

7.
Analysis of protein structures based on backbone structural patterns known as structural alphabets have been shown to be very useful. Among them, a set of 16 pentapeptide structural motifs known as protein blocks (PBs) has been identified and upon which backbone model of most protein structures can be built. PBs allows simplification of 3D space onto 1D space in the form of sequence of PBs. Here, for the first time, substitution probabilities of PBs in a large number of aligned homologous protein structures have been studied and are expressed as a simplified 16 x 16 substitution matrix. The matrix was validated by benchmarking how well it can align sequences of PBs rather like amino acid alignment to identify structurally equivalent regions in closely or distantly related proteins using dynamic programming approach. The alignment results obtained are very comparable to well established structure comparison methods like DALI and STAMP. Other interesting applications of the matrix have been investigated. We first show that, in variable regions between two superimposed homologous proteins, one can distinguish between local conformational differences and rigid-body displacement of a conserved motif by comparing the PBs and their substitution scores. Second, we demonstrate, with the example of aspartic proteinases, that PBs can be efficiently used to detect the lobe/domain flexibility in the multidomain proteins. Lastly, using protein kinase as an example, we identify regions of conformational variations and rigid body movements in the enzyme as it is changed to the active state from an inactive state.  相似文献   

8.
We developed a novel approach for predicting local protein structure from sequence. It relies on the Hybrid Protein Model (HPM), an unsupervised clustering method we previously developed. This model learns three-dimensional protein fragments encoded into a structural alphabet of 16 protein blocks (PBs). Here, we focused on 11-residue fragments encoded as a series of seven PBs and used HPM to cluster them according to their local similarities. We thus built a library of 120 overlapping prototypes (mean fragments from each cluster), with good three-dimensional local approximation, i.e., a mean accuracy of 1.61 A Calpha root-mean-square distance. Our prediction method is intended to optimize the exploitation of the sequence-structure relations deduced from this library of long protein fragments. This was achieved by setting up a system of 120 experts, each defined by logistic regression to optimize the discrimination from sequence of a given prototype relative to the others. For a target sequence window, the experts computed probabilities of sequence-structure compatibility for the prototypes and ranked them, proposing the top scorers as structural candidates. Predictions were defined as successful when a prototype <2.5 A from the true local structure was found among those proposed. Our strategy yielded a prediction rate of 51.2% for an average of 4.2 candidates per sequence window. We also proposed a confidence index to estimate prediction quality. Our approach predicts from sequence alone and will thus provide valuable information for proteins without structural homologs. Candidates will also contribute to global structure prediction by fragment assembly.  相似文献   

9.
Current methods for identification of domains within protein sequences require either structural information or the identification of homologous domain sequences in different sequence contexts. Knowledge of structural domain boundaries is important for fold recognition experiments and structural determination by X-ray crystallography or nuclear magnetic resonance spectroscopy using the divide-and-conquer approach. Here, a new and conceptually simple method for the identification of structural domain boundaries in multiple protein sequence alignments is presented. Analysis of covariance at positions within the alignment is first used to predict 3D contacts. By the nature of the domain as an independent folding unit, inter-domain predicted contacts are fewer than intra-domain predicted contacts. By analysing all possible domain boundaries and constructing a smoothed profile of predicted contact density (PCD), true structural domain boundaries are predicted as local profile minima associated with low PCD. A training data set is constructed from 52 non-homologous two-domain protein sequences of known 3D structure and used to determine optimal parameters for the profile analysis. The alignments in the training data set contained 48 +/- 17 (mean +/- SD) sequences and lengths of 257 +/- 121 residues. Of the 47 alignments yielding predictions, 35% of true domain boundaries are predicted to within 15 amino acids by the local profile minimum with the lowest profile value. Including predictions from the second- and third-lowest local minima increases the correct domain boundary coverage to 60%, whereas the lowest five local minima cover 79% of correct domain boundaries. Through further profile analysis, criteria are presented which reliably identify subsets of more accurate predictions. Retrospective analysis of CASP3 targets shows predictions of sufficient accuracy to enable dramatically improved fold recognition results. Finally, a prediction is made for geminivirus AL1 protein which is in full agreement with biochemical data, yielding a plausible, novel threading result.  相似文献   

10.
The function of a protein molecule is greatly influenced by its three-dimensional (3D) structure and therefore structure prediction will help identify its biological function. We have updated Sequence, Motif and Structure (SMS), the database of structurally rigid peptide fragments, by combining amino acid sequences and the corre-sponding 3D atomic coordinates of non-redundant (25%) and redundant (90%) protein chains available in the Protein Data Bank (PDB). SMS 2.0 provides information pertaining to the peptide fragments of length 5-14 resi-dues. The entire dataset is divided into three categories, namely, same sequence motifs having similar, intermedi-ate or dissimilar 3D structures. Further, options are provided to facilitate structural superposition using the pro-gram structural alignment of multiple proteins (STAMP) and the popular JAVA plug-in (Jmol) is deployed for visualization. In addition, functionalities are provided to search for the occurrences of the sequence motifs in other structural and sequence databases like PDB, Genome Database (GDB), Protein Information Resource (PIR) and Swiss-Prot. The updated database along with the search engine is available over the World Wide Web through the following URL http://cluster.physics.iisc.ernet.in/sms/.  相似文献   

11.
12.
Proteins in the intracellular lipid-binding protein (iLBP) family show remarkably high structural conservation despite their low-sequence identity. A multiple-sequence alignment using 52 sequences of iLBP family members revealed 15 fully conserved positions, with a disproportionately high number of these (n=7) located in the relatively small helical region. The conserved positions displayed high structural conservation based on comparisons of known iLBP crystal structures. It is striking that the beta-sheet domain had few conserved positions, despite its high structural conservation. This observation prompted us to analyze pair-wise interactions within the beta-sheet region to ask whether structural information was encoded in interacting amino acid pairs. We conducted this analysis on the iLBP family member, cellular retinoic acid-binding protein I (CRABP I), whose folding mechanism is under study in our laboratory. Indeed, an analysis based on a simple classification of hydrophobic and polar amino acids revealed a network of conserved interactions in CRABP I that cluster spatially, suggesting a possible nucleation site for folding. Significantly, a small number of residues participated in multiple conserved interactions, suggesting a key role for these sites in the structure and folding of CRABP I. The results presented here correlate well with available experimental evidence on folding of CRABPs and their family members and suggest future experiments. The analysis also shows the usefulness of considering pair-wise conservation based on a simple classification of amino acids, in analyzing sequences and structures to find common core regions among homologues.  相似文献   

13.
Secondary structure formation and stability are essential features in the knowledge of complex folding topology of biomolecules. To better understand the relationships between preferred conformations and functional properties of beta-homo-amino acids, the synthesis and conformational characterization by X-ray diffraction analysis of peptides containing conformationally constrained Calpha,alpha-dialkylated amino acid residues, such as alpha-aminoisobutyric acid or 1-aminocyclohexane-1-carboxylic acid and a single beta-homoamino acid, differently displaced along the peptide sequence have been carried out. The peptides investigated are: Boc-betaHLeu-(Ac6c)2-OMe, Boc-Ac6c-betaHLeu-(Ac6c)2-OMe and Boc-betaHVal-(Aib)5-OtBu, together with the C-protected beta-homo-residue HCl.H-betaHVal-OMe. The results indicate that the insertion of a betaH-residue at position 1 or 2 of peptides containing strong helix-inducing, bulky Calpha,alpha-disubstituted amino acid residues does not induce any specific conformational preferences. In the crystal state, most of the NH groups of beta-homo residues of tri- and tetrapeptides are not involved in intramolecular hydrogen bonds, thus failing to achieve helical structures similar to those of peptides exclusively constituted of Calpha,alpha-disubstituted amino acid residues. However, by repeating the structural motifs observed in the molecules investigated, a beta-pleated sheet secondary structure, and a new helical structure, named (14/15)-helix, were generated, corresponding to calculated minimum-energy conformations. Our findings, as well as literature data, strongly indicate that conformations of betaH-residues, with the micro torsion angle equal to -60 degrees, are very unlikely.  相似文献   

14.
The local environment of an amino acid in a folded protein determines the acceptability of mutations at that position. In order to characterize and quantify these structural constraints, we have made a comparative analysis of families of homologous proteins. Residues in each structure are classified according to amino acid type, secondary structure, accessibility of the side chain, and existence of hydrogen bonds from the side chains. Analysis of the pattern of observed substitutions as a function of local environment shows that there are distinct patterns, especially for buried polar residues. The substitution data tables are available on diskette with Protein Science. Given the fold of a protein, one is able to predict sequences compatible with the fold (profiles or templates) and potentially to discriminate between a correctly folded and misfolded protein. Conversely, analysis of residue variation across a family of aligned sequences in terms of substitution profiles can allow prediction of secondary structure or tertiary environment.  相似文献   

15.
Protein phosphorylation is widely used in biological regulatory processes. The study of spatial features related to phosphorylation sites is necessary to increase the efficacy of recognition of phosphorylation patterns in protein sequences. Using the data on phosphosites found in amino acid sequences, we mapped these sites onto 3D structures and studied the structural variability of the same sites in different PDB entries related to the same proteins. Solvent accessibility was calculated for the residues known to be phosphorylated. A significant change in accessibility was shown for many sites, but several ones were determined as buried in all the structures considered. Most phosphosites were found in coil regions. However, a significant portion was located in the structurally stable ordered regions. Comparison of structures with the same sites in modified and unmodified states showed that the region surrounding a site could be significantly shifted due to phosphorylation. Comparison between non‐modified structures (as well as between the modified ones) suggested that phosphorylation stabilizes one of the possible conformations. The local structure around the site could be changed due to phosphorylation, but often the initial conformation of the site surrounding is not altered within bounds of a rather large substructure. In this case, we can observe an extensive displacement within a protein domain. Phosphorylation without structural alteration seems to provide the interface for domain‐domain or protein‐protein interactions. Accounting for structural features is important for revealing more specific patterns of phosphorylation. It is also necessary for explaining structural changes as a basis for regulatory processes.  相似文献   

16.
Nuclear magnetic resonance (NMR) methods have been used to address issues regarding the relevance and feasibility of zinc binding to "zinc finger-like" sequences of the type C-X2-C-X4-H-X4-C [referred to as CCHC or retroviral-type (RT) zinc finger sequences]. One-dimensional (1D) NMR experiments with an 18-residue synthetic peptide containing the amino acid sequence of an HIV-1 RT-zinc finger domain (HIV1-F1) indicate that the sequences are capable of binding zinc tightly and stoichiometrically. 1H-113Cd spin echo difference NMR data confirm that the Cys and His amino acids are coordinated to metal in the 113Cd adduct. The 3D structure of the zinc adduct [Zn(HIV1-F1)] was determined to high atomic resolution by a new NMR-based approach that utilizes 2D-NOESY back-calculations as a measure of the consistency between the structures and the experimental data. Several interesting structural features were observed, including (1) the presence of extensive internal hydrogen bonding, and (2) the similarity of the folding of the first six residues to the folding observed by X-ray crystallography for related residues in the iron domain of rubredoxin. Structural constraints associated with conservatively substituted glycines provide further rationale for the physiological relevance of the zinc adduct. Similar NMR and structural results have been obtained for the second HIV-1 RT-zinc finger peptide, Zn(HIV1-F2). NMR studies of the zinc adduct with the NCP isolated directly from HIV-1 particles provide solid evidence that zinc finger domains are formed that are conformationally similar (if not identical) to the peptide structures.(ABSTRACT TRUNCATED AT 250 WORDS)  相似文献   

17.
Analysis of Mgm101p isolated from mitochondria shows that the mature protein of 27.6 kDa lacks 22 amino acids from the N-terminus. This mitochondrial targeting sequence has been incorporated in the design of oligonucleotides used to determine a functional core of Mgm101p. Progressive deletions, although retaining the targeting sequence, reveal that 76 N-terminal and six C-terminal amino acids of Mgm101p can be removed without altering the ability to complement an mgm101-1(ts) temperature-sensitive mutant. However, this active core is unable to complement mgm101 null mutants, suggesting that the Mgm101p might need to form a dimer or multimer to be functional in vivo. The active core, enriched in basic residues, contains 165 amino acids with a pI of 9.2. Alignment with 22 Mgm101p sequences from other lower eukaryotes shows that a number of amino acids are highly conserved in this region. Random mutagenesis confirms that certain critical amino acids required for function are invariant across the 23 proteins. Searches in the PFAM database revealed a low level of structural similarity between the active core and the Rad52 protein family.  相似文献   

18.
To facilitate investigation of the molecular and biochemical functions of the adenovirus E4 Orf6 protein, we sought to derive three-dimensional structural information using computational methods, particularly threading and comparative protein modeling. The amino acid sequence of the protein was used for secondary structure and hidden Markov model (HMM) analyses, and for fold recognition by the ProCeryon program. Six alternative models were generated from the top-scoring folds identified by threading. These models were examined by 3D-1D analysis and evaluated in the light of available experimental evidence. The final model of the E4 protein derived from these and additional threading calculations was a chimera, with the tertiary structure of its C-terminal 226 residues derived from a TIM barrel template and a mainly alpha-nonbundle topology for its poorly conserved N-terminal 68 residues. To assess the accuracy of this model, additional threading calculations were performed with E4 Orf6 sequences altered as in previous experimental studies. The proposed structural model is consistent with the reported secondary structure of a functionally important C-terminal sequence and can account for the properties of proteins carrying alterations in functionally important sequences or of those that disrupt an unusual zinc-coordination motif.  相似文献   

19.
A statistical analysis of the Protein Databank (PDB) structures had led us to define a set of small 3D structural prototypes called Protein Blocks (PBs). This structural alphabet includes 16 PBs, each one defined by the (phi, psi) dihedral angles of 5 consecutive residues. Here, we analyze the effect of the enlargement of the PDB on the PBs' definition. The results highlight the quality of the 3D approximation ensured by the PBs. These last could be of great interest in ab initio modeling.  相似文献   

20.
We present a thorough analysis of the relation between amino acid sequence and local three-dimensional structure in proteins. A library of overlapping local structural prototypes was built using an unsupervised clustering approach called “hybrid protein model” (HPM). The HPM carries out a multiple structural alignment of local folds from a non-redundant protein structure databank encoded into a structural alphabet composed of 16 protein blocks (PBs). Following previous research focusing on the HPM protocol, we have considered gaps in the local structure prototype. This methodology allows to have variable length fragments. Hence, 120 local structure prototypes were obtained. Twenty-five percent of the protein fragments learnt by HPM had gaps.An investigation of tight turns suggested that they are mainly derived from three PB series with precise locations in the HPM. The amino acid information content of the whole conformational classes was tackled by multivariate methods, e.g., canonical correlation analysis. It points out the presence of seven amino acid equivalence classes showing high propensities for preferential local structures. In the same way, definition of “contrast factors” based on sequence-structure properties underline the specificity of certain structural prototypes, e.g., the dependence of Gly or Asn-rich turns to a limited number of PBs, or, the opposition between Pro-rich coils to those enriched in Ser, Thr, Asn and Glu. These results are so useful to analyze the sequence-structure relationships, but could also be used to improve fragment-based method for protein structure prediction from sequence.  相似文献   

设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号