首页 | 本学科首页   官方微博 | 高级检索  
相似文献
 共查询到20条相似文献,搜索用时 15 毫秒
1.
There is currently a gap in knowledge between complexes of known three-dimensional structure and those known from other experimental methods such as affinity purifications or the two-hybrid system. This gap can sometimes be bridged by methods that extrapolate interaction information from one complex structure to homologues of the interacting proteins. To do this, it is important to know if and when proteins of the same type (e.g. family, superfamily or fold) interact in the same way. Here, we study interactions of known structure to address this question. We found all instances within the structural classification of proteins database of the same domain pairs interacting in different complexes, and then compared them with a simple measure (interaction RMSD). When plotted against sequence similarity we find that close homologues (30-40% or higher sequence identity) almost invariably interact the same way. Conversely, similarity only in fold (i.e. without additional evidence for a common ancestor) is only rarely associated with a similarity in interaction. The results suggest that there is a twilight zone of sequence similarity where it is not possible to say whether or not domains will interact similarly. We also discuss the rare instances of fold similarities interacting the same way, and those where obviously homologous proteins interact differently.  相似文献   

2.
Li M  Huang Y  Xiao Y 《Proteins》2008,72(4):1161-1170
Proteins with symmetric structures are ideal models to investigate the sequence-structure relations. We investigate proteins with beta-trefoil fold and find they have different degrees of sequence symmetries although they show similar symmetric structures. To understand this, we calculate the strength of interactions of the beta-trefoil folds with surrounding environments and find the low degrees of sequence symmetries are often correlated with large external interactions. Our results give an additional confirmation of Anfinsen's thermodynamic hypothesis that protein structures are not only determined by their sequences but also by their surrounding environments. We suggest the external interactions should be considered additionally in protein structure prediction through ab initio folding.  相似文献   

3.
Based on previous studies of interleukin-1beta (IL-1beta) and both acidic and basic fibroblast growth factors (FGFs), it has been suggested that the folding of beta-trefoil proteins is intrinsically slow and may occur via the formation of essential intermediates. Using optical and NMR-detected quenched-flow hydrogen/deuterium exchange methods, we have measured the folding kinetics of hisactophilin, another beta-trefoil protein that has < 10% sequence identity and unrelated function to IL-1beta and FGFs. We find that hisactophilin can fold rapidly and with apparently two-state kinetics, except under the most stabilizing conditions investigated where there is evidence for formation of a folding intermediate. The hisactophilin intermediate has significant structural similarities to the IL-1beta intermediate that has been observed experimentally and predicted theoretically using a simple, topology-based folding model; however, it appears to be different from the folding intermediate observed experimentally for acidic FGF. For hisactophilin and acidic FGF, intermediates are much less prominent during folding than for IL-1beta. Considering the structures of the different beta-trefoil proteins, it appears that differences in nonconserved loops and hydrophobic interactions may play an important role in differential stabilization of the intermediates for these proteins.  相似文献   

4.
BACKGROUND: Structures that have diverged from a common ancestor often retain functional and sequence similarity, although the latter may be very reduced. Even so, the overall fold of the structure is generally highly conserved. Now however, several have been identified of proteins that have been identified that have different functions but which have converged to a similar fold. These proteins will also have low sequence identities. RESULTS: By comparing the complete structure databank against itself, using sequence and structure alignment techniques, we have been able to identify six new examples of structurally related folds that have no apparent sequence or functional similarity. These related proteins include a family of crambin-like folds and a family of ferredoxin II folds. We found that all the similarities between structures are present in small proteins and occur as motifs within the core of a larger protein. CONCLUSION: The low sequence similarity and the lack of any obvious functional relationship between proteins with similar structures suggest that the proteins have diverged from independent ancestors. The similarities may therefore be of interest for understanding the various stereochemical and physical criteria that operate to generate a favourable fold.  相似文献   

5.
6.
Feng J  Li M  Huang Y  Xiao Y 《PloS one》2010,5(11):e14138
To understand how symmetric structures of many proteins are formed from asymmetric sequences, the proteins with two repeated beta-trefoil domains in Plant Cytotoxin B-chain family and all presently known beta-trefoil proteins are analyzed by structure-based multi-sequence alignments. The results show that all these proteins have similar key structural residues that are distributed symmetrically in their structures. These symmetric key structural residues are further analyzed in terms of inter-residues interaction numbers and B-factors. It is found that they can be distinguished from other residues and have significant propensities for structural framework. This indicates that these key structural residues may conduct the formation of symmetric structures although the sequences are asymmetric.  相似文献   

7.
The ability to predict structure from sequence is particularly important for toxins, virulence factors, allergens, cytokines, and other proteins of public health importance. Many such functions are represented in the parallel beta-helix and beta-trefoil families. A method using pairwise beta-strand interaction probabilities coupled with evolutionary information represented by sequence profiles is developed to tackle these problems for the beta-helix and beta-trefoil folds. The algorithm BetaWrapPro employs a "wrapping" component that may capture folding processes with an initiation stage followed by processive interaction of the sequence with the already-formed motifs. BetaWrapPro outperforms all previous motif recognition programs for these folds, recognizing the beta-helix with 100% sensitivity and 99.7% specificity and the beta-trefoil with 100% sensitivity and 92.5% specificity, in crossvalidation on a database of all nonredundant known positive and negative examples of these fold classes in the PDB. It additionally aligns 88% of residues for the beta-helices and 86% for the beta-trefoils accurately (within four residues of the exact position) to the structural template, which is then used with the side-chain packing program SCWRL to produce 3D structure predictions. One striking result has been the prediction of an unexpected parallel beta-helix structure for a pollen allergen, and its recent confirmation through solution of its structure. A Web server running BetaWrapPro is available and outputs putative PDB-style coordinates for sequences predicted to form the target folds.  相似文献   

8.
Several studies based on the known three-dimensional (3-D) structures of proteins show that two homologous proteins with insignificant sequence similarity could adopt a common fold and may perform same or similar biochemical functions. Hence, it is appropriate to use similarities in 3-D structure of proteins rather than the amino acid sequence similarities in modelling evolution of distantly related proteins. Here we present an assessment of using 3-D structures in modelling evolution of homologous proteins. Using a dataset of 108 protein domain families of known structures with at least 10 members per family we present a comparison of extent of structural and sequence dissimilarities among pairs of proteins which are inputs into the construction of phylogenetic trees. We find that correlation between the structure-based dissimilarity measures and the sequence-based dissimilarity measures is usually good if the sequence similarity among the homologues is about 30% or more. For protein families with low sequence similarity among the members, the correlation coefficient between the sequence-based and the structure-based dissimilarities are poor. In these cases the structure-based dendrogram clusters proteins with most similar biochemical functional properties better than the sequence-similarity based dendrogram. In multi-domain protein families and disulphide-rich protein families the correlation coefficient for the match of sequence-based and structure-based dissimilarity (SDM) measures can be poor though the sequence identity could be higher than 30%. Hence it is suggested that protein evolution is best modelled using 3-D structures if the sequence similarities (SSM) of the homologues are very low.  相似文献   

9.
Reddy BV  Li WW  Shindyalov IN  Bourne PE 《Proteins》2001,42(2):148-163
An all-against-all protein structure comparison using the Combinatorial Extension (CE) algorithm applied to a representative set of PDB structures revealed a gallery of common substructures in proteins (http://cl.sdsc.edu/ce.html). These substructures represent commonly identified folds, domains, or components thereof. Most of the subsequences forming these similar substructures have no significant sequence similarity. We present a method to identify conserved amino acid positions and residue-dependent property clusters within these subsequences starting with structure alignments. Each of the subsequences is aligned to its homologues in SWALL, a nonredundant protein sequence database. The most similar sequences are purged into a common frequency matrix, and weighted homologues of each one of the subsequences are used in scoring for conserved key amino acid positions (CKAAPs). We have set the top 20% of the high-scoring positions in each substructure to be CKAAPs. It is hypothesized that CKAAPs may be responsible for the common folding patterns in either a local or global view of the protein-folding pathway. Where a significant number of structures exist, CKAAPs have also been identified in structure alignments of complete polypeptide chains from the same protein family or superfamily. Evidence to support the presence of CKAAPs comes from other computational approaches and experimental studies of mutation and protein-folding experiments, notably the Paracelsus challenge. Finally, the structural environment of CKAAPs versus non-CKAAPs is examined for solvent accessibility, hydrogen bonding, and secondary structure. The identification of CKAAPs has important implications for protein engineering, fold recognition, modeling, and structure prediction studies and is dependent on the availability of structures and an accurate structure alignment methodology. Proteins 2001;42:148-163.  相似文献   

10.
Hisactophilin is a histidine-rich pH-dependent actin-binding protein from Dictyostelium discoideum. The structure of hisactophilin is typical of the beta-trefoil fold, a common structure adopted by diverse proteins with unrelated primary sequences and functions. The thermodynamics of denaturation of hisactophilin have been measured using fluorescence- and CD-monitored equilibrium urea denaturation curves, pH-denaturation, and thermal denaturation curves, as well as differential scanning calorimetry. Urea denaturation is reversible from pH 5.7 to pH 9.7; however, thermal denaturation is highly reversible only below pH approximately 6.2. Reversible denaturation by urea and heat is well fit using a two-state transition between the native and the denatured states. Urea denaturation curves are best fit using a quadratic dependence of the Gibbs free energy of unfolding upon urea concentration. Hisactophilin has moderate, roughly constant stability from pH 7.7 to pH 9.7; however, below pH 7.7, stability decreases markedly, most likely due to protonation of histidine residues. Enthalpic effects of histidine ionization upon unfolding also appear to be involved in the occurrence of cold unfolding of hisactophilin under relatively mild solution conditions. The stability data for hisactophilin are compared with data on hisactophilin function, and with data for two other beta-trefoil proteins, human interleukin-1beta, and basic fibroblast growth factor.  相似文献   

11.
The question of protein homology versus analogy arises when proteins share a common function or a common structural fold without any statistically significant amino acid sequence similarity. Even though two or more proteins do not have similar sequences but share a common fold and the same or closely related function, they are assumed to be homologs, descendant from a common ancestor. The problem of homolog identification is compounded in the case of proteins of 100 or less amino acids. This is due to a limited number of basic single domain folds and to a likelihood of identifying by chance sequence similarity. The latter arises from two conditions: first, any search of the currently very large protein database is likely to identify short regions of chance match; secondly, a direct sequence comparison among a small set of short proteins sharing a similar fold can detect many similar patterns of hydrophobicity even if proteins do not descend from a common ancestor. In an effort to identify distant homologs of the many ubiquitin proteins, we have developed a combined structure and sequence similarity approach that attempts to overcome the above limitations of homolog identification. This approach results in the identification of 90 probable ubiquitin-related proteins, including examples from the two prokaryotic domains of life, Archaea and Bacteria.  相似文献   

12.
The information required to generate a protein structure is contained in its amino acid sequence, but how three-dimensional information is mapped onto a linear sequence is still incompletely understood. Multiple structure alignments of similar protein structures have been used to investigate conserved sequence features but contradictory results have been obtained, due, in large part, to the absence of subjective criteria to be used in the construction of sequence profiles and in the quantitative comparison of alignment results. Here, we report a new procedure for multiple structure alignment and use it to construct structure-based sequence profiles for similar proteins. The definition of "similar" is based on the structural alignment procedure and on the protein structural distance (PSD) described in paper I of this series, which offers an objective measure for protein structure relationships. Our approach is tested in two well-studied groups of proteins; serine proteases and Ig-like proteins. It is demonstrated that the quality of a sequence profile generated by a multiple structure alignment is quite sensitive to the PSD used as a threshold for the inclusion of proteins in the alignment. Specifically, if the proteins included in the aligned set are too distant in structure from one another, there will be a dilution of information and patterns that are relevant to a subset of the proteins are likely to be lost.In order to understand better how the same three-dimensional information can be encoded in seemingly unrelated sequences, structure-based sequence profiles are constructed for subsets of proteins belonging to nine superfolds. We identify patterns of relatively conserved residues in each subset of proteins. It is demonstrated that the most conserved residues are generally located in the regions where tertiary interactions occur and that are relatively conserved in structure. Nevertheless, the conservation patterns are relatively weak in all cases studied, indicating that structure-determining factors that do not require a particular sequential arrangement of amino acids, such as secondary structure propensities and hydrophobic interactions, are important in encoding protein fold information. In general, we find that similar structures can fold without having a set of highly conserved residue clusters or a well-conserved sequence profile; indeed, in some cases there is no apparent conservation pattern common to structures with the same fold. Thus, when a group of proteins exhibits a common and well-defined sequence pattern, it is more likely that these sequences have a close evolutionary relationship rather than the similarities having arisen from the structural requirements of a given fold.  相似文献   

13.
We have used GRATH, a graph-based structure comparison algorithm, to map the similarities between the different folds observed in the CATH domain structure database. Statistical analysis of the distributions of the fold similarities has allowed us to assess the significance for any similarity. Therefore we have examined whether it is best to represent folds as discrete entities or whether, in fact, a more accurate model would be a continuum wherein folds overlap via common motifs. To do this we have introduced a new statistical measure of fold similarity, termed gregariousness. For a particular fold, gregariousness measures how many other folds have a significant structural overlap with that fold, typically comprising 40% or more of the larger structure. Gregarious folds often contain commonly occurring super-secondary structural motifs, such as beta-meanders, greek keys, alpha-beta plait motifs or alpha-hairpins, which are matching similar motifs in other folds. Apart from one example, all the most gregarious folds matching 20% or more of the other folds in the database, are alpha-beta proteins. They also occur in highly populated architectural regions of fold space, adopting sandwich-like arrangements containing two or more layers of alpha-helices and beta-strands.Domains that exhibit a low gregariousness, are those that have very distinctive folds, with few common motifs or motifs that are packed in unusual arrangements. Most of the superhelices exhibit low gregariousness despite containing some commonly occurring super-secondary structural motifs. In these folds, these common motifs are combined in an unusual way and represent a small proportion of the fold (<10%). Our results suggest that fold space may be considered as continuous for some architectural arrangements (e.g. alpha-beta sandwiches), in that super-secondary motifs can be used to link neighbouring fold groups. However, in other regions of fold space much more discrete topologies are observed with little similarity between folds.  相似文献   

14.
Previous crystallographic analyses of the Kunitz inhibitors from soybean. Erythrina caffra and wheat, the interleukins-1 beta and 1 alpha and the acidic and basic fibroblast growth factors have shown that they contain a most unusual fold. It is formed by six two-stranded hairpins. Three of these form a barrel structure and the other three are in a triangular array that caps the barrel. The arrangement of the secondary structures gives the molecules a pseudo 3-fold axis. Although the different proteins have very similar structures, many of their sequences have no significant similarities overall. The structural determinants of this fold are described and discussed in this paper. The barrels in the different proteins have the same geometrical features: six strands tilted at 56 degrees to the barrel axis; a barrel diameter of 16 A, and the beta-sheet hydrogen bonded so that it is staggered with a shear number of 12. These features fit McLachlan's equations for ideal barrels formed by beta-sheets. The wide diameter of the barrels is filled by layers of residues that, while not identical in the different proteins, are, in almost all cases, large. The structure of the triangular array of hairpins is determined by the coiling of the strands and the packing of hairpin residues against each other and against residues from the interior of the barrel. The major sequence requirements of this fold are large or medium hydrophobic residues at 18 buried sites. In the different structures the total volume of these residues is 3000 (+/- 120) A. The polyhedron model of protein architecture is used to demonstrate that the main, and in particular the symmetrical, features of this fold arise from the ideal and equal packing of six hairpins, modified only slightly to form hydrogen bonds between the hairpins.  相似文献   

15.
Proteins might have considerable structural similarities even when no evolutionary relationship of their sequences can be detected. This property is often referred to as the proteins sharing only a "fold". Of course, there are also sequences of common origin in each fold, called a "superfamily", and in them groups of sequences with clear similarities, designated "family". Developing algorithms to reliably identify proteins related at any level is one of the most important challenges in the fast growing field of bioinformatics today. However, it is not at all certain that a method proficient at finding sequence similarities performs well at the other levels, or vice versa.Here, we have compared the performance of various search methods on these different levels of similarity. As expected, we show that it becomes much harder to detect proteins as their sequences diverge. For family related sequences the best method gets 75% of the top hits correct. When the sequences differ but the proteins belong to the same superfamily this drops to 29%, and in the case of proteins with only fold similarity it is as low as 15%. We have made a more complete analysis of the performance of different algorithms than earlier studies, also including threading methods in the comparison. Using this method a more detailed picture emerges, showing multiple sequence information to improve detection on the two closer levels of relationship. We have also compared the different methods of including this information in prediction algorithms.For lower specificities, the best scheme to use is a linking method connecting proteins through an intermediate hit. For higher specificities, better performance is obtained by PSI-BLAST and some procedures using hidden Markov models. We also show that a threading method, THREADER, performs significantly better than any other method at fold recognition.  相似文献   

16.
ThiS is a sulfur carrier protein that plays a central role in thiamin biosynthesis in Escherichia coli. Here we report the solution NMR structure of ThiS, the first for this class of sulfur carrier proteins. Although ThiS shares only 14% sequence identity with ubiquitin, it possesses the ubiquitin fold. This structural homology, combined with established functional similarities involving sulfur chemistry, demonstrates that the eukaryotic ubiquitin and the prokaryotic ThiS evolved from a common ancestor. This illustrates how structure determination is essential in establishing evolutionary links between proteins in which structure and function have been conserved through eons of evolution despite loss of sequence identity. The ThiS structure reveals both hydrophobic and electrostatic surface features that are likely determinants for interactions with binding partners. Comparison with surface features of ubiquitin and ubiquitin homologs SUMO-1, RUB-1 and NEDD8 suggest how Nature has utilized this single fold to incorporate similar chemistry into a broad array of highly specific biological processes.  相似文献   

17.
Most homologous pairs of proteins have no significant sequence similarity to each other and are not identified by direct sequence comparison or profile-based strategies. However, multiple sequence alignments of low similarity homologues typically reveal a limited number of positions that are well conserved despite diversity of function. It may be inferred that conservation at most of these positions is the result of the importance of the contribution of these amino acids to the folding and stability of the protein. As such, these amino acids and their relative positions may define a structural signature. We demonstrate that extraction of this fold template provides the basis for the sequence database to be searched for patterns consistent with the fold, enabling identification of homologs that are not recognized by global sequence analysis. The fold template method was developed to address the need for a tool that could comprehensively search the midnight and twilight zones of protein sequence similarity without reliance on global statistical significance. Manual implementations of the fold template method were performed on three folds--immunoglobulin, c-lectin and TIM barrel. Following proof of concept of the template method, an automated version of the approach was developed. This automated fold template method was used to develop fold templates for 10 of the more populated folds in the SCOP database. The fold template method developed three-dimensional structural motifs or signatures that were able to return a diverse collection of proteins, while maintaining a low false positive rate. Although the results of the manual fold template method were more comprehensive than the automated fold template method, the diversity of the results from the automated fold template method surpassed those of current methods that rely on statistical significance to infer evolutionary relationships among divergent proteins.  相似文献   

18.
The genomes of over 60 organisms from all three kingdoms of life are now entirely sequenced. In many respects, the inventory of proteins used in different kingdoms appears surprisingly similar. However, eukaryotes differ from other kingdoms in that they use many long proteins, and have more proteins with coiled-coil helices and with regions abundant in regular secondary structure. Particular structural domains are used in many pathways. Nevertheless, one domain tends to occur only once in one particular pathway. Many proteins do not have close homologues in different species (orphans) and there could even be folds that are specific to one species. This view implies that protein fold space is discrete. An alternative model suggests that structure space is continuous and that modern proteins evolved by aggregating fragments of ancient proteins. Either way, after having harvested proteomes by applying standard tools, the challenge now seems to be to develop better methods for comparative proteomics.  相似文献   

19.
Plants and metazoans share many similarities in terms of conserved proteins. Antibodies have been used extensively to detect remote homologues, many of which are yet to be identified conclusively. Genome sequencing and the creation of novel sequence or structure comparison programs have assisted greatly in the identification of distant protein homologues. The continuing development of new software algorithms and the combining of bioinformatics with proteomics offer hope that remaining homologues will be soon identified.  相似文献   

20.
Issues in predicting protein function from sequence   总被引:1,自引:0,他引:1  
Identifying homologues, defined as genes that arose from a common evolutionary ancestor, is often a relatively straightforward task, thanks to recent advances made in estimating the statistical significance of sequence similarities found from database searches. The extent by which homologues possess similarities in function, however, is less amenable to statistical analysis. Consequently, predicting function by homology is a qualitative, rather than quantitative, process and requires particular care to be taken. This review focuses on the various approaches that have been developed to predict function from the scale of the atom to that of the organism. Similarities in homologues' functions differ considerably at each of these different scales and also vary for different domain families. It is argued that due attention should be paid to all available clues to function, including orthologue identification, conservation of particular residue types, and the co-occurrence of domains in proteins. Pitfalls in database searching methods arising from amino acid compositional bias and database size effects are also discussed.  相似文献   

设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号