首页 | 本学科首页   官方微博 | 高级检索  
相似文献
 共查询到20条相似文献,搜索用时 15 毫秒
1.
The problem of estimating haplotype frequencies from population data has been considered by numerous investigators, resulting in a wide variety of possible algorithmic and statistical solutions. We propose a relatively unique approach that employs an artificial neural network (ANN) to predict the most likely haplotype frequencies from a sample of population genotype data. Through an innovative ANN design for mapping genotype patterns to diplotypes, we have produced a prototype that demonstrates the feasibility of this approach, with provisional results that correlate well with estimates produced by the expectation maximization algorithm for haplotype frequency estimation. Given the computational demands of estimating haplotype frequencies for 20 or more single-nucleotide polymorphisms, the ANN approach is promising because its design fits well with parallel computing architectures.  相似文献   

2.
We propose a method, SDpop, able to infer sex-linkage caused by recombination suppression typical of sex chromosomes. The method is based on the modeling of the allele and genotype frequencies of individuals of known sex in natural populations. It is implemented in a hierarchical probabilistic framework, accounting for different sources of error. It allows statistical testing for the presence or absence of sex chromosomes, and detection of sex-linked genes based on the posterior probabilities in the model. Furthermore, for gametologous sequences, the haplotype and level of nucleotide polymorphism of each copy can be inferred, as well as the divergence between them. We test the method using simulated data, as well as data from both a relatively recent and an old sex chromosome system (the plant Silene latifolia and humans) and show that, for most cases, robust predictions are obtained with 5 to 10 individuals per sex.  相似文献   

3.
Single nucleotide polymorphisms (SNPs) are plentiful in most genomes and amenable to high throughput genotyping, but they are not yet popular for parentage or paternity analysis. The markers are bi-allelic, so individually they contain little information about parentage, and in nonmodel organisms the process of identifying large numbers of unlinked SNPs can be daunting. We explore the possibility of using blocks of between three and 26 linked SNPs as highly polymorphic molecular markers for reconstructing male genotypes in polyandrous organisms with moderate (five offspring) to large (25 offspring) clutches of offspring. Haplotypes are inferred for each block of linked SNPs using the programs Haplore and Phase 2.1. Each multi-SNP haplotype is then treated as a separate allele, producing a highly polymorphic, 'microsatellite-like' marker. A simulation study is performed using haplotype frequencies derived from empirical data sets from Drosophila melanogaster and Mus musculus populations. We find that the markers produced are competitive with microsatellite loci in terms of single parent exclusion probabilities, particularly when using six or more linked SNPs to form a haplotype. These markers contain only modest rates of missing data and genotyping or phasing errors and thus should be seriously considered as molecular markers for parentage analysis, particularly when the study is interested in the functional significance of polymorphisms across the genome.  相似文献   

4.
Because current molecular haplotyping methods are expensive and not amenable to automation, many researchers rely on statistical methods to infer haplotype pairs from multilocus genotypes, and subsequently treat these inferred haplotype pairs as observations. These procedures are prone to haplotype misclassification. We examine the effect of these misclassification errors on the false-positive rate and power for two association tests. These tests include the standard likelihood ratio test (LRTstd) and a likelihood ratio test that employs a double-sampling approach to allow for the misclassification inherent in the haplotype inference procedure (LRTae). We aim to determine the cost-benefit relationship of increasing the proportion of individuals with molecular haplotype measurements in addition to genotypes to raise the power gain of the LRTae over the LRTstd. This analysis should provide a guideline for determining the minimum number of molecular haplotypes required for desired power. Our simulations under the null hypothesis of equal haplotype frequencies in cases and controls indicate that (1) for each statistic, permutation methods maintain the correct type I error; (2) specific multilocus genotypes that are misclassified as the incorrect haplotype pair are consistently misclassified throughout each entire dataset; and (3) our simulations under the alternative hypothesis showed a significant power gain for the LRTae over the LRTstd for a subset of the parameter settings. Permutation methods should be used exclusively to determine significance for each statistic. For fixed cost, the power gain of the LRTae over the LRTstd varied depending on the relative costs of genotyping, molecular haplotyping, and phenotyping. The LRTae showed the greatest benefit over the LRTstd when the cost of phenotyping was very high relative to the cost of genotyping. This situation is likely to occur in a replication study as opposed to a whole-genome association study.  相似文献   

5.
Zhao J  Jin L  Xiong M 《Genetics》2006,174(3):1529-1538
As millions of single-nucleotide polymorphisms (SNPs) have been identified and high-throughput genotyping technologies have been rapidly developed, large-scale genomewide association studies are soon within reach. However, since a genomewide association study involves a large number of SNPs it is therefore nearly impossible to ensure a genomewide significance level of 0.05 using the available statistics, although the multiple-test problems can be alleviated, but not sufficiently, by the use of tagging SNPs. One strategy to circumvent the multiple-test problem associated with genome-wide association tests is to develop novel test statistics with high power. In this report, we introduce several nonlinear tests, which are based on nonlinear transformation of allele or haplotype frequencies. We investigate the power of the nonlinear test statistics and demonstrate that under certain conditions, some nonlinear test statistics have much higher power than the standard chi2-test statistic. Type I error rates of the nonlinear tests are validated using simulation studies. We also show that a class of similarity measure-based test statistics is based on the quadratic function of allele or haplotype frequencies, and thus they belong to nonlinear tests. To evaluate their performance, the nonlinear test statistics are also applied to three real data sets. Our study shows that nonlinear test statistics have great potential in association studies of complex diseases.  相似文献   

6.
We analyzed allele frequencies and pairwise linkage disequilibria of 13 variants in the EDN1 gene of 298 young males, the majority of German ancestry. Our analysis comprises all common variants in the five exons and flanking intronic regions, as well as known polymorphisms in the promoter sequence. In addition to previously analyzed polymorphisms, our haplotype reconstruction included five recently described variants and was done by using three different algorithms to allow inference of result stability. More than 30 haplotypes were predicted. All haplotypes with frequencies > or = 1% were inferred by all three methods and can be described by seven haplotype tagging single-nucleotide polymorphisms (htSNPs), reducing the genotyping load to 65%. Three of these haplotypes with frequencies of about 11%, 9%, and 4% had been mistaken for one haplotype in the previous analysis, which included only six polymorphisms, some of them not being htSNPs. Systematic analysis of sequence variability and comprehensive haplotype analysis of the EDN1 gene determined a substantial part of its genetic variability for further association studies and helped to reduce the genotyping load for common phenotypes.  相似文献   

7.
To optimize the strategies for population-based pharmacogenetic studies, we extensively analyzed single-nucleotide polymorphisms (SNPs) and haplotypes in 199 drug-related genes, through use of 4,190 SNPs in 752 control subjects. Drug-related genes, like other genes, have a haplotype-block structure, and a few haplotype-tagging SNPs (htSNPs) could represent most of the major haplotypes constructed with common SNPs in a block. Because our data included 860 uncommon (frequency <0.1) SNPs with frequencies that were accurately estimated, we analyzed the relationship between haplotypes and uncommon SNPs within the blocks (549 SNPs). We inferred haplotype frequencies through use of the data from all htSNPs and one of the uncommon SNPs within a block and calculated four joint probabilities for the haplotypes. We show that, irrespective of the minor-allele frequency of an uncommon SNP, the majority (mean +/- SD frequency 0.943+/-0.117) of the minor alleles were assigned to a single haplotype tagged by htSNPs if the uncommon SNP was within the block. These results support the hypothesis that recombinations occur only infrequently within blocks. The proportion of a single haplotype tagged by htSNPs to which the minor alleles of an uncommon SNP were assigned was positively correlated with the minor-allele frequency when the frequency was <0.03 (P<.000001; n=233 [Spearman's rank correlation coefficient]). The results of simulation studies suggested that haplotype analysis using htSNPs may be useful in the detection of uncommon SNPs associated with phenotypes if the frequencies of the SNPs are higher in affected than in control populations, the SNPs are within the blocks, and the frequencies of the SNPs are >0.03.  相似文献   

8.
MOTIVATION: Haplotype reconstruction is an essential step in genetic linkage and association studies. Although many methods have been developed to estimate haplotype frequencies and reconstruct haplotypes for a sample of unrelated individuals, haplotype reconstruction in large pedigrees with a large number of genetic markers remains a challenging problem. METHODS: We have developed an efficient computer program, HAPLORE (HAPLOtype REconstruction), to identify all haplotype sets that are compatible with the observed genotypes in a pedigree for tightly linked genetic markers. HAPLORE consists of three steps that can serve different needs in applications. In the first step, a set of logic rules is used to reduce the number of compatible haplotypes of each individual in the pedigree as much as possible. After this step, the haplotypes of all individuals in the pedigree can be completely or partially determined. These logic rules are applicable to completely linked markers and they can be used to impute missing data and check genotyping errors. In the second step, a haplotype-elimination algorithm similar to the genotype-elimination algorithms used in linkage analysis is applied to delete incompatible haplotypes derived from the first step. All superfluous haplotypes of the pedigree members will be excluded after this step. In the third step, the expectation-maximization (EM) algorithm combined with the partition and ligation technique is used to estimate haplotype frequencies based on the inferred haplotype configurations through the first two steps. Only compatible haplotype configurations with haplotypes having frequencies greater than a threshold are retained. RESULTS: We test the effectiveness and the efficiency of HAPLORE using both simulated and real datasets. Our results show that, the rule-based algorithm is very efficient for completely genotyped pedigree. In this case, almost all of the families have one unique haplotype configuration. In the presence of missing data, the number of compatible haplotypes can be substantially reduced by HAPLORE, and the program will provide all possible haplotype configurations of a pedigree under different circumstances, if such multiple configurations exist. These inferred haplotype configurations, as well as the haplotype frequencies estimated by the EM algorithm, can be used in genetic linkage and association studies. AVAILABILITY: The program can be downloaded from http://bioinformatics.med.yale.edu.  相似文献   

9.
Innan H  Zhang K  Marjoram P  Tavaré S  Rosenberg NA 《Genetics》2005,169(3):1763-1777
Several tests of neutral evolution employ the observed number of segregating sites and properties of the haplotype frequency distribution as summary statistics and use simulations to obtain rejection probabilities. Here we develop a “haplotype configuration test” of neutrality (HCT) based on the full haplotype frequency distribution. To enable exact computation of rejection probabilities for small samples, we derive a recursion under the standard coalescent model for the joint distribution of the haplotype frequencies and the number of segregating sites. For larger samples, we consider simulation-based approaches. The utility of the HCT is demonstrated in simulations of alternative models and in application to data from Drosophila melanogaster.  相似文献   

10.
We analyzed flavin-containing monooxygenase 3 (FMO3) polymorphisms, haplotype structure, and linkage disequilibrium (LD) in 256 Han Chinese and 50 African-American individuals to compare their haplotype frequencies and LD with other world populations. For the Han Chinese, genotyping of three haplotype tag single nucleotide polymorphisms (E158K, V257M, and E308G) was performed by polymerase chain reaction (PCR)-restriction fragment length polymorphism. For the African-Americans, genotyping of all coding exons was performed by modified PCR-single strand conformational polymorphism. Haplotype frequencies, LD, and evolutionary rates were inferred and estimated computationally. There were significant differences in haplotype frequency distribution and LD pattern among Han Chinese, African-Americans, and other world populations. Four major haplotypes of Han Chinese were EVE, KVE, EME, and EVG. Two major haplotypes of African-Americans were EVE and KVE. We found that sites 158 and 257 are in significant LD in both populations. This is the first report comparing FMO haplotypes and LD of Han Chinese with African-Americans. The data presented here justify further pharmacogenetic studies for potentially optimizing recommended drug dosages and evaluating relationships with disease processes.  相似文献   

11.
Huang ZS  Ji YJ  Zhang DX 《Molecular ecology》2008,17(8):1930-1947
Single copy nuclear polymorphic (scnp) DNA is potentially a powerful molecular marker for evolutionary studies of populations. However, a practical obstacle to its employment is the general problem of haplotype determination due to the common occurrence of heterozygosity in diploid organisms. We explore here a 'consensus vote' (CV) approach to this question, combining statistical haplotype reconstruction and experimental verification using as an example an indel-free scnp DNA marker from the flanking region of a microsatellite locus of the migratory locust. The raw data comprise 251-bp sequences from 526 locust individuals (1052 chromosomes), with 71 (28.3%) polymorphic nucleotide sites (including seven triallelic sites) and 141 distinct genotypes (with frequencies ranging from 0.2 to 25.5%). Six representative statistical haplotype reconstruction algorithms are employed in our CV approach, including one parsimony method, two expectation-maximization (EM) methods and three Bayesian methods. The phases of 116 ambiguous individuals inferred by this approach are verified by molecular cloning experiments. We demonstrate the effectiveness of the CV approach compared to inferences based on individual statistical algorithms. First, it has the unique power to partition the inferrals into a reliable group and an uncertain group, thereby allowing the identification of the inferrals with greater uncertainty (12.7% of the total sample in this case). This considerably reduces subsequent efforts of experimental verification. Second, this approach is capable of handling genotype data pooled from many geographical populations, thus tolerating heterogeneity of genetic diversity among populations. Third, the performance of the CV approach is not influenced by the number of heterozygous sites in the ambiguous genotypes. Therefore, the CV approach is potentially a reliable strategy for effective haplotype determination of nuclear DNA markers. Our results also show that rare variations and rare inferrals tend to be more vulnerable to inference error, and hence deserve extra surveillance.  相似文献   

12.
Furihata S  Ito T  Kamatani N 《Genetics》2006,174(3):1505-1516
The use of haplotype information in case-control studies is an area of focus for the research on the association between phenotypes and genetic polymorphisms. We examined the validity of the application of the likelihood-based algorithm, which was originally developed to analyze the data from cohort studies or clinical trials, to the data from case-control studies. This algorithm was implemented in a computer program called PENHAPLO. In this program, haplotype frequencies and penetrances are estimated using the expectation-maximization algorithm, and the haplotype-phenotype association is tested using the generalized likelihood ratio. We show that this algorithm was useful not only for cohort studies but also for case-control studies. Simulations under the null hypothesis (no association between haplotypes and phenotypes) have shown that the type I error rates were accurately estimated. The simulations under alternative hypotheses showed that PENHAPLO is a robust method for the analysis of the data from case-control studies even when the haplotypes were not in HWE, although real penetrances cannot be estimated. The power of PENHAPLO was higher than that of other methods using the likelihood-ratio test for the comparison of haplotype frequencies. Results of the analysis of real data indicated that a significant association between haplotypes in the SAA1 gene and AA-amyloidosis phenotype was observed in patients with rheumatoid arthritis, thereby suggesting the validity of the application of PENHAPLO for case-control data.  相似文献   

13.
We analyzed flavin-containing monooxygenase 3 (FMO3) polymorphisms, haplotype structure, and linkage disequilibrium (LD) in 256 Han Chinese and 50 African-American individuals to compare their haplotype frequencies and LD with other world populations. For the Han Chinese, genotyping of three haplotype tag single nucleotide polymorphisms (E158K, V257M, and E308G) was performed by polymerase chain reaction (PCR)-restriction fragment length polymorphism. For the African-Americans, genotyping of all coding exons was performed by modified PCR-single strand conformational polymorphism. Haplotype frequencies, LD, and evolutionary rates were inferred and estimated computationally. There were significant differences in haplotype frequency distribution and LD pattern among Han Chinese, African-Americans, and other world populations. Four major haplotypes of Han Chinese were EVE, KVE, EME, and EVG. Two major haplotypes of African-Americans were EVE and KVE. We found that sites 158 and 257 are in significant LD in both populations. This is the first report comparing FMO haplotypes and LD of Han Chinese with African-Americans. The data presented here justify further pharmacogenetic studies for potentially optimizing recommended drug dosages and evaluating relationships with disease processes.  相似文献   

14.
Family data teamed with the transmission/disequilibrium test (TDT), which simultaneously evaluates linkage and association, is a powerful means of detecting disease-liability alleles. To increase the information provided by the test, various researchers have proposed TDT-based methods for haplotype transmission. Haplotypes indeed produce more-definitive transmissions than do the alleles comprising them, and this tends to increase power. However, the larger number of haplotypes, relative to alleles at individual loci, tends to decrease power, because of the additional degrees of freedom required for the test. An optimal strategy would focus the test on particular haplotypes or groups of haplotypes. In this report we develop such an approach by combining the theory of TDT with that of measured haplotype analysis (MHA). MHA uses the evolutionary relationships among haplotypes to produce a limited set of hypothesis tests and to increase the interpretability of these tests. The theory of our approach, called the "evolutionary tree" (ET)-TDT, is developed for two cases: when haplotype transmission is certain and when it is not. Simulations show the ET-TDT can be more powerful than other proposed methods under reasonable conditions. More importantly, our results show that, when multiple polymorphisms are found within the gene, the ET-TDT can be useful for determining which polymorphisms affect liability.  相似文献   

15.
This study compares the properties of dominant markers, such as amplified fragment length polymorphisms (AFLPs), with those of codominant multiallelic markers, such as microsatellites, in reconstructing parentage. These two types of markers were used to search for both parents of an individual without prior knowledge of their relationships, by calculating likelihood ratios based on genotypic data, including mistyping. Experimental data on 89 oak trees genotyped for six microsatellite markers and 159 polymorphic AFLP loci were used as a starting point for simulations and tests. Both sets of markers produced high exclusion probabilities, and among dominant markers those with dominant allele frequencies in the range 0.1-0.4 were more informative. Such codominant and dominant markers can be used to construct powerful statistical tests to decide whether a genotyped individual (or two individuals) can be considered as the true parent (or parent pair). Gene flow from outside the study stand (GFO), inferred from parentage analysis with microsatellites, overestimated the true GFO, whereas with AFLPs it was underestimated. As expected, dominant markers are less efficient than codominant markers for achieving this, but can still be used with good confidence, especially when loci are deliberately selected according to their allele frequencies.  相似文献   

16.
Haplotype analysis of single nucleotide polymorphisms (SNPs) is an important and rapidly growing approach for association studies. In recent years, statistical procedures to haplotype determination from genotypic information have employed in population studies. These procedures, even though some advantages for estimation of haplotype frequencies in large population samples, have limitations in the accuracy of the analysis. In this study, we have designed a reliable method for direct haplotyping of polymorphic sites using the amplification refractory mutation system (ARMS) and restriction fragment length polymorphism (RFLP) analysis techniques. We applied the method to determination of haplotypes composed of three SNPs within the paraoxonase1 gene promoter and found the approach can be used in many studies in population and in a variety of clinical settings.  相似文献   

17.
Once genetic linkage has been identified for a complex disease, the next step is often association analysis, in which single-nucleotide polymorphisms (SNPs) within the linkage region are genotyped and tested for association with the disease. If a SNP shows evidence of association, it is useful to know whether the linkage result can be explained, in part or in full, by the candidate SNP. We propose a novel approach that quantifies the degree of linkage disequilibrium (LD) between the candidate SNP and the putative disease locus through joint modeling of linkage and association. We describe a simple likelihood of the marker data conditional on the trait data for a sample of affected sib pairs, with disease penetrances and disease-SNP haplotype frequencies as parameters. We estimate model parameters by maximum likelihood and propose two likelihood-ratio tests to characterize the relationship of the candidate SNP and the disease locus. The first test assesses whether the candidate SNP and the disease locus are in linkage equilibrium so that the SNP plays no causal role in the linkage signal. The second test assesses whether the candidate SNP and the disease locus are in complete LD so that the SNP or a marker in complete LD with it may account fully for the linkage signal. Our method also yields a genetic model that includes parameter estimates for disease-SNP haplotype frequencies and the degree of disease-SNP LD. Our method provides a new tool for detecting linkage and association and can be extended to study designs that include unaffected family members.  相似文献   

18.
We compared the accuracy of haplotype inferences at a 6 Mb region on chromosome 7 where significant linkage between a brain oscillation phenotype and a cholinergic muscarinic receptor gene was previously reported. Individual haplotype assignments and haplotype frequencies were estimated using 5, 10, and 14 consecutive Illumina single-nucleotide polymorphisms (SNPs) within the 1-LOD unit support interval of the chromosome 7 linkage peak. Initially, haplotypes were constructed incorporating phase information provided by relatives using the pedigree analysis package MERLIN. Population-based haplotypes were inferred using the haplotype estimation software HAPLO.STATS and PHASE, using unrelated individuals. The 14 SNPs within this region exhibited markedly low linkage disequilibrium, and the average D' estimate between SNPs was 0.18 (range: 0.01-0.97). In comparison to the family-based haplotypes calculated in MERLIN, the computational inferences of individual haplotype assignments were most accurate when considering 5 consecutive SNPs, but decayed dramatically when considering 10 or 14 SNPs in both PHASE and HAPLO.STATS. When comparing the two haplotype inference methods, both PHASE and HAPLO.STATS performed poorly. These analyses underscore the difficulties of haplotype estimation in the presence of low linkage disequilibrium and stress the importance of careful consideration of confidence measures when using estimated haplotype frequencies and individual assignments in biomedical research.  相似文献   

19.
Statistical properties of new neutrality tests against population growth   总被引:2,自引:0,他引:2  
A number of statistical tests for detecting population growth are described. We compared the statistical power of these tests with that of others available in the literature. The tests evaluated fall into three categories: those tests based on the distribution of the mutation frequencies, on the haplotype distribution, and on the mismatch distribution. We found that, for an extensive variety of cases, the most powerful tests for detecting population growth are Fu's F(S) test and the newly developed R(2) test. The behavior of the R(2) test is superior for small sample sizes, whereas F(S) is better for large sample sizes. We also show that some popular statistics based on the mismatch distribution are very conservative.  相似文献   

20.
A retrospective likelihood-based approach was proposed to test and estimate the effect of haplotype on disease risk using unphased genotype data with adjustment for environmental covariates. The proposed method was also extended to handle the data in which the haplotype and environmental covariates are not independent. Likelihood ratio tests were constructed to test the effects of haplotype and gene-environment interaction. The model parameters such as haplotype effect size was estimated using an Expectation Conditional-Maximization (ECM) algorithm developed by Meng and Rubin (1993). Model-based variance estimates were derived using the observed information matrix. Simulation studies were conducted for three different genetic effect models, including dominant effect, recessive effect, and additive effect. The results showed that the proposed method generated unbiased parameter estimates, proper type I error, and true beta coverage probabilities. The model performed well with small or large sample sizes, as well as short or long haplotypes.  相似文献   

设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号