首页 | 本学科首页   官方微博 | 高级检索  
相似文献
 共查询到20条相似文献,搜索用时 281 毫秒
1.
Several programs are currently available for the detection of genotyping error that may or may not be Mendelianly inconsistent. However, no systematic study exists that evaluates their performance under varying pedigree structures and sizes, marker spacing, and allele frequencies. Our simulation study compares four multipoint methods: Merlin, Mendel4, SimWalk2, and Sibmed. We look at empirical thresholds, power, and false-positive rates on 7 small pedigree structures that included sibships with and without genotyped parents, and a three-generation pedigree, using 11 microsatellite markers with 3 different map spacings. Simulated data includes 5,000 replicates of each pedigree structure and marker map, with random genotyping errors in about 4% of the middle marker's genotypes. We found that the default thresholds used by these programs provide low power (47-72%). Power is improved more by adding genotyped siblings than by using more closely spaced markers. Some mistyping methods are sensitive to the frequencies of the observed alleles. Siblings of mistyped individuals have elevated false-positive rates, as do markers close to the mistyped marker. We conclude that thresholds should be decided based on the pedigree and marker data and that greater focus should be placed on modeling genotyping error when computing likelihoods, rather than on detecting and eliminating genotyping errors.  相似文献   

2.
The identification of genes contributing to complex diseases and quantitative traits requires genetic data of high fidelity, because undetected errors and mutations can profoundly affect linkage information. The recent emphasis on the use of the sibling-pair design eliminates or decreases the likelihood of detection of genotyping errors and marker mutations through apparent Mendelian incompatibilities or close double recombinants. In this article, we describe a hidden Markov method for detecting genotyping errors and mutations in multilocus linkage data. Specifically, we calculate the posterior probability of genotyping error or mutation for each sibling-pair-marker combination, conditional on all marker data and an assumed genotype-error rate. The method is designed for use with sibling-pair data when parental genotypes are unavailable. Through Monte Carlo simulation, we explore the effects of map density, marker-allele frequencies, marker position, and genotype-error rate on the accuracy of our error-detection method. In addition, we examine the impact of genotyping errors and error detection and correction on multipoint linkage information. We illustrate that even moderate error rates can result in substantial loss of linkage information, given efforts to fine-map a putative disease locus. Although simulations suggest that our method detects 相似文献   

3.
It is well known that genotyping errors lead to loss of power in gene-mapping studies and underestimation of the strength of correlations between trait- and marker-locus genotypes. In two-point linkage analysis, these errors can be absorbed in an inflated recombination-fraction estimate, leaving the test statistic quite robust. In multipoint analysis, however, genotyping errors can easily result in false exclusion of the true location of a disease-predisposing gene. In a companion article, we described a "complex-valued" extension of the recombination fraction to accommodate errors in the assignment of trait-locus genotypes, leading to a multipoint LOD score with the same robustness to errors in trait-locus genotypes that is seen with the conventional two-point LOD score. Here, a further extension of this model to "hypercomplex-valued" recombination fractions (hereafter referred to as "hypercomplex recombination fractions") is presented, to handle random and systematic sources of marker-locus genotyping errors. This leads to a multipoint method (either "model-based" or "model-free") with the same robustness to marker-locus genotyping errors that is seen with conventional two-point analysis but with the advantage that multiple marker loci can be used jointly to increase meiotic informativeness. The cost of this increased robustness is a decrease in fine-scale resolution of the estimated map location of the trait locus, in comparison with traditional multipoint analysis. This probability model further leads to algorithms for the estimation of the lower bounds for the error rates for genomewide and locus-specific genotyping, based on the null-hypothesis distribution of the LOD-score statistic in the presence of such errors. It is argued that those genome scans in which the LOD score is 0 for >50% of the genome are likely to be characterized by high rates of genotyping errors in general.  相似文献   

4.
Determining population sizes can be difficult, but is essential for conservation. By counting distinct microsatellite genotypes, DNA from noninvasive samples (hair, faeces) allows estimation of population size. Problems arise because genotypes from noninvasive samples are error-prone, but genotyping errors can be reduced by multiple polymerase chain reaction (PCR). For faecal genotypes from wolves in Yellowstone National Park, error rates varied substantially among samples, often above the 'worst-case threshold' suggested by simulation. Consequently, a substantial proportion of multilocus genotypes held one or more errors, despite multiple PCR. These genotyping errors created several genotypes per individual and caused overestimation (up to 5.5-fold) of population size. We propose a 'matching approach' to eliminate this overestimation bias.  相似文献   

5.
6.
Genotyping errors are present in almost all genetic data and can affect biological conclusions of a study, particularly for studies based on individual identification and parentage. Many statistical approaches can incorporate genotyping errors, but usually need accurate estimates of error rates. Here, we used a new microsatellite data set developed for brown rockfish (Sebastes auriculatus) to estimate genotyping error using three approaches: (i) repeat genotyping 5% of samples, (ii) comparing unintentionally recaptured individuals and (iii) Mendelian inheritance error checking for known parent–offspring pairs. In each data set, we quantified genotyping error rate per allele due to allele drop‐out and false alleles. Genotyping error rate per locus revealed an average overall genotyping error rate by direct count of 0.3%, 1.5% and 1.7% (0.002, 0.007 and 0.008 per allele error rate) from replicate genotypes, known parent–offspring pairs and unintentionally recaptured individuals, respectively. By direct‐count error estimates, the recapture and known parent–offspring data sets revealed an error rate four times greater than estimated using repeat genotypes. There was no evidence of correlation between error rates and locus variability for all three data sets, and errors appeared to occur randomly over loci in the repeat genotypes, but not in recaptures and parent–offspring comparisons. Furthermore, there was no correlation in locus‐specific error rates between any two of the three data sets. Our data suggest that repeat genotyping may underestimate true error rates and may not estimate locus‐specific error rates accurately. We therefore suggest using methods for error estimation that correspond to the overall aim of the study (e.g. known parent–offspring comparisons in parentage studies).  相似文献   

7.
Errors while genotyping are inevitable and can reduce the power to detect linkage. However, does genotyping error have the same impact on linkage results for single-nucleotide polymorphism (SNP) and microsatellite (MS) marker maps? To evaluate this question we detected genotyping errors that are consistent with Mendelian inheritance using large changes in multipoint identity-by-descent sharing in neighboring markers. Only a small fraction of Mendelian consistent errors were detectable (e.g., 18% of MS and 2.4% of SNP genotyping errors). More SNP genotyping errors are Mendelian consistent compared to MS genotyping errors, so genotyping error may have a greater impact on linkage results using SNP marker maps. We also evaluated the effect of genotyping error on the power and type I error rate using simulated nuclear families with missing parents under 0, 0.14, and 2.8% genotyping error rates. In the presence of genotyping error, we found that the power to detect a true linkage signal was greater for SNP (75%) than MS (67%) marker maps, although there were also slightly more false-positive signals using SNP marker maps (5 compared with 3 for MS). Finally, we evaluated the usefulness of accounting for genotyping error in the SNP data using a likelihood-based approach, which restores some of the power that is lost when genotyping error is introduced.  相似文献   

8.
Microsatellite markers are important tools in population, conservation and forensic studies and are frequently used for species delineation, the detection of hybridization and introgression. Therefore, marker sets that amplify variable DNA regions in two species are required; however, cross-species amplification is often difficult, as genotyping errors such as null alleles may occur. To estimate the level of potential misidentifications based on genotyping errors, we compared the occurrence of parental alleles in laboratory and natural Daphnia hybrids (Daphnia longispina group). We tested a set of 12 microsatellite loci with regard to their suitability for unambiguous species and hybrid class identification using F(1) hybrids bred in the laboratory. Further, a large set of 44 natural populations of D. cucullata, D. galeata and D. longispina (1715 individuals) as well as their interspecific hybrids were genotyped to validate the discriminatory power of different marker combinations. Species delineation using microsatellite multilocus genotypes produced reliable results for all three studied species using assignment tests. Daphnia galeata × cucullata hybrid detection was limited due to three loci exhibiting D. cucullata-specific null alleles, which most likely are caused by differences in primer-binding sites of parental species. Overall, discriminatory power in hybrid detection was improved when a subset of markers was identified that amplifies equally well in both species.  相似文献   

9.
The occurrence of clonality in threatened plants can have important implications for their conservation. In this study, allozymes and RAPDs were used to determine the extent of clonality in the endangered shrub Haloragodendron lucasii (Haloragaceae), which is known from only four sites within an 8 km range. Allozyme markers identified only six multilocus genotypes among the 53 ramets sampled across the four sites, although a total of 54 different genotypes were possible with the three polymorphic allozyme loci detected. The polymorphic bands detected in the RAPD analysis were capable of producing 246 genotypes, but again only six multilocus genotypes were delineated. The allozyme and RAPD data were congruent at three of the four sites. At the fourth site two genotypes were detected by each marker; however, once combined, three multilocus genotypes were observed. The probabilities that the observed number of replicates of each combined allozyme and RAPD genotype could be generated by sexual reproduction were less than 10–18, leaving little doubt that clonality is the explanation for the observed patterns of genotypes. The genetic conclusions are supported by root excavations which show potential for vegetative reproduction and the observation of no sexual reproduction in the species. The recognition of extensive clonality in H. lucasii has had immediate implications for the conservation management of the species and resulted in changes to the management priorities for the species. Thus it is clear that appropriate genetic studies can play an important role in the management of threatened species.  相似文献   

10.
ABSTRACT Use of non-invasive sources of DNA, such as hair or scat, to obtain a genetic mark for population estimates is becoming commonplace. Unfortunately, with such marks, potentials for genotyping errors and for the shadow effect have resulted in use of many loci and amplification of each specimen many times at each locus, drastically increasing time and cost of obtaining a population estimate. We proposed a method, the Genotyping Uncertainty Added Variance Adjustment (GUAVA), which statistically adjusts for genotyping errors and the shadow effect, thereby allowing use of fewer loci and one amplification of each specimen per locus. Using allele frequencies and estimates of genotyping error rates, we determined, for each pair of specimens, the probability that the pair was obtained from the same individual, whether or not their observed genotypes match. Using these probabilities, we reconstructed possible capture history matrices and used this distribution to obtain a population estimate. With simulated data, we consistently found our estimates had lower bias and smaller variance than estimates based on single amplifications in which genotyping error was ignored and that were comparable to estimates based on data free of genotyping errors. We also demonstrated the method on a fecal DNA data set from a population of red wolves (Canis rufus). The GUAVA estimate based on only one amplification genotypes compares favorably to the estimate based on consensus genotypes. A program to conduct the analysis is available from the first author for UNIX or Windows platforms. Application of GUAVA may allow for increased accuracy in population estimates at reduced cost.  相似文献   

11.
In the context of parentage assignment using genomic markers, key issues are genotyping errors and an absence of parent genotypes because of sampling, traceability or genotyping problems. Most likelihood‐based parentage assignment software programs require a priori estimates of genotyping errors and the proportion of missing parents to set up meaningful assignment decision rules. We present here the R package APIS, which can assign offspring to their parents without any prior information other than the offspring and parental genotypes, and a user‐defined, acceptable error rate among assigned offspring. Assignment decision rules use the distributions of average Mendelian transmission probabilities, which enable estimates of the proportion of offspring with missing parental genotypes. APIS has been compared to other software (CERVUS, VITASSIGN), on a real European seabass (Dicentrarchus labrax) single nucleotide polymorphism data set. The type I error rate (false positives) was lower with APIS than with other software, especially when parental genotypes were missing, but the true positive rate was also lower, except when the theoretical exclusion power reached 0.99999. In general, APIS provided assignments that satisfied the user‐set acceptable error rate of 1% or 5%, even when tested on simulated data with high genotyping error rates (1% or 3%) and up to 50% missing sires. Because it uses the observed distribution of Mendelian transmission probabilities, APIS is best suited to assigning parentage when numerous offspring (>200) are genotyped. We have demonstrated that APIS is an easy‐to‐use and reliable software for parentage assignment, even when up to 50% of sires are missing.  相似文献   

12.
Although it is clear that errors in genotyping data can lead to severe errors in linkage analysis, there is as yet no consensus strategy for identification of genotyping errors. Strategies include comparison of duplicate samples, independent calling of alleles, and Mendelian-inheritance-error checking. This study aimed to develop a better understanding of error types associated with microsatellite genotyping, as a first step toward development of a rational error-detection strategy. Two microsatellite marker sets (a commercial genomewide set and a custom-designed fine-resolution mapping set) were used to generate 118,420 and 22,500 initial genotypes and 10,088 and 8,328 duplicates, respectively. Mendelian-inheritance errors were identified by PedManager software, and concordance was determined for the duplicate samples. Concordance checking identifies only human errors, whereas Mendelian-inheritance-error checking is capable of detection of additional errors, such as mutations and null alleles. Neither strategy is able to detect all errors. Inheritance checking of the commercial marker data identified that the results contained 0.13% human errors and 0.12% other errors (0.25% total error), whereas concordance checking found 0.16% human errors. Similarly, Mendelian-inheritance-error checking of the custom-set data identified 1.37% errors, compared with 2.38% human errors identified by concordance checking. A greater variety of error types were detected by Mendelian-inheritance-error checking than by duplication of samples or by independent reanalysis of gels. These data suggest that Mendelian-inheritance-error checking is a worthwhile strategy for both types of genotyping data, whereas fine-mapping studies benefit more from concordance checking than do studies using commercial marker data. Maximization of error identification increases the likelihood of linkage when complex diseases are analyzed.  相似文献   

13.
Identifying marker typing incompatibilities in linkage analysis.   总被引:3,自引:3,他引:0       下载免费PDF全文
A common problem encountered in linkage analyses is that execution of the computer program is halted because of genotypes in the data that are inconsistent with Mendelian inheritance. Such inconsistencies may arise because of pedigree errors or errors in typing. In some cases, the source of the inconsistencies is easily identified by examining the pedigree. In others, the error is not obvious, and substantial time and effort are required to identify the responsible genotypes. We have developed two methods for automatically identifying those individuals whose genotypes are most likely the cause of the inconsistencies. First, we calculate the posterior probability of genotyping error for each member of the pedigree, given the marker data on all pedigree members and allowing anyone in the pedigree to have an error. Second, we identify those individuals whose genotypes could be solely responsible for the inconsistency in the pedigree. We illustrate these methods with two examples: one a pedigree error, the second a genotyping error. These methods have been implemented as a module of the pedigree analysis program package MENDEL.  相似文献   

14.
We obtained fresh dung samples from 202 (133 mother-offspring pairs) savannah elephants (Loxodonta africana) in Samburu, Kenya, and genotyped them at 20 microsatellite loci to assess genotyping success and errors. A total of 98.6% consensus genotypes was successfully obtained, with allelic dropout and false allele rates at 1.6% (n = 46) and 0.9% (n = 37) of heterozygous and total consensus genotypes, respectively, and an overall genotyping error rate of 2.5% based on repeat typing. Mendelian analysis revealed consistent inheritance in all but 38 allelic pairs from mother-offspring, giving an average mismatch error rate of 2.06%, a possible result of null alleles, mutations, genotyping errors, or inaccuracy in maternity assignment. We detected no evidence for large allele dropout, stuttering, or scoring error in the dataset and significant Hardy-Weinberg deviations at only two loci due to heterozygosity deficiency. Across loci, null allele frequencies were low (range: 0.000-0.042) and below the 0.20 threshold that would significantly bias individual-based studies. The high genotyping success and low errors observed in this study demonstrate reliability of the method employed and underscore the application of simple pedigrees in noninvasive studies. Since none of the sires were included in this study, the error rates presented are just estimates.  相似文献   

15.
Errors in genotype calling can have perverse effects on genetic analyses, confounding association studies, and obscuring rare variants. Analyses now routinely incorporate error rates to control for spurious findings. However, reliable estimates of the error rate can be difficult to obtain because of their variance between studies. Most studies also report only a single estimate of the error rate even though genotypes can be miscalled in more than one way. Here, we report a method for estimating the rates at which different types of genotyping errors occur at biallelic loci using pedigree information. Our method identifies potential genotyping errors by exploiting instances where the haplotypic phase has not been faithfully transmitted. The expected frequency of inconsistent phase depends on the combination of genotypes in a pedigree and the probability of miscalling each genotype. We develop a model that uses the differences in these frequencies to estimate rates for different types of genotype error. Simulations show that our method accurately estimates these error rates in a variety of scenarios. We apply this method to a dataset from the whole-genome sequencing of owl monkeys (Aotus nancymaae) in three-generation pedigrees. We find significant differences between estimates for different types of genotyping error, with the most common being homozygous reference sites miscalled as heterozygous and vice versa. The approach we describe is applicable to any set of genotypes where haplotypic phase can reliably be called and should prove useful in helping to control for false discoveries.  相似文献   

16.
Because current molecular haplotyping methods are expensive and not amenable to automation, many researchers rely on statistical methods to infer haplotype pairs from multilocus genotypes, and subsequently treat these inferred haplotype pairs as observations. These procedures are prone to haplotype misclassification. We examine the effect of these misclassification errors on the false-positive rate and power for two association tests. These tests include the standard likelihood ratio test (LRTstd) and a likelihood ratio test that employs a double-sampling approach to allow for the misclassification inherent in the haplotype inference procedure (LRTae). We aim to determine the cost-benefit relationship of increasing the proportion of individuals with molecular haplotype measurements in addition to genotypes to raise the power gain of the LRTae over the LRTstd. This analysis should provide a guideline for determining the minimum number of molecular haplotypes required for desired power. Our simulations under the null hypothesis of equal haplotype frequencies in cases and controls indicate that (1) for each statistic, permutation methods maintain the correct type I error; (2) specific multilocus genotypes that are misclassified as the incorrect haplotype pair are consistently misclassified throughout each entire dataset; and (3) our simulations under the alternative hypothesis showed a significant power gain for the LRTae over the LRTstd for a subset of the parameter settings. Permutation methods should be used exclusively to determine significance for each statistic. For fixed cost, the power gain of the LRTae over the LRTstd varied depending on the relative costs of genotyping, molecular haplotyping, and phenotyping. The LRTae showed the greatest benefit over the LRTstd when the cost of phenotyping was very high relative to the cost of genotyping. This situation is likely to occur in a replication study as opposed to a whole-genome association study.  相似文献   

17.
In recent years, numerous outbreaks of multidrug-resistant Pseudomonas aeruginosa have been reported across the world. Once an outbreak occurs, besides routinely testing isolates for susceptibility to antimicrobials, it is required to check their virulence genotypes and clonality profiles. Replacing pulsed-field gel electrophoresis DNA fingerprinting are faster, easier-to-use, and less expensive polymerase chain reaction (PCR)-based methods for characterizing hospital isolates. P. aeruginosa possesses a mosaic genome structure and a highly conserved core genome displaying low sequence diversity and a highly variable accessory genome that communicates with other Pseudomonas species via horizontal gene transfer. Multiple-locus variable-number tandem-repeat analysis and multilocus sequence typing methods allow for phylogenetic analysis of isolates by PCR amplification of target genes with the support of Internet-based services. The target genes located in the core genome regions usually contain low-frequency mutations, allowing the resulting phylogenetic trees to infer evolutionary processes. The multiplex PCR-based open reading frame typing (POT) method, integron PCR, and exoenzyme genotyping can determine a genotype by PCR amplifying a specific insertion gene in the accessory genome region using a single or a multiple primer set. Thus, analyzing P. aeruginosa isolates for their clonality, virulence factors, and resistance characteristics is achievable by combining the clonality evaluation of the core genome based on multiple-locus targeting methods with other methods that can identify specific virulence and antimicrobial genes. Software packages such as eBURST, R, and Dendroscope, which are powerful tools for phylogenetic analyses, enable researchers and clinicians to visualize clonality associations in clinical isolates.  相似文献   

18.
Moskvina V  Schmidt KM 《Biometrics》2006,62(4):1116-1123
With the availability of fast genotyping methods and genomic databases, the search for statistical association of single nucleotide polymorphisms with a complex trait has become an important methodology in medical genetics. However, even fairly rare errors occurring during the genotyping process can lead to spurious association results and decrease in statistical power. We develop a systematic approach to study how genotyping errors change the genotype distribution in a sample. The general M-marker case is reduced to that of a single-marker locus by recognizing the underlying tensor-product structure of the error matrix. Both method and general conclusions apply to the general error model; we give detailed results for allele-based errors of size depending both on the marker locus and the allele present. Multiple errors are treated in terms of the associated diffusion process on the space of genotype distributions. We find that certain genotype and haplotype distributions remain unchanged under genotyping errors, and that genotyping errors generally render the distribution more similar to the stable one. In case-control association studies, this will lead to loss of statistical power for nondifferential genotyping errors and increase in type I error for differential genotyping errors. Moreover, we show that allele-based genotyping errors do not disturb Hardy-Weinberg equilibrium in the genotype distribution. In this setting we also identify maximally affected distributions. As they correspond to situations with rare alleles and marker loci in high linkage disequilibrium, careful checking for genotyping errors is advisable when significant association based on such alleles/haplotypes is observed in association studies.  相似文献   

19.
In non‐model organisms, evolutionary questions are frequently addressed using reduced representation sequencing techniques due to their low cost, ease of use, and because they do not require genomic resources such as a reference genome. However, evidence is accumulating that such techniques may be affected by specific biases, questioning the accuracy of obtained genotypes, and as a consequence, their usefulness in evolutionary studies. Here, we introduce three strategies to estimate genotyping error rates from such data: through the comparison to high quality genotypes obtained with a different technique, from individual replicates, or from a population sample when assuming Hardy‐Weinberg equilibrium. Applying these strategies to data obtained with Restriction site Associated DNA sequencing (RAD‐seq), arguably the most popular reduced representation sequencing technique, revealed per‐allele genotyping error rates that were much higher than sequencing error rates, particularly at heterozygous sites that were wrongly inferred as homozygous. As we exemplify through the inference of genome‐wide and local ancestry of well characterized hybrids of two Eurasian poplar (Populus) species, such high error rates may lead to wrong biological conclusions. By properly accounting for these error rates in downstream analyses, either by incorporating genotyping errors directly or by recalibrating genotype likelihoods, we were nevertheless able to use the RAD‐seq data to support biologically meaningful and robust inferences of ancestry among Populus hybrids. Based on these findings, we strongly recommend carefully assessing genotyping error rates in reduced representation sequencing experiments, and to properly account for these in downstream analyses, for instance using the tools presented here.  相似文献   

20.
The purpose of this work is to quantify the effects that errors in genotyping have on power and the sample size necessary to maintain constant asymptotic Type I and Type II error rates (SSN) for case-control genetic association studies between a disease phenotype and a di-allelic marker locus, for example a single nucleotide polymorphism (SNP) locus. We consider the effects of three published models of genotyping errors on the chi-square test for independence in the 2 x 3 table. After specifying genotype frequencies for the marker locus conditional on disease status and error model in both a genetic model-based and a genetic model-free framework, we compute the asymptotic power to detect association through specification of the test's non-centrality parameter. This parameter determines the functional dependence of SSN on the genotyping error rates. Additionally, we study the dependence of SSN on linkage disequilibrium (LD), marker allele frequencies, and genotyping error rates for a dominant disease model. Increased genotyping error rate requires a larger SSN. Every 1% increase in sum of genotyping error rates requires that both case and control SSN be increased by 2-8%, with the extent of increase dependent upon the error model. For the dominant disease model, SSN is a nonlinear function of LD and genotyping error rate, with greater SSN for lower LD and higher genotyping error rate. The combination of lower LD and higher genotyping error rates requires a larger SSN than the sum of the SSN for the lower LD and for the higher genotyping error rate.  相似文献   

设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号