首页 | 本学科首页   官方微博 | 高级检索  
相似文献
 共查询到20条相似文献,搜索用时 15 毫秒
1.
2.

Background

Transposable elements form a significant proportion of eukaryotic genomes. Recently, Lexa et al. (Nucleic Acids Res 42:968-978, 2014) reported that plant long terminal repeat (LTR) retrotransposons often contain potential quadruplex sequences (PQSs) in their LTRs and experimentally confirmed their ability to adopt four-stranded DNA conformations.

Results

Here, we searched for PQSs in human retrotransposons and found that PQSs are specifically localized in the 3’-UTR of LINE-1 elements, in LTRs of HERV elements and are strongly accumulated in specific regions of SVA elements. Circular dichroism spectroscopy confirmed that most PQSs had adopted monomolecular or bimolecular guanine quadruplex structures. Evolutionarily young SVA elements contained more PQSs than older elements and their propensity to form quadruplex DNA was higher. Full-length L1 elements contained more PQSs than truncated elements; the highest proportion of PQSs was found inside transpositionally active L1 elements (PA2 and HS families).

Conclusions

Conservation of quadruplexes at specific positions of transposable elements implies their importance in their life cycle. The increasing quadruplex presence in evolutionarily young LINE-1 and SVA families makes these elements important contributors toward present genome-wide quadruplex distribution.

Electronic supplementary material

The online version of this article (doi:10.1186/1471-2164-15-1032) contains supplementary material, which is available to authorized users.  相似文献   

3.

Background

Prophages, integral components of many bacterial genomes, play significant roles in cognate host bacteria, such as virulence, toxin biosynthesis and secretion, fitness cost, genomic variations, and evolution. Many prophages and prophage-like elements present in sequenced bacterial genomes, such as Bifidobacteria, Lactococcus and Streptococcus, have been described. However, information for the prophage of Mycobacterium remains poorly defined.

Results

In this study, based on the search of the complete genome database from GenBank, the Whole Genome Shotgun (WGS) databases, and some published literatures, thirty-three prophages were described in detail. Eleven of them were full-length prophages, and others were prophage-like elements. Eleven prophages were firstly revealed. They were phiMAV_1, phiMAV_2, phiMmcs_1, phiMmcs_2, phiMkms_1, phiMkms_2, phiBN42_1, phiBN44_1, phiMCAN_1, phiMycsm_1, and phiW7S_1. Their genomes and gene contents were firstly analyzed. Furthermore, comparative genomics analyses among mycobacterioprophages showed that full-length prophage phi172_2 belonged to mycobacteriophage Cluster A and the phiMmcs_1, phiMkms_1, phiBN44_1, and phiMCAN_1 shared high homology and could be classified into one group.

Conclusions

To our knowledge, this is the first systematic characterization of mycobacterioprophages, their genomic organization and phylogeny. This information will afford more understanding of the biology of Mycobacterium.

Electronic supplementary material

The online version of this article (doi:10.1186/1471-2164-15-243) contains supplementary material, which is available to authorized users.  相似文献   

4.
5.

Background

Analyzing the integration profile of retroviral vectors is a vital step in determining their potential genotoxic effects and developing safer vectors for therapeutic use. Identifying retroviral vector integration sites is also important for retroviral mutagenesis screens.

Results

We developed VISA, a vector integration site analysis server, to analyze next-generation sequencing data for retroviral vector integration sites. Sequence reads that contain a provirus are mapped to the human genome, sequence reads that cannot be localized to a unique location in the genome are filtered out, and then unique retroviral vector integration sites are determined based on the alignment scores of the remaining sequence reads.

Conclusions

VISA offers a simple web interface to upload sequence files and results are returned in a concise tabular format to allow rapid analysis of retroviral vector integration sites.

Electronic supplementary material

The online version of this article (doi:10.1186/s12859-015-0653-6) contains supplementary material, which is available to authorized users.  相似文献   

6.

Background

The siRNA and piRNA pathways have been shown in insects to be essential for regulation of gene expression and defence against exogenous and endogenous genetic elements (viruses and transposable elements). The vast majority of endogenous small RNAs produced by the siRNA and piRNA pathways originate from repetitive or transposable elements (TE). In D. melanogaster, TE-derived endogenous siRNAs and piRNAs are involved in genome surveillance and maintenance of genome integrity. In the medically relevant malaria mosquito Anopheles gambiae TEs constitute 12-16% of the genome size. Genetic variations induced by TE activities are known to shape the genome landscape and to alter the fitness in An. gambiae.

Results

Here, using bioinformatics approaches we analyzed the small RNA data sets from 6 libraries formally reported in a previous study and examined the expression of the mixed germline/somatic siRNAs and piRNAs produced in adult An. gambiae females. We characterized a large population of TE-derived endogenous siRNAs and piRNAs, which constitutes 56-60% of the total siRNA and piRNA reads in the analysed libraries. Moreover, we identified a number of protein coding genes producing gene-specific siRNAs and piRNAs that were generally expressed at much lower levels than the TE-associated small RNAs. Detailed sequence analysis revealed that An. gambiae piRNAs were produced by both “ping-pong” dependent (TE-associated piRNAs) and independent mechanisms (genic piRNAs). Similarly to D. melanogaster, more than 90% of the detected piRNAs were produced from TE-associated clusters in An. gambiae. We also found that biotic stress as blood feeding and infection with Plasmodium parasite, the etiological agent of malaria, modulated the expression levels of the endogenous siRNAs and piRNAs in An. gambiae.

Conclusions

We identified a large and diverse set of the endogenously derived siRNAs and piRNAs that share common and distinct aspects of small RNA expression across insect species, and inferred their impact on TE and gene activity in An. gambiae. The TE-specific small RNAs produced by both the siRNA and piRNA pathways represent an important aspect of genome stability and genetic variation, which might have a strong impact on the evolution of the genome and vector competence in the malaria mosquitoes.

Electronic supplementary material

The online version of this article (doi:10.1186/s12864-015-1436-1) contains supplementary material, which is available to authorized users.  相似文献   

7.

Background

Mobile elements are active in the human genome, both in the germline and cancers, where they can mutate driver genes.

Results

While analysing whole genome paired-end sequencing of oesophageal adenocarcinomas to find genomic rearrangements, we identified three ways in which new mobile element insertions appear in the data, resembling translocation or insertion junctions: inserts where unique sequence has been transduced by an L1 (Long interspersed element 1) mobile element; novel inserts that are confidently, but often incorrectly, mapped by alignment software to L1s or polyA tracts in the reference sequence; and a combination of these two ways, where different sequences within one insert are mapped to different loci. We identified nine unique sequences that were transduced by neighbouring L1s, both L1s in the reference genome and L1s not present in the reference. Many of the resulting inserts were small fragments that include little or no recognisable mobile element sequence. We found 6 loci in the reference genome to which sequence reads from inserts were frequently mapped, probably erroneously, by alignment software: these were either L1 sequence or particularly long polyA runs. Inserts identified from such apparent rearrangement junctions averaged 16 inserts/tumour, range 0–153 insertions in 43 tumours. However, many inserts would not be detected by mapping the sequences to the reference genome, because they do not include sufficient mappable sequence. To estimate total somatic inserts we searched for polyA sequences that were not present in the matched normal or other normals from the same tumour batch, and were not associated with known polymorphisms. Samples of these candidate inserts were verified by sequencing across them or manual inspection of surrounding reads: at least 85 % were somatic and resembled L1-mediated events, most including L1Hs sequence. Approximately 100 such inserts were detected per tumour on average (range zero to approximately 700).

Conclusions

Somatic mobile elements insertions are abundant in these tumours, with over 75 % of cases having a number of novel inserts detected. The inserts create a variety of problems for the interpretation of paired-end sequencing data.

Electronic supplementary material

The online version of this article (doi:10.1186/s12864-015-1685-z) contains supplementary material, which is available to authorized users.  相似文献   

8.

Background

The publicly available Laccaria bicolor genome sequence has provided a considerable genomic resource allowing systematic identification of transposable elements (TEs) in this symbiotic ectomycorrhizal fungus. Using a TE-specific annotation pipeline we have characterized and analyzed TEs in the L. bicolor S238N-H82 genome.

Methodology/Principal Findings

TEs occupy 24% of the 60 Mb L. bicolor genome and represent 25,787 full-length and partial copy elements distributed within 171 families. The most abundant elements were the Copia-like. TEs are not randomly distributed across the genome, but are tightly nested or clustered. The majority of TEs exhibits signs of ancient transposition except some intact copies of terminal inverted repeats (TIRS), long terminal repeats (LTRs) and a large retrotransposon derivative (LARD) element. There were three main periods of TE expansion in L. bicolor: the first from 57 to 10 Mya, the second from 5 to 1 Mya and the most recent from 0.5 Mya ago until now. LTR retrotransposons are closely related to retrotransposons found in another basidiomycete, Coprinopsis cinerea.

Conclusions

This analysis 1) represents an initial characterization of TEs in the L. bicolor genome, 2) contributes to improve genome annotation and a greater understanding of the role TEs played in genome organization and evolution and 3) provides a valuable resource for future research on the genome evolution within the Laccaria genus.  相似文献   

9.

Background

Piwi-interacting RNAs (piRNAs) are a recently discovered class of small non-coding RNAs whose best-understood function is to repress mobile element (ME) activity in animal germline. To date, nearly all piRNA studies have been conducted in model organisms and little is known about piRNA diversity, target specificity and biological function in human.

Results

Here we performed high-throughput sequencing of piRNAs from three human adult testis samples. We found that more than 81% of the ~17 million putative piRNAs mapped to ~6,000 piRNA-producing genomic clusters using a relaxed definition of clusters. A set of human protein-coding genes produces a relatively large amount of putative piRNAs from their 3’UTRs, and are significantly enriched for certain biological processes, suggestive of non-random sampling by the piRNA biogenesis machinery. Up to 16% of putative piRNAs mapped to a few hundred annotated long non-coding RNA (lncRNA) genes, suggesting that some lncRNA genes can act as piRNA precursors. Among major ME families, young families of LTR and endogenous retroviruses have a greater association with putative piRNAs than other MEs. In addition, piRNAs preferentially mapped to specific regions in the consensus sequences of several ME (sub)families and some piRNA mapping peaks showed patterns consistent with the “ping-pong” cycle of piRNA targeting and amplification.

Conclusions

Overall our data provide a comprehensive analysis and improved annotation of human piRNAs in adult human testes and shed new light into the relationship of piRNAs with protein-coding genes, lncRNAs, and mobile genetic elements in human.

Electronic supplementary material

The online version of this article (doi:10.1186/1471-2164-15-545) contains supplementary material, which is available to authorized users.  相似文献   

10.
11.

Background

The 17 Gb bread wheat genome has massively expanded through the proliferation of transposable elements (TEs) and two recent rounds of polyploidization. The assembly of a 774 Mb reference sequence of wheat chromosome 3B provided us with the opportunity to explore the impact of TEs on the complex wheat genome structure and evolution at a resolution and scale not reached so far.

Results

We develop an automated workflow, CLARI-TE, for TE modeling in complex genomes. We delineate precisely 56,488 intact and 196,391 fragmented TEs along the 3B pseudomolecule, accounting for 85% of the sequence, and reconstruct 30,199 nested insertions. TEs have been mostly silent for the last one million years, and the 3B chromosome has been shaped by a succession of bursts that occurred between 1 to 3 million years ago. Accelerated TE elimination in the high-recombination distal regions is a driving force towards chromosome partitioning. CACTAs overrepresented in the high-recombination distal regions are significantly associated with recently duplicated genes. In addition, we identify 140 CACTA-mediated gene capture events with 17 genes potentially created by exon shuffling and show that 19 captured genes are transcribed and under selection pressure, suggesting the important role of CACTAs in the recent wheat adaptation.

Conclusion

Accurate TE modeling uncovers the dynamics of TEs in a highly complex and polyploid genome. It provides novel insights into chromosome partitioning and highlights the role of CACTA transposons in the high level of gene duplication in wheat.

Electronic supplementary material

The online version of this article (doi:10.1186/s13059-014-0546-4) contains supplementary material, which is available to authorized users.  相似文献   

12.
13.
14.

Background

Like other structural variants, transposable element insertions can be highly polymorphic across individuals. Their functional impact, however, remains poorly understood. Current genome-wide approaches for genotyping insertion-site polymorphisms based on targeted or whole-genome sequencing remain very expensive and can lack accuracy, hence new large-scale genotyping methods are needed.

Results

We describe a high-throughput method for genotyping transposable element insertions and other types of structural variants that can be assayed by breakpoint PCR. The method relies on next-generation sequencing of multiplex, site-specific PCR amplification products and read count-based genotype calls. We show that this method is flexible, efficient (it does not require rounds of optimization), cost-effective and highly accurate.

Conclusions

This method can benefit a wide range of applications from the routine genotyping of animal and plant populations to the functional study of structural variants in humans.

Electronic supplementary material

The online version of this article (doi:10.1186/s12864-015-1700-4) contains supplementary material, which is available to authorized users.  相似文献   

15.

Background

Galileo is one of three members of the P superfamily of DNA transposons. It was originally discovered in Drosophila buzzatii, in which three segregating chromosomal inversions were shown to have been generated by ectopic recombination between Galileo copies. Subsequently, Galileo was identified in six of 12 sequenced Drosophila genomes, indicating its widespread distribution within this genus. Galileo is strikingly abundant in Drosophila willistoni, a neotropical species that is highly polymorphic for chromosomal inversions, suggesting a role for this transposon in the evolution of its genome.

Results

We carried out a detailed characterization of all Galileo copies present in the D. willistoni genome. A total of 191 copies, including 133 with two terminal inverted repeats (TIRs), were classified according to structure in six groups. The TIRs exhibited remarkable variation in their length and structure compared to the most complete copy. Three copies showed extended TIRs due to internal tandem repeats, the insertion of other transposable elements (TEs), or the incorporation of non-TIR sequences into the TIRs. Phylogenetic analyses of the transposase (TPase)-encoding and TIR segments yielded two divergent clades, which we termed Galileo subfamilies V and W. Target-site duplications (TSDs) in D. willistoni Galileo copies were 7- or 8-bp in length, with the consensus sequence GTATTAC. Analysis of the region around the TSDs revealed a target site motif (TSM) with a 15-bp palindrome that may give rise to a stem-loop secondary structure.

Conclusions

There is a remarkable abundance and diversity of Galileo copies in the D. willistoni genome, although no functional copies were found. The TIRs in particular have a dynamic structure and extend in different ways, but their ends (required for transposition) are more conserved than the rest of the element. The D. willistoni genome harbors two Galileo subfamilies (V and W) that diverged ~9 million years ago and may have descended from an ancestral element in the genome. Galileo shows a significant insertion preference for a 15-bp palindromic TSM.

Electronic supplementary material

The online version of this article (doi:10.1186/1471-2164-15-792) contains supplementary material, which is available to authorized users.  相似文献   

16.

Background

Comparative evolutionary analysis of whole genomes requires not only accurate annotation of gene space, but also proper annotation of the repetitive fraction which is often the largest component of most if not all genomes larger than 50 kb in size.

Results

Here we present the Rice TE database (RiTE-db) - a genus-wide collection of transposable elements and repeated sequences across 11 diploid species of the genus Oryza and the closely-related out-group Leersia perrieri. The database consists of more than 170,000 entries divided into three main types: (i) a classified and curated set of publicly-available repeated sequences, (ii) a set of consensus assemblies of highly-repetitive sequences obtained from genome sequencing surveys of 12 species; and (iii) a set of full-length TEs, identified and extracted from 12 whole genome assemblies.

Conclusions

This is the first report of a repeat dataset that spans the majority of repeat variability within an entire genus, and one that includes complete elements as well as unassembled repeats. The database allows sequence browsing, downloading, and similarity searches. Because of the strategy adopted, the RiTE-db opens a new path to unprecedented direct comparative studies that span the entire nuclear repeat content of 15 million years of Oryza diversity.

Electronic supplementary material

The online version of this article (doi:10.1186/s12864-015-1762-3) contains supplementary material, which is available to authorized users.  相似文献   

17.

Background

The community composition of the human microbiome is known to vary at distinct anatomical niches. But little is known about the nature of variations, if any, at the genome/sub-genome levels of a specific microbial community across different niches. The present report aims to explore, as a case study, the variations in gene repertoire of 28 Prevotella reference genomes derived from different body-sites of human, as reported earlier by the Human Microbiome Consortium.

Results

The pan-genome for Prevotella remains “open”. On an average, 17% of predicted protein-coding genes of any particular Prevotella genome represent the conserved core genes, while the remaining 83% contribute to the flexible and singletons. The study reveals exclusive presence of 11798, 3673, 3348 and 934 gene families and exclusive absence of 17, 221, 115 and 645 gene families in Prevotella genomes derived from human oral cavity, gastro-intestinal tracts (GIT), urogenital tract (UGT) and skin, respectively. Distribution of various functional COG categories differs significantly among the habitat-specific genes. No niche-specific variations could be observed in distribution of KEGG pathways.

Conclusions

Prevotella genomes derived from different body sites differ appreciably in gene repertoire, suggesting that these microbiome components might have developed distinct genetic strategies for niche adaptation within the host. Each individual microbe might also have a component of its own genetic machinery for host adaptation, as appeared from the huge number of singletons.

Electronic supplementary material

The online version of this article (doi:10.1186/s12864-015-1350-6) contains supplementary material, which is available to authorized users.  相似文献   

18.

Background

Lactobacillus hokkaidonensis is an obligate heterofermentative lactic acid bacterium, which is isolated from Timothy grass silage in Hokkaido, a subarctic region of Japan. This bacterium is expected to be useful as a silage starter culture in cold regions because of its remarkable psychrotolerance; it can grow at temperatures as low as 4°C. To elucidate its genetic background, particularly in relation to the source of psychrotolerance, we constructed the complete genome sequence of L. hokkaidonensis LOOC260T using PacBio single-molecule real-time sequencing technology.

Results

The genome of LOOC260T comprises one circular chromosome (2.28 Mbp) and two circular plasmids: pLOOC260-1 (81.6 kbp) and pLOOC260-2 (41.0 kbp). We identified diverse mobile genetic elements, such as prophages, integrated and conjugative elements, and conjugative plasmids, which may reflect adaptation to plant-associated niches. Comparative genome analysis also detected unique genomic features, such as genes involved in pentose assimilation and NADPH generation.

Conclusions

This is the first complete genome in the L. vaccinostercus group, which is poorly characterized, so the genomic information obtained in this study provides insight into the genetics and evolution of this group. We also found several factors that may contribute to the ability of L. hokkaidonensis to grow at cold temperatures. The results of this study will facilitate further investigation for the cold-tolerance mechanism of L. hokkaidonensis.

Electronic supplementary material

The online version of this article (doi:10.1186/s12864-015-1435-2) contains supplementary material, which is available to authorized users.  相似文献   

19.

Introduction

The aim of the study was to interrogate the genetic architecture and autoimmune pleiotropy of scleroderma susceptibility in the Australian population.

Methods

We genotyped individuals from a well-characterized cohort of Australian scleroderma patients with the Immunochip, a custom array enriched for single nucleotide polymorphisms (SNPs) at immune loci. Controls were taken from the 1958 British Birth Cohort. After data cleaning and adjusting for population stratification the final dataset consisted of 486 cases, 4,458 controls and 146,525 SNPs. Association analyses were conducted using logistic regression in PLINK. A replication study was performed using 833 cases and 1,938 controls.

Results

A total of eight loci with suggestive association (P <10-4.5) were identified, of which five showed significant association in the replication cohort (HLA-DRB1, DNASE1L3, STAT4, TNP03-IRF5 and VCAM1). The most notable findings were at the DNASE1L3 locus, previously associated with systemic lupus erythematosus, and VCAM1, a locus not previously associated with human disease. This study identified a likely functional variant influencing scleroderma susceptibility at the DNASE1L3 locus; a missense polymorphism rs35677470 in DNASE1L3, with an odds ratio of 2.35 (P = 2.3 × 10−10) in anti-centromere antibody (ACA) positive cases.

Conclusions

This pilot study has confirmed previously reported scleroderma associations, revealed further genetic overlap between scleroderma and systemic lupus erythematosus, and identified a putative novel scleroderma susceptibility locus.

Electronic supplementary material

The online version of this article (doi:10.1186/s13075-014-0438-8) contains supplementary material, which is available to authorized users.  相似文献   

20.

Background

Transposable elements are mobile DNA repeat sequences, known to have high impact on genes, genome structure and evolution. This has stimulated broad interest in the detailed biological studies of transposable elements. Hence, we have developed an easy-to-use tool for the comparative analysis of the structural organization and functional relationships of transposable elements, to help understand their functional role in genomes.

Results

We named our new software VisualTE and describe it here. VisualTE is a JAVA stand-alone graphical interface that allows users to visualize and analyze all occurrences of transposable element families in annotated genomes. VisualTE reads and extracts transposable elements and genomic information from annotation and repeat data. Result analyses are displayed in several graphical panels that include location and distribution on the chromosome, the occurrence of transposable elements in the genome, their size distribution, and neighboring genes’ features and ontologies. With these hallmarks, VisualTE provides a convenient tool for studying transposable element copies and their functional relationships with genes, at the whole-genome scale, and in diverse organisms.

Conclusions

VisualTE graphical interface makes possible comparative analyses of transposable elements in any annotated sequence as well as structural organization and functional relationships between transposable elements and other genetic object. This tool is freely available at: http://lcb.cnrs-mrs.fr/spip.php?article867.

Electronic supplementary material

The online version of this article (doi:10.1186/s12864-015-1351-5) contains supplementary material, which is available to authorized users.  相似文献   

设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号