circular DNA virus genomes

New publication provides novel method for analysis of circular DNA virus genomes

A recent publication by VIRTIGATION partners KU Leuven and EMWEB proposes a novel method for in-depth analysis of circular DNA virus genomes.

Article by Marianna Granatier from the VIRTIGATION project

Background and aims of the study

The study, published recently in the prestigious Nucleic Acids Research (NAR) journal, presents a new pipeline design that accurately identifies and characterises full-length sequence variants of viral circular DNA genomes, using Nanopore sequencing technology and the Genome Detective bioinformatics tool. Project researchers outlined the basis for this publication also in the VIRTIGATION Deliverable D2.3.

Current challenges in detecting rare variants in circular DNA plant virus genomes

Currently, researchers sequence entire virus genomes and discover new viral species via second-generation high-throughput sequencing (SGS) technologies, such as Illumina. The SGS produces reads between 50 and 300 base pairs. However, it limits its ability to deeply profile full-length virus genomes, which is particularly problematic for viruses that exist as complex populations of closely related sequences (i.e. quasispecies) where accurate assembly from short reads is difficult.

In contrast, third-generation sequencing platforms such as Oxford Nanopore (ONT) offer long reads that are better suited for reconstructing full-length haplotypes of these quasispecies. Recent improvements even allow for high-quality reads with 99.9% accuracy and minimal bias from deletions or indels. Nevertheless, a major challenge in using such data for profiling genetic variants and mutations in viruses such as Geminiviridae and Nanoviridae (most significant circular DNA virus families affecting plants) is distinguishing true mutations from sequencing errors. Although some bioinformatics tools exist to reconstruct virus haplotypes from low-accuracy TGS reads, their sensitivity is insufficient to detect haplotypes with low abundance (1%) and low diversity (<0.3%).

How the new VIRTIGATION sequencing Method reveals hidden plant virus variants

To overcome existing limitations, VIRTIGATION partners KU Leuven and EMWEB propose in their NAR publication a new pipeline that generates and profiles high-quality full genome sequences of circular ssDNA virus variants with single nucleotide differences and a relative abundance of at least 1%. Their study highlights the value of long-read sequencing technologies for in-depth analysis of plant virus populations (monopartite Tomato yellow leaf curl Sardinia virus (TYLCSV) and multipartite Banana bunchy top virus (BBTV). By enriching circular DNA viruses and generating high-quality sequencing data, their pipeline successfully reconstructed full-length genomic sequences of both monopartite and multipartite viruses.

Firstly, project researchers sequenced viral circular DNAs enriched from total DNA samples using the Nanopore sequencing kit (R10.4.1). They then deconcatenated the resulting long RCA (rolling circle amplification) viral reads using the TideHunter tool to yield highly accurate, intact full-genome viral sequences. The scientists then clustered the sequences by similarity, and corrected and quantified random errors using the Genome Detective online tool. Within the pipeline, Genome Detective facilitates deconcatenation of RCA raw or duplex data, correction, and annotation of reads to produce high-quality full-genome sequences and their relative abundances.

VIRTIGATION Dengue
The Genome Detective software of VIRTIGATION partner EMWEB © 2024 EMWEB

Genome Detective enables accurate reconstruction of original virus clone sequences

The use of duplex reads significantly enhances the reliability and accuracy of the analysis, as demonstrated by the detection of authentic TYLCSV variants confirmed by SGS. However, using raw reads led to numerous false positives. Since the proportion of duplex reads can vary across libraries, total reads may be necessary to detect low-abundance haplotypes in samples with few duplex reads, though independent validation is required in such cases. The error-correction step in Genome Detective also enables accurate reconstruction of original virus clone sequences without indels or deletions in homopolymeric regions, although challenges in sequencing these regions with TGS remain and need careful review.

Using field samples of a multipartite virus, the KU Leuven and EMWEB pipeline effectively identified variants from different genomic components and enabled the clear detection of partial genomic components. These partial sequences appear to originate from the most abundant TYLCSV and BBTV variants. Although VIRTIGATION researchers note in their publication that their functions is not yet fully understood, they seem to influence infection and contribute to viral genetic diversity.

DNA RNA plant viruses
KU Leuven and EMWEB pipeline for full genome sequencing of circular DNA plant viruses © Victor Golyaev & Sam Dierickx

A new method to help study virus adaptation across environments

Overall, the ability of KU Leuven and EMWEB’s pipeline to accurately profile circular DNA virus populations, including low-abundance variants and partial genomic components, offers significant potential for monitoring virus emergence and understanding the complex dynamics of virus populations in relation to host genetics and environmental factors. Further analysis of emerging virus populations under varying conditions (e.g. time, geography, host plant, environment) using this pipeline could provide insights into viral evolution and adaptation to host antiviral responses or environmental changes. Additionally, the precise identification of virus variants could support the isolation of mild variants for use in cross-protection strategies against more severe parental strains.

More info about KU Leuven and EMWEB's new study in NAR

The full version of the study in Nucleic Acids Research (NAR) titled ‘A method for in-depth analysis of circular DNA virus populations by unambiguously profiling the low abundant virus variants and partial genomic components’, authored by KU Leuven and EMWEB, is available online here since 2 April 2025. The paper has been written by some of KU Leuven and EMWEB principal investigators in VIRTIGATION: Hervé Vanderschuren and Victor Golyaev (both KU Leuven), and Wim Dumon, Koen Deforche and Sam Dierickx (all EMWEB). The datasets underlying their peer-reviewed, open access scientific publication are available here. Find out more about VIRTIGATION’s scientific publications on our website here.

© Cover image: Shutterstock