A K-mer based approach to detect sequence variations and define haplotypes in bread wheat Abstract uri icon

abstract

  • The release of the wheat whole-genome reference sequence has significantly accelerated the association studies between genetic variations and agricultural traits. Currently, the workflow for identifying genetic variations in wheat is derived from those established for rice and maize by mapping the resequencing sequences to the reference genome.

    However, the allohexaploid nature and high sequence similarity among subgenomes of wheat challenge the conventional workflows. Moreover, bread wheat has many wild relatives, which are an important gene pool for expanding its genetic diversity, yet the larger genetic distances between them hampers the identification of sequence variations with conventional workflows.

    This study employs a strategy that directly queries the presence/absence of the K-mers within a given 50 Kb window of the reference genome sequence in resequencing data and records the number of reference-specific K-mers to assess the extent of sequence variations between them.

    Compared to conventional workflows, this strategy:

    Eliminates the need to align resequencing sequences to the reference genome;

    Requires no sufficient sequence alignment coverage, enabling detection of genetic variation between distant genotypes.

    Reduces sequence complexity by partitioning the reference genome into smaller windows.

    Integrates various types of sequence variation into a single value, simplifying downstream data analysis. Building upon this, unsupervised clustering analysis using genetic variation values from consecutive 20 windows (1 Mb) allows for rapid haplotype assignment within the 1 Mb interval.

    Based on the identified genetic variation values and haplotypes, the following applications have been developed:

    Direct identification of introgressions by comparing genetic variation values between hexaploid accessions and the wild relatives;

    Uncover the parental origins for a given progeny genome segment;

    Identification of the donor for a given accession's genome segment within the investigated population, enabling pedigree reconstruction independent of original documents; 4) Develop a haplotype-GWAS model to associate haplotypes and agricultural traits.

publication date

  • September 2024