Understanding the bread wheat genome: from ESTs to graph pangenomes Abstract uri icon

abstract

  • DNA sequencing technology continues to advance and support the development of improved varieties. It has been 20 years since capillary-based Sanger sequencing was applied to a transcriptome resource for wheat functional genomics. This identified genes and where they were expressed, and importantly provided a resource for large scale single nucleotide molecular marker discovery at a time when these markers were starting to replace simple sequence repeat markers.

    Ambitious plans were developed to sequence the whole genome of wheat, with the variety Chinese Spring selected due to the availability of cytogenetic stocks; and chromosome arm specific BAC libraries, sequenced using the Sanger method decided as the approach. As next generation DNA sequencing technology advanced, it was becoming feasible to shotgun assemble genomes using this data. Whole genome sequencing was first attempted using Roche 454 technology, and the sequencing of isolated chromosome arms was attempted using Illumina paired reads. The successful sequencing of the gene content for 7DS and ordering of these genes based on synteny with other grasses in 2011 demonstrated that it was possible to assemble initial sequence-based maps for wheat.

    Chromosome 7DS was quickly followed by 7BS in 2012 and 7AS in 2013, permitting comparison of gene contents between these homoeologous chromosome arms. The general approach was adopted by the IWGSC to sequence the remaining arms to produce the first rough draft of the wheat genome in 2014. As technology developed it became feasible to shotgun the whole genome of bread wheat to provide an improved reference in 2018.

    By 2016 a significant amount of gene presence/absence variation was found in plant genomes, and it was becoming understood that a single reference cannot represent the species, leading to the construction of pangenomes. An earlier Australian project in 2012 had generated low coverage sequencing of 16 bread wheat varieties for diversity analysis, and this data was reused to construct the first bread wheat pangenome in 2017 using the iterative assembly approach and Chinese Spring as the base reference. This highlighted the huge difference in gene content between Chinese Spring and modern varieties.

    As it becomes both cheaper and easier to sequence wheat genomes, with significant improvements in their assembly quality, many more genomes are being sequenced and compared. These high-quality assemblies were used to make the first wheat graph pangenome in 2022, providing a reference for genome analysis that represents modern germplasm. This reference, together with the sequencing of large diversity sets allows us to understand the genomic basis of many important wheat traits and identify genes and haplotypes that can be used for further crop improvement.

publication date

  • September 2024