abstract
-
Xianran Li: xianran.li@usda.gov
A substantial portion of characterized natural functional polymorphisms are large indels, which can alter gene structure and expression. One purpose of developing a pangenome is to accelerate gene identification and characterization by pinpointing such large indels. However, for specific genes, surveying and graphing large indels across assemblies are challenging and painstaking tasks.
To overcome the challenge, we have devised two unsupervised machine learning algorithms, CHOICE (Clustering HSPs for Ortholog Identification via Coordinates and Equivalence) and CLIPS (Clustering via Large-Indel Permuted Slopes). CHOICE autonomously retrieves the segments harboring the ortholog from each assembly for comprehensive All-vs-All comparisons, while CLIPS aggregates accessions sharing identical indels into haplotypes for concisely graphing the indel patterns.
We constructed an interactive webapp BRIDGEcereal (https://bridgecereal.scinet.usda.gov/) streamlining these two algorithms to expedite the process. With a just a gene model ID as input, BRIDGEcereal rapidly surveys the presence of large indels among pan-genomes, typically completing the task in under 30 seconds. To showcase the utility of BRIDGEcereal, we demonstrate its effectiveness in identifying potential causal genes associated with large indel polymorphisms using three QTLs (Rc-D1, B1, and Hooded) identified from a wheat RIL population.