abstract
-
email: amidou.ndiaye@usask.ca
Wheat is the most widely cultivated crop in the world and supplies 20% of food calories and protein to the world’s population. Wheat breeders have been striving to improve the quality standards of the food industry by screening 1000’s of breeding lines each year. However, the direct measurement of various end-use quality traits such as milling qualities is daunting and requires a large quantity of grain, traits-specific instruments, and sample throughput is low, all of which limits the screening process.
Advancements in genotyping technologies and the development of new computational approaches has spearheaded advancements in genomic selection (GS) strategies for the prediction of quality attributes with high accuracy. To optimize GS strategies, the selection of informative markers can be an effective strategy to reduce the number of markers and genotyping costs for the practical implementation of GS in wheat breeding.
While several studies have described the impacts of marker density on GS, none of them have addressed strategies to select the most predictive makers to maximize prediction accuracy. In the present study, we used feature selection (FS), a machine learning-based technique, to extract the most informative markers for implementing GS prediction models for end-use quality.
Three independent populations were genotyped with high-throughput SNP arrays and evaluated for various agronomic and quality traits, including yield, protein content, gluten index, dough tenacity, extensibility, and strength. Using FS, we substantially reduced the number of informative markers for all the traits in all populations.
Selecting only the most informative markers resulted in a substantial increase in prediction accuracy for all traits, as compared to using all available markers.For example, using 252 markers improved the prediction accuracy of gluten index by up to 42% in one population, while the improvement in gluten index prediction accuracy reached 164% with only 136 in a second population.
Because plant breeders usually evaluate the performance of breeding lines based on multiple traits, we also applied FS to multi-trait genomic prediction. Our approach gave similar or higher accuracies than the traditional multi-trait model which relied on all available markers, demonstrating the effectiveness of FS for both single- and multi-trait genomic prediction.
Thus, our study presents a novel FS strategy to effectively reduce marker density while maximizing trait prediction accuracy in breeding programs.