文章基本信息

标题：Multiple similarly effective solutions exist for biomedical feature selection and classification problems
本地全文：下载
作者：Jiamei Liu ; Cheng Xu ; Weifeng Yang 等
期刊名称：Scientific Reports
电子版ISSN：2045-2322
出版年度：2017
卷号：7
期号：1
DOI：10.1038/s41598-017-13184-8
语种：English
出版社：Springer Nature
摘要：Binary classification is a widely employed problem to facilitate the decisions on various biomedical big data questions, such as clinical drug trials between treated participants and controls, and genome-wide association studies (GWASs) between participants with or without a phenotype. A machine learning model is trained for this purpose by optimizing the power of discriminating samples from two groups. However, most of the classification algorithms tend to generate one locally optimal solution according to the input dataset and the mathematical presumptions of the dataset. Here we demonstrated from the aspects of both disease classification and feature selection that multiple different solutions may have similar classification performances. So the existing machine learning algorithms may have ignored a horde of fishes by catching only a good one. Since most of the existing machine learning algorithms generate a solution by optimizing a mathematical goal, it may be essential for understanding the biological mechanisms for the investigated classification question, by considering both the generated solution and the ignored ones.