首页    期刊浏览 2024年12月03日 星期二
登录注册

文章基本信息

  • 标题:SELECTION OF SNP MARKERS: ANALYZING GAW17 DATA USING DIFFERENT METHODOLOGIES
  • 本地全文:下载
  • 作者:Mariana Pavan IÓCA ; Daiane Aparecida ZUANETTI
  • 期刊名称:Revista Brasileira de Biometria
  • 印刷版ISSN:0102-0811
  • 电子版ISSN:1983-0823
  • 出版年度:2021
  • 卷号:39
  • 期号:1
  • 页码:71-88
  • DOI:10.28951/rbb.v39i1.499
  • 出版社:Universidade Federal de Lavras
  • 摘要:The quantity and complexity of generated data due to advances in genetic sequencing technologies has made statistical analysis an essential tool for their correct study and interpretation. However, there is still no agreement about which methodologies are more appropriate for those data, especially for the selection of genetic features that influence a specic phenotype. Genetic data are usually characterized by having a number of variables which is much greater than the number of observations. These variables exhibit little variability and high correlation. These characteristics hinder the application of traditional methodologies for variable selection. In this work (i.) we present dierent methodologies for selecting variables - Random Forest, LASSO and the traditional Stepwise method; (ii.) we apply them to genetic data to select SNP markers that characterize the presence or absence of a disease and (iii.) we compare their performances. Random Forest and Lasso show similar prediction performance, however none of them correctly select the relevant SNPs.
  • 关键词:LASSO; Random Forest; SNP markers; variable selection.
国家哲学社会科学文献中心版权所有