首页    期刊浏览 2024年12月02日 星期一
登录注册

文章基本信息

  • 标题:Content-based SMS Classification: Statistical Analysis for the Relationship between Number of Features and Classification Performance
  • 本地全文:下载
  • 作者:Waddah Waheeb ; Rozaida Ghazali
  • 期刊名称:Computación y Sistemas
  • 印刷版ISSN:1405-5546
  • 出版年度:2017
  • 卷号:21
  • 期号:4
  • 页码:771-785
  • 语种:English
  • 出版社:Instituto Politécnico Nacional
  • 其他摘要:High dimensionality of the feature space is one of the difficulty that affect short message service (SMS) classification performance. Some studies used feature selection methods to pick up some features, while other studies used the full extracted features. In this work, we aim to analyse the relationship between features size and classification performance. For that, a classification performance comparison was carried out between ten features sizes selected by varies feature selection methods. The used methods were chi-square, Gini index and information gain (IG). Support vector machine was used as a classifier. Area Under the ROC (Receiver Operating Characteristics) Curve between true positive rate and false positive rate was used to measure the classification performance. We used the repeated measures ANOVA at p < 0.05 level to analyse the performance. Experimental results showed that IG method outperformed the other methods in all features sizes. The best result was with 50% of the extracted features. Furthermore, the results explicitly showed that using larger features size in the classification does not mean superior performance but sometimes leads to less classification performance. Therefore, feature selection step should be used. By reducing the used features for the classification, without degrading the classification performance, it means reducing memory usage and classification time.
  • 其他关键词:Short text classification; content-based SMS spam filtering; SMS classification; dimension reduction; feature selection; support vector machine; ANOVA.
国家哲学社会科学文献中心版权所有