首页    期刊浏览 2024年12月03日 星期二
登录注册

文章基本信息

  • 标题:Streamed Sampling on Dynamic data as Support for Classification Model
  • 本地全文:下载
  • 作者:Astried Silvanie ; Taufik Djatna ; Heru Sukoco
  • 期刊名称:TELKOMNIKA (Telecommunication Computing Electronics and Control)
  • 印刷版ISSN:2302-9293
  • 出版年度:2013
  • 卷号:11
  • 期号:4
  • 页码:855-863
  • DOI:10.12928/telkomnika.v11i4.1210
  • 语种:English
  • 出版社:Universitas Ahmad Dahlan
  • 摘要:Data mining process on dynamically changing data have several problems, such as unknown data size and changing of class distribution . Random sampling method commonly applied for extracting general synopsis from very large database. In this research, Vitter’s reservoir algorithm is used to retrieve k records of data from the database and put into the sample. Sample is used as input for classification task in data mining. Sample type is backing sample and it saved as table contains value of id, priority and timestamp. Priority indicates the probability of how long data retained in the sample. Kullback-Leibler divergence applied to measure the similarity between database and sample distribution. Result of this research is showed that continuously taken samples randomly is possible when transaction occurs. Kullback-Leibler divergence with interval from 0 to 0.0001, is a very good measure to maintain similar class distribution between database and sample. Sample results are always up to date on new transactions with similar class distribution. Classifier built from balance class distribution showed to have better performance than from imbalance one.
国家哲学社会科学文献中心版权所有