首页    期刊浏览 2025年02月28日 星期五
登录注册

文章基本信息

  • 标题:A Sketch Algorithm for Estimating Two-Way and Multi-Way Associations
  • 本地全文:下载
  • 作者:Ping Li ; Kenneth W. Church
  • 期刊名称:Computational Linguistics
  • 印刷版ISSN:0891-2017
  • 电子版ISSN:1530-9312
  • 出版年度:2007
  • 卷号:33
  • 期号:3
  • 页码:305-354
  • DOI:10.1162/coli.2007.33.3.305
  • 语种:English
  • 出版社:MIT Press
  • 摘要:We should not have to look at the entire corpus (e.g., the Web) to know if two (or more) words are strongly associated or not. One can often obtain estimates of associations from a small sample. We develop a sketch-based algorithm that constructs a contingency table for a sample. One can estimate the contingency table for the entire population using straightforward scaling. However, one can do better by taking advantage of the margins (also known as document frequencies). The proposed method cuts the errors roughly in half over Broder's sketches.
国家哲学社会科学文献中心版权所有