首页    期刊浏览 2025年01月06日 星期一
登录注册

文章基本信息

  • 标题:Scalable Evaluation and Improvement of Document Set Expansion via Neural Positive-Unlabeled Learning
  • 本地全文:下载
  • 作者:Alon Jacovi ; Gang Niu ; Yoav Goldberg
  • 期刊名称:Conference on European Chapter of the Association for Computational Linguistics (EACL)
  • 出版年度:2021
  • 卷号:2021
  • 页码:581-592
  • DOI:10.18653/v1/2021.eacl-main.47
  • 语种:English
  • 出版社:ACL Anthology
  • 摘要:We consider the situation in which a user has collected a small set of documents on a cohesive topic, and they want to retrieve additional documents on this topic from a large collection. Information Retrieval (IR) solutions treat the document set as a query, and look for similar documents in the collection. We propose to extend the IR approach by treating the problem as an instance of positive-unlabeled (PU) learning—i.e., learning binary classifiers from only positive (the query documents) and unlabeled (the results of the IR engine) data. Utilizing PU learning for text with big neural networks is a largely unexplored field. We discuss various challenges in applying PU learning to the setting, showing that the standard implementations of state-of-the-art PU solutions fail. We propose solutions for each of the challenges and empirically validate them with ablation tests. We demonstrate the effectiveness of the new method using a series of experiments of retrieving PubMed abstracts adhering to fine-grained topics, showing improvements over the common IR solution and other baselines.
国家哲学社会科学文献中心版权所有