首页    期刊浏览 2024年12月12日 星期四
登录注册

文章基本信息

  • 标题:CCGbank: A Corpus of CCG Derivations and Dependency Structures Extracted from the Penn Treebank
  • 本地全文:下载
  • 作者:Julia Hockenmaier ; Mark Steedman
  • 期刊名称:Computational Linguistics
  • 印刷版ISSN:0891-2017
  • 电子版ISSN:1530-9312
  • 出版年度:2007
  • 卷号:33
  • 期号:3
  • 页码:355-396
  • DOI:10.1162/coli.2007.33.3.355
  • 语种:English
  • 出版社:MIT Press
  • 摘要:This article presents an algorithm for translating the Penn Treebank into a corpus of Combinatory Categorial Grammar (CCG) derivations augmented with local and long-range word-word dependencies. The resulting corpus, CCGbank, includes 99.4% of the sentences in the Penn Treebank. It is available from the Linguistic Data Consortium, and has been used to train wide-coverage statistical parsers that obtain state-of-the-art rates of dependency recovery. In order to obtain linguistically adequate CCG analyses, and to eliminate noise and inconsistencies in the original annotation, an extensive analysis of the constructions and annotations in the Penn Treebank was called for, and a substantial number of changes to the Treebank were necessary. We discuss the implications of our findings for the extraction of other linguistically expressive grammars from the Treebank, and for the design of future treebanks.
国家哲学社会科学文献中心版权所有