首页    期刊浏览 2024年12月04日 星期三
登录注册

文章基本信息

  • 标题:Levenshtein Distances Fail to Identify Language Relationships Accurately
  • 本地全文:下载
  • 作者:Simon J. Greenhill
  • 期刊名称:Computational Linguistics
  • 印刷版ISSN:0891-2017
  • 电子版ISSN:1530-9312
  • 出版年度:2011
  • 卷号:37
  • 期号:4
  • 页码:689-698
  • DOI:10.1162/COLI_a_00073
  • 语种:English
  • 出版社:MIT Press
  • 摘要:The Levenshtein distance is a simple distance metric derived from the number of edit operations needed to transform one string into another. This metric has received recent attention as a means of automatically classifying languages into genealogical subgroups. In this article I test the performance of the Levenshtein distance for classifying languages by subsampling three language subsets from a large database of Austronesian languages. Comparing the classification proposed by the Levenshtein distance to that of the comparative method shows that the Levenshtein classification is correct only 40% of time. Standardizing the orthography increases the performance, but only to a maximum of 65% accuracy within language subgroups. The accuracy of the Levenshtein classification decreases rapidly with phylogenetic distance, failing to discriminate homology and chance similarity across distantly related languages.This poor performance suggests the need for more linguistically nuanced methods for automated language classification tasks.
国家哲学社会科学文献中心版权所有