The MADAR Arabic Dialect Corpus and Lexicon
LREC(2018)
摘要
In this paper, we present two resources that were created as part of the Multi Arabic Dialect Applications and Resources (MADAR) project. The first is a large parallel corpus of 25 Arabic city dialects in the travel domain. The second is a lexicon of 1,045 concepts with an average of 45 words from 25 cities per concept. These resources are the first of their kind in terms of the breadth of their coverage and the fine location granularity. The focus on cities, as opposed to regions in studying Arabic dialects, opens new avenues to many areas of research from dialectology to dialect identification and machine translation.
更多查看译文
关键词
Arabic Dialects, Parallel Corpus, Lexicon
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络