ParaCrawl Corpus version 1.0

user-5fe1a78c4c775e6ec07359f9(2018)

引用 0|浏览25
暂无评分
摘要
The January 2018 release of the ParaCrawl is the first version of the corpus. It contains parallel corpora for 11 languages paired with English, crawled from a large number of web sites. The selection of websites is based on CommonCrawl, but ParaCrawl is extracted from a brand new crawl which has much higher coverage of these selected websites than CommonCrawl. Since the data is fairly raw, it is released with two quality metrics that can be used for corpus filtering. An official "clean" version of each corpus uses one of the metrics. For more details and raw data download please visit: http://paracrawl.eu/releases.html
更多
查看译文
关键词
Text corpus,Machine translation,Raw data,Natural language processing,Download,Computer science,Artificial intelligence,Parallel corpora
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要