Optical character recognition quality affects subjective user perception of historical newspaper clippings

JOURNAL OF DOCUMENTATION(2023)

引用 0|浏览1
暂无评分
摘要
PurposeThis study aims to identify user perception of different qualities of optical character recognition (OCR) in texts. The purpose of this paper is to study the effect of different quality OCR on users' subjective perception through an interactive information retrieval task with a collection of one digitized historical Finnish newspaper.Design/methodology/approachThis study is based on the simulated work task model used in interactive information retrieval. Thirty-two users made searches to an article collection of Finnish newspaper Uusi Suometar 1869-1918 which consists of ca. 1.45 million autosegmented articles. The article search database had two versions of each article with different quality OCR. Each user performed six pre-formulated and six self-formulated short queries and evaluated subjectively the top 10 results using a graded relevance scale of 0-3. Users were not informed about the OCR quality differences of the otherwise identical articles.FindingsThe main result of the study is that improved OCR quality affects subjective user perception of historical newspaper articles positively: higher relevance scores are given to better-quality texts.Originality/valueTo the best of the authors' knowledge, this simulated interactive work task experiment is the first one showing empirically that users' subjective relevance assessments are affected by a change in the quality of an optically read text.
更多
查看译文
关键词
Historical newspapers,OCR quality,Interactive information retrieval,Simulated work task,Evaluation,Finnish
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要