CrowdCleaner: Data cleaning for multi-version data on the web via crowdsourcing

ICDE(2014)

引用 85|浏览125
暂无评分
摘要
Multi-version data is often one of the most concerned information on the Web since this type of data is usually updated frequently. Even though there exist some Web information integration systems that try to maintain the latest update version, the maintained multi-version data usually includes inaccurate and invalid information due to the data integration or update delay errors. In this demo, we present CrowdCleaner, a smart data cleaning system for cleaning multi-version data on the Web, which utilizes crowdsourcing-based approaches for detecting and repairing errors that usually cannot be solved by traditional data integration and cleaning techniques. In particular, CrowdCleaner blends active and passive crowdsourcing methods together for rectifying errors for multi-version data. We demonstrate the following four facilities provided by CrowdCleaner: (1) an error-monitor to find out which items (e.g., submission date, price of real estate, etc.) are wrong versions according to the reports from the crowds, which belongs to a passive crowdsourcing strategy; (2) a task-manager to allocate the tasks to human workers intelligently; (3) a smart-decision-maker to identify which answer from the crowds is correct with active crowdsourcing methods; and (4) a whom-to-ask-finder to discover which users (or human workers) should be the most credible according to their answer records.
更多
查看译文
关键词
smart data cleaning system,web information integration systems,error-monitor facility,update delay error,crowdsourcing methods,smart-decision-maker facility,information retrieval,crowdsourcing-based approach,data integration error,multiversion data,whom-to-ask-finder facility,internet,crowdcleaner,task-manager facility,passive crowdsourcing strategy,data integration,maintenance engineering,entropy,uncertainty
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要