A Logic for Document Spanners

ICDT(2018)

引用 33|浏览43
暂无评分
摘要
Document spanners are a formal framework for information extraction that was introduced by Fagin, Kimelfeld, Reiss, and Vansummeren (PODS 2013, JACM 2015). One of the central models in this framework are core spanners, which formalize the query language AQL that is used in IBM’s SystemT. As shown by Freydenberger and Holldack (ICDT 2016, ToCS 2018), there is a connection between core spanners and i̇𝖤𝖢^reg , the existential theory of concatenation with regular constraints. The present paper further develops this connection by defining i̇𝖲𝗉𝖫𝗈𝗀 , a fragment of i̇𝖤𝖢^reg that has the same expressive power as core spanners. This equivalence extends beyond equivalence of expressive power, as we show the existence of polynomial time conversions between i̇𝖲𝗉𝖫𝗈𝗀 and core spanners. Consequences and applications include an alternative way of defining relations for spanners, a pumping lemma for core spanners, and insights into the relative succinctness of various classes of spanner representations and their connection to graph querying languages. We also briefly discuss the connection between i̇𝖲𝗉𝖫𝗈𝗀 with negation and core spanners with a difference operator.
更多
查看译文
关键词
Information extraction,Document spanners,Word equations,xregex,Descriptional complexity,CRPQs with string equality
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要