A Two-Level Topic Model Towards Knowledge Discovery from Citation Networks

IEEE Transactions on Knowledge and Data Engineering(2014)

引用 46|浏览61
暂无评分
摘要
Knowledge discovery from scientific articles has received increasing attention recently since huge repositories are made available by the development of the Internet and digital databases. In a corpus of scientific articles such as a digital library, documents are connected by citations and one document plays two different roles in the corpus: document itself and a citation of other documents. In the existing topic models, little effort is made to differentiate these two roles. We believe that the topic distributions of these two roles are different and related in a certain way. In this paper, we propose a Bernoulli process topic (BPT) model which considers the corpus at two levels: document level and citation level. In the BPT model, each document has two different representations in the latent topic space associated with its roles. Moreover, the multi-level hierarchical structure of citation network is captured by a generative process involving a Bernoulli process. The distribution parameters of the BPT model are estimated by a variational approximation approach. An efficient computation algorithm is proposed to overcome the difficulty of matrix inverse operation. In addition to conducting the experimental evaluations on the document modeling and document clustering tasks, we also apply the BPT model to well known corpora to discover the latent topics, recommend important citations, detect the trends of various research areas in computer science between 1991 and 1998, and to investigate the interactions among the research areas. The comparisons against state-of-the-art methods demonstrate a very promising performance. The implementations and the data sets are available online .
更多
查看译文
关键词
multilevel hierarchical structure,database management systems,citation networks,two-level topic model,approximation theory,internet,variational approximation approach,knowledge discovery,data mining,computer science,digital databases,bpt model,scientific articles,bernoulli process topic model,unsupervised learning,citation analysis,text mining,latent models,data models,probabilistic logic,computational modeling,vectors,context modeling
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要