Automatically extracting cancer disease characteristics from pathology reports into a Disease Knowledge Representation Model.

Journal of Biomedical Informatics(2009)

引用 186|浏览0
暂无评分
摘要
We introduce an extensible and modifiable knowledge representation model to represent cancer disease characteristics in a comparable and consistent fashion. We describe a system, MedTAS/P which automatically instantiates the knowledge representation model from free-text pathology reports. MedTAS/P is based on an open-source framework and its components use natural language processing principles, machine learning and rules to discover and populate elements of the model. To validate the model and measure the accuracy of MedTAS/P, we developed a gold-standard corpus of manually annotated colon cancer pathology reports. MedTAS/P achieves F1-scores of 0.97-1.0 for instantiating classes in the knowledge representation model such as histologies or anatomical sites, and F1-scores of 0.82-0.93 for primary tumors or lymph nodes, which require the extractions of relations. An F1-score of 0.65 is reported for metastatic tumors, a lower score predominantly due to a very small number of instances in the training and test sets.
更多
查看译文
关键词
analysis system,knowledge representation model,cancer disease knowledge representation model,disease knowledge representation model,anatomical site,medical records,information retrieval,annotated colon cancer pathology,gold-standard corpus,modifiable knowledge representation model,lower score,cancer disease characteristic,free-text pathology report,instantiating class,natural language processing,concept formation,consistent fashion,colon cancer,machine learning,gold standard,knowledge representation
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要