[Citation needed] Data usage and citation practices in medical imaging conferences
CoRR(2024)
摘要
Medical imaging papers often focus on methodology, but the quality of the
algorithms and the validity of the conclusions are highly dependent on the
datasets used. As creating datasets requires a lot of effort, researchers often
use publicly available datasets, there is however no adopted standard for
citing the datasets used in scientific papers, leading to difficulty in
tracking dataset usage. In this work, we present two open-source tools we
created that could help with the detection of dataset usage, a pipeline
using
OpenAlex and full-text analysis, and a PDF annotation software
used in our study to
manually label the presence of datasets. We applied both tools on a study of
the usage of 20 publicly available medical datasets in papers from MICCAI and
MIDL. We compute the proportion and the evolution between 2013 and 2023 of 3
types of presence in a paper: cited, mentioned in the full text, cited and
mentioned. Our findings demonstrate the concentration of the usage of a limited
set of datasets. We also highlight different citing practices, making the
automation of tracking difficult.
更多查看译文
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要