A comparison study on patient-psychologist voice diarization.

Rachid Riad,Hadrien Titeux, Laurie Lemoine, Justine Montillot, Agnes Sliwinski,Jennifer Bagnou, Xuan Cao,Anne-Catherine Bachoud-Levi,Emmanuel Dupoux

Workshop on Speech and Language Processing for Assistive Technologies (SLPAT)(2022)

引用 0|浏览20
暂无评分
摘要
Conversations between a clinician and a patient, in natural conditions, are valuable sources of information for medical follow-up. The automatic analysis of these dialogues could help extract new language markers and speed up the clinicians’ reports. Yet, it is not clear which model is the most efficient to detect and identify the speaker turns, especially for individuals with speech disorders. Here, we proposed a split of the data that allows conducting a comparative evaluation of different diarization methods. We designed and trained end-to-end neural network architectures to directly tackle this task from the raw signal and evaluate each approach under the same metric. We also studied the effect of fine-tuning models to find the best performance. Experimental results are reported on naturalistic clinical conversations between Psychologists and Interviewees, at different stages of Huntington’s disease, displaying a large panel of speech disorders. We found out that our best end-to-end model achieved 19.5 % IER on the test set, compared to 23.6% achieved by the finetuning of the X-vector architecture. Finally, we observed that we could extract clinical markers directly from the automatic systems, highlighting the clinical relevance of our methods.
更多
查看译文
关键词
voice,patient-psychologist
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要