The Impact of Word Splitting on the Semantic Content of Contextualized Word Representations
CoRR(2024)
摘要
When deriving contextualized word representations from language models, a
decision needs to be made on how to obtain one for out-of-vocabulary (OOV)
words that are segmented into subwords. What is the best way to represent these
words with a single vector, and are these representations of worse quality than
those of in-vocabulary words? We carry out an intrinsic evaluation of
embeddings from different models on semantic similarity tasks involving OOV
words. Our analysis reveals, among other interesting findings, that the quality
of representations of words that are split is often, but not always, worse than
that of the embeddings of known words. Their similarity values, however, must
be interpreted with caution.
更多查看译文
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要