12-in-1: Multi-Task Vision and Language Representation Learning

CVPR(2020)

引用 500|浏览469
暂无评分
摘要
Much of vision-and-language research focuses on a small but diverse set of independent tasks and supporting datasets often studied in isolation; however, the visually-grounded language understanding skills required for success at these tasks overlap significantly. In this work, we investigate these relationships between vision-and-language tasks by developing a large-scale, multi-task training regime. Our approach culminates in a single model on 12 datasets from four broad categories of task including visual question answering, caption-based image retrieval, grounding referring expressions, and multi-modal verification. Compared to independently trained single-task models, this represents a reduction from approximately 3 billion parameters to 270 million while simultaneously improving performance by 2.05 points on average across tasks. We use our multi-task framework to perform in-depth analysis of the effect of joint training diverse tasks. Further, we show that finetuning task-specific models from our single multi-task model can lead to further improvements, achieving performance at or above the state-of-the-art.
更多
查看译文
关键词
visual question answering,grounding referring expressions,vision-and-language tasks,visually-grounded language understanding skills,multitask vision,single multitask model,finetuning task-specific models,joint training diverse tasks,multitask framework,single-task models,caption-based image retrieval,independent tasks,language representation learning
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要