SynSin: End-to-end View Synthesis from a Single Image

CVPR(2020)

引用 370|浏览267
暂无评分
摘要
Single image view synthesis allows for the generation of new views of a scene given a single input image. This is challenging, as it requires comprehensively understanding the 3D scene from a single image. As a result, current methods typically use multiple images, train on ground-truth depth, or are limited to synthetic data. We propose a novel end-to-end model for this task; it is trained on real images without any ground-truth 3D information. To this end, we introduce a novel differentiable point cloud renderer that is used to transform a latent 3D point cloud of features into the target view. The projected features are decoded by our refinement network to inpaint missing regions and generate a realistic output image. The 3D component inside of our generative model allows for interpretable manipulation of the latent feature space at test time, e.g. we can animate trajectories from a single image. Unlike prior work, we can generate high resolution images and generalise to other input resolutions. We outperform baselines and prior work on the Matterport, Replica, and RealEstate10K datasets.
更多
查看译文
关键词
SynSin,Replica datasets,RealEstate10K datasets,Matterport datasets,latent feature space,interpretable manipulation,3D component,realistic output image,target view,differentiable point cloud renderer,ground-truth 3D information,end-to-end model,ground-truth depth,multiple images,end-to-end view synthesis,high resolution images,single image,generative model
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要