干旱区生态环境知识资源中心(Arid): OASIS: Only Adversarial Supervision for Semantic Image Synthesis

Arid

DOI	10.1007/s11263-022-01673-x
	OASIS: Only Adversarial Supervision for Semantic Image Synthesis
	Sushko, Vadim; Schoenfeld, Edgar; Zhang, Dan; Gall, Juergen; Schiele, Bernt; Khoreva, Anna
通讯作者	Sushko, V
来源期刊	INTERNATIONAL JOURNAL OF COMPUTER VISION
ISSN	0920-5691
EISSN	1573-1405
出版年	2022
卷号	130 期号:12 页码:2903-2923
英文摘要	Despite their recent successes, generative adversarial networks (GANs) for semantic image synthesis still suffer from poor image quality when trained with only adversarial supervision. Previously, additionally employing the VGG-based perceptual loss has helped to overcome this issue, significantly improving the synthesis quality, but at the same time limited the progress of GAN models for semantic image synthesis. In this work, we propose a novel, simplified GAN model, which needs only adversarial supervision to achieve high quality results. We re-design the discriminator as a semantic segmentation network, directly using the given semantic label maps as the ground truth for training. By providing stronger supervision to the discriminator as well as to the generator through spatially- and semantically-aware discriminator feedback, we are able to synthesize images of higher fidelity and with a better alignment to their input label maps, making the use of the perceptual loss superfluous. Furthermore, we enable high-quality multi-modal image synthesis through global and local sampling of a 3D noise tensor injected into the generator, which allows complete or partial image editing. We show that images synthesized by our model are more diverse and follow the color and texture distributions of real images more closely. We achieve a strong improvement in image synthesis quality over prior state-of-the-art models across the commonly used ADE20K, Cityscapes, and COCO-Stuff datasets using only adversarial supervision. In addition, we investigate semantic image synthesis under severe class imbalance and sparse annotations, which are common aspects in practical applications but were overlooked in prior works. To this end, we evaluate our model on LVIS, a dataset originally introduced for long-tailed object recognition. We thereby demonstrate high performance of our model in the sparse and unbalanced data regimes, achieved by means of the proposed 3D noise and the ability of our discriminator to balance class contributions directly in the loss function. Our code and pretrained models are available at https://github.com/boschresearch/OASIS.
英文关键词	Semantic image synthesis GAN Semantic segmentation Label-to-image translation Image editing
类型	Article
语种	英语
开放获取类型	hybrid
收录类别	SCI-E
WOS记录号	WOS:000854437300001
WOS类目	Computer Science, Artificial Intelligence
WOS研究方向	Computer Science
资源类型	期刊论文
条目标识符	http://119.78.100.177/qdio/handle/2XILL650/393138
推荐引用方式 GB/T 7714	Sushko, Vadim,Schoenfeld, Edgar,Zhang, Dan,et al. OASIS: Only Adversarial Supervision for Semantic Image Synthesis[J],2022,130(12):2903-2923.
APA	Sushko, Vadim,Schoenfeld, Edgar,Zhang, Dan,Gall, Juergen,Schiele, Bernt,&Khoreva, Anna.(2022).OASIS: Only Adversarial Supervision for Semantic Image Synthesis.INTERNATIONAL JOURNAL OF COMPUTER VISION,130(12),2903-2923.
MLA	Sushko, Vadim,et al."OASIS: Only Adversarial Supervision for Semantic Image Synthesis".INTERNATIONAL JOURNAL OF COMPUTER VISION 130.12(2022):2903-2923.