Skip to main navigation Skip to search Skip to main content

Fashion Captioning: Towards Generating Accurate Descriptions with Semantic Rewards

  • Xuewen Yang
  • , Heming Zhang
  • , Di Jin
  • , Yingru Liu
  • , Chi Hao Wu
  • , Jianchao Tan
  • , Dongliang Xie
  • , Jue Wang
  • , Xin Wang
  • Stony Brook University
  • University of Southern California
  • Massachusetts Institute of Technology
  • Kwai Inc.
  • Beijing University of Posts and Telecommunications
  • Megvii

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

37 Scopus citations

Abstract

Generating accurate descriptions for online fashion items is important not only for enhancing customers’ shopping experiences, but also for the increase of online sales. Besides the need of correctly presenting the attributes of items, the expressions in an enchanting style could better attract customer interests. The goal of this work is to develop a novel learning framework for accurate and expressive fashion captioning. Different from popular work on image captioning, it is hard to identify and describe the rich attributes of fashion items. We seed the description of an item by first identifying its attributes, and introduce attribute-level semantic (ALS) reward and sentence-level semantic (SLS) reward as metrics to improve the quality of text descriptions. We further integrate the training of our model with maximum likelihood estimation (MLE), attribute embedding, and Reinforcement Learning (RL). To facilitate the learning, we build a new FAshion CAptioning Dataset (FACAD), which contains 993K images and 130K corresponding enchanting and diverse descriptions. Experiments on FACAD demonstrate the effectiveness of our model (Code and data: https://github.com/xuewyang/Fashion_Captioning).

Original languageEnglish
Title of host publicationComputer Vision – ECCV 2020 - 16th European Conference, 2020, Proceedings
EditorsAndrea Vedaldi, Horst Bischof, Thomas Brox, Jan-Michael Frahm
PublisherSpringer Science and Business Media Deutschland GmbH
Pages1-17
Number of pages17
ISBN (Print)9783030586003
DOIs
StatePublished - 2020
Event16th European Conference on Computer Vision, ECCV 2020 - Glasgow, United Kingdom
Duration: Aug 23 2020Aug 28 2020

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume12358 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference16th European Conference on Computer Vision, ECCV 2020
Country/TerritoryUnited Kingdom
CityGlasgow
Period08/23/2008/28/20

Keywords

  • Captioning
  • Fashion
  • Reinforcement Learning
  • Semantics

Fingerprint

Dive into the research topics of 'Fashion Captioning: Towards Generating Accurate Descriptions with Semantic Rewards'. Together they form a unique fingerprint.

Cite this