Skip to main navigation Skip to search Skip to main content

From Within to Between: Knowledge Distillation for Cross Modality Retrieval

  • Stony Brook University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

We propose a novel loss function for training text-to-video and video-to-text retrieval networks based on knowledge distillation. This loss function addresses an important drawback of the max-margin loss function often used in existing cross-modality retrieval methods, in which a fixed margin is used in training to separate matching video-and-caption pairs from non-matching pairs, treating all non-matching pairs the same and failing to account for the different degrees of non-matching. We address this drawback by introducing a novel loss for the non-matching pairs; this loss leverages the knowledge within one domain to train a better network for matching between two domains. This proposed loss does not require extra annotation. It is complementary to the existing max-margin loss, and it can be integrated into the training pipeline of any cross-modality retrieval method. Experimental results on four cross-modal retrieval datasets namely MSRVTT, ActivityNet, DiDeMo, and MSVD show the effectiveness of the proposed method. Code is available at: https://github.com/tqvinhcs/CrossKD.

Original languageEnglish
Title of host publicationComputer Vision – ACCV 2022 - 16th Asian Conference on Computer Vision, Proceedings
EditorsLei Wang, Juergen Gall, Tat-Jun Chin, Imari Sato, Rama Chellappa
PublisherSpringer Science and Business Media Deutschland GmbH
Pages605-622
Number of pages18
ISBN (Print)9783031263156
DOIs
StatePublished - 2023
Event16th Asian Conference on Computer Vision, ACCV 2022 - Hybrid, Macao, China
Duration: Dec 4 2022Dec 8 2022

Publication series

NameLecture Notes in Computer Science
Volume13844 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference16th Asian Conference on Computer Vision, ACCV 2022
Country/TerritoryChina
CityHybrid, Macao
Period12/4/2212/8/22

Keywords

  • Knowledge distillation
  • Text-video retrieval

Fingerprint

Dive into the research topics of 'From Within to Between: Knowledge Distillation for Cross Modality Retrieval'. Together they form a unique fingerprint.

Cite this