Skip to main navigation Skip to search Skip to main content

Cross-modal Manifold Cutmix for Self-supervised Video Representation Learning

  • University of North Carolina at Charlotte

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

1 Scopus citations

Abstract

In this paper, we address the challenge of obtaining large-scale unlabelled video datasets for contrastive representation learning in real-world applications. We present a novel video augmentation technique for self-supervised learning, called Cross-Modal Manifold Cutmix (CMMC), which generates augmented samples by combining different modalities in videos. By embedding a video tesseract into another across two modalities in the feature space, our method enhances the quality of learned video representations. We perform extensive experiments on two small-scale video datasets, UCF101 and HMDB51, for action recognition and video retrieval tasks. Our approach is also shown to be effective on the NTU dataset with limited domain knowledge. Our CMMC achieves comparable performance to other self-supervised methods while using less training data for both downstream tasks.

Original languageEnglish
Title of host publicationProceedings of MVA 2023 - 18th International Conference on Machine Vision and Applications
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9784885523434
DOIs
StatePublished - 2023
Event18th International Conference on Machine Vision and Applications, MVA 2023 - Hamamatsu, Japan
Duration: Jul 23 2023Jul 25 2023

Publication series

NameProceedings of MVA 2023 - 18th International Conference on Machine Vision and Applications

Conference

Conference18th International Conference on Machine Vision and Applications, MVA 2023
Country/TerritoryJapan
CityHamamatsu
Period07/23/2307/25/23

Fingerprint

Dive into the research topics of 'Cross-modal Manifold Cutmix for Self-supervised Video Representation Learning'. Together they form a unique fingerprint.

Cite this