TY - GEN
T1 - Cross-modal Manifold Cutmix for Self-supervised Video Representation Learning
AU - Das, Srijan
AU - Ryoo, Michael
N1 - Publisher Copyright:
© 2023 IEICE.
PY - 2023
Y1 - 2023
N2 - In this paper, we address the challenge of obtaining large-scale unlabelled video datasets for contrastive representation learning in real-world applications. We present a novel video augmentation technique for self-supervised learning, called Cross-Modal Manifold Cutmix (CMMC), which generates augmented samples by combining different modalities in videos. By embedding a video tesseract into another across two modalities in the feature space, our method enhances the quality of learned video representations. We perform extensive experiments on two small-scale video datasets, UCF101 and HMDB51, for action recognition and video retrieval tasks. Our approach is also shown to be effective on the NTU dataset with limited domain knowledge. Our CMMC achieves comparable performance to other self-supervised methods while using less training data for both downstream tasks.
AB - In this paper, we address the challenge of obtaining large-scale unlabelled video datasets for contrastive representation learning in real-world applications. We present a novel video augmentation technique for self-supervised learning, called Cross-Modal Manifold Cutmix (CMMC), which generates augmented samples by combining different modalities in videos. By embedding a video tesseract into another across two modalities in the feature space, our method enhances the quality of learned video representations. We perform extensive experiments on two small-scale video datasets, UCF101 and HMDB51, for action recognition and video retrieval tasks. Our approach is also shown to be effective on the NTU dataset with limited domain knowledge. Our CMMC achieves comparable performance to other self-supervised methods while using less training data for both downstream tasks.
UR - https://www.scopus.com/pages/publications/85170518882
U2 - 10.23919/MVA57639.2023.10216260
DO - 10.23919/MVA57639.2023.10216260
M3 - Conference contribution
AN - SCOPUS:85170518882
T3 - Proceedings of MVA 2023 - 18th International Conference on Machine Vision and Applications
BT - Proceedings of MVA 2023 - 18th International Conference on Machine Vision and Applications
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 18th International Conference on Machine Vision and Applications, MVA 2023
Y2 - 23 July 2023 through 25 July 2023
ER -