Skip to main navigation Skip to search Skip to main content

Sequence-to-Segments Networks for Detecting Segments in Videos

  • Zijun Wei
  • , Boyu Wang
  • , Minh Hoai
  • , Jianming Zhang
  • , Xiaohui Shen
  • , Zhe Lin
  • , Radomir Mech
  • , Dimitris Samaras
  • Stony Brook University
  • Adobe Systems Incorporated
  • ByteDance Ltd.

Research output: Contribution to journalArticlepeer-review

13 Scopus citations

Abstract

Detecting segments of interest from videos is a common problem for many applications. And yet it is a challenging problem as it often requires not only knowledge of individual target segments, but also contextual understanding of the entire video and the relationships between the target segments. To address this problem, we propose the Sequence-to-Segments Network (S2N), a novel and general end-to-end sequential encoder-decoder architecture. S2N first encodes the input video into a sequence of hidden states that capture information progressively, as it appears in the video. It then employs the Segment Detection Unit (SDU), a novel decoding architecture, that sequentially detects segments. At each decoding step, the SDU integrates the decoder state and encoder hidden states to detect a target segment. During training, we address the problem of finding the best assignment of predicted segments to ground truth using the Hungarian Matching Algorithm with Lexicographic Cost. Additionally we propose to use the squared Earth Mover's Distance to optimize the localization errors of the segments. We show the state-of-the-art performance of S2N across numerous tasks, including video highlighting, video summarization, and human action proposal generation.

Original languageEnglish
Article number8827968
Pages (from-to)1009-1021
Number of pages13
JournalIEEE Transactions on Pattern Analysis and Machine Intelligence
Volume43
Issue number3
DOIs
StatePublished - Mar 1 2021

Keywords

  • Segment detection
  • video analysis
  • video highlighting
  • video summarization
  • video temporal action proposal

Fingerprint

Dive into the research topics of 'Sequence-to-Segments Networks for Detecting Segments in Videos'. Together they form a unique fingerprint.

Cite this