Skip to main navigation Skip to search Skip to main content

Unsupervised Discovery of Actions in Instructional Videos

  • Alphabet Inc.

Research output: Contribution to conferencePaperpeer-review

1 Scopus citations

Abstract

In this paper we address the problem of automatically discovering atomic actions from instructional videos. Instructional videos contain complex activities and are a rich source of information for intelligent agents, such as, autonomous robots or virtual assistants, which can, for example, automatically 'read' the steps from an instructional video and execute them. However, videos are rarely annotated with atomic activities, their boundaries or duration. We present an unsupervised approach to learn atomic actions of structured human tasks from a variety of instructional videos. We propose a sequential stochastic autoregressive model for temporal segmentation of videos, which learns to represent and discover the sequential relationship between different actions of the task, and provides automatic and unsupervised self-labeling. We evaluate on the breakfast, 50-salads and narrated instructional videos datasets. Code will be open sourced.

Original languageEnglish
StatePublished - 2021
Event32nd British Machine Vision Conference, BMVC 2021 - Virtual, Online
Duration: Nov 22 2021Nov 25 2021

Conference

Conference32nd British Machine Vision Conference, BMVC 2021
CityVirtual, Online
Period11/22/2111/25/21

Fingerprint

Dive into the research topics of 'Unsupervised Discovery of Actions in Instructional Videos'. Together they form a unique fingerprint.

Cite this