Object-Centric Representation Learning from Unlabeled Videos

12/01/2016
by   Ruohan Gao, et al.
0

Supervised (pre-)training currently yields state-of-the-art performance for representation learning for visual recognition, yet it comes at the cost of (1) intensive manual annotations and (2) an inherent restriction in the scope of data relevant for learning. In this work, we explore unsupervised feature learning from unlabeled video. We introduce a novel object-centric approach to temporal coherence that encourages similar representations to be learned for object-like regions segmented from nearby frames. Our framework relies on a Siamese-triplet network to train a deep convolutional neural network (CNN) representation. Compared to existing temporal coherence methods, our idea has the advantage of lightweight preprocessing of the unlabeled video (no tracking required) while still being able to extract object-level regions from which to learn invariances. Furthermore, as we show in results on several standard datasets, our method typically achieves substantial accuracy gains over competing unsupervised methods for image classification and retrieval tasks.

READ FULL TEXT

page 2

page 5

page 6

page 11

page 12

research
05/04/2015

Unsupervised Learning of Visual Representations using Videos

Is strong supervision necessary for learning a good visual representatio...
research
08/03/2017

Unsupervised Representation Learning by Sorting Sequences

We present an unsupervised representation learning approach using videos...
research
06/15/2015

Slow and steady feature analysis: higher order temporal coherence in video

How can unlabeled video augment visual learning? Existing methods perfor...
research
01/24/2018

Unsupervised learning from videos using temporal coherency deep networks

In this work we address the challenging problem of unsupervised learning...
research
11/19/2022

Efficient Video Representation Learning via Masked Video Modeling with Motion-centric Token Selection

Self-supervised Video Representation Learning (VRL) aims to learn transf...
research
03/28/2016

Shuffle and Learn: Unsupervised Learning using Temporal Order Verification

In this paper, we present an approach for learning a visual representati...
research
03/30/2016

Unsupervised Learning of Visual Representations by Solving Jigsaw Puzzles

In this paper we study the problem of image representation learning with...

Please sign up or login with your details

Forgot password? Click here to reset