SVFormer: Semi-supervised Video Transformer for Action Recognition

11/23/2022
by   Zhen Xing, et al.
0

Semi-supervised action recognition is a challenging but critical task due to the high cost of video annotations. Existing approaches mainly use convolutional neural networks, yet current revolutionary vision transformer models have been less explored. In this paper, we investigate the use of transformer models under the SSL setting for action recognition. To this end, we introduce SVFormer, which adopts a steady pseudo-labeling framework (ie, EMA-Teacher) to cope with unlabeled video samples. While a wide range of data augmentations have been shown effective for semi-supervised image classification, they generally produce limited results for video recognition. We therefore introduce a novel augmentation strategy, Tube TokenMix, tailored for video data where video clips are mixed via a mask with consistent masked tokens over the temporal axis. In addition, we propose a temporal warping augmentation to cover the complex temporal variation in videos, which stretches selected frames to various temporal durations in the clip. Extensive experiments on three datasets Kinetics-400, UCF-101, and HMDB-51 verify the advantage of SVFormer. In particular, SVFormer outperforms the state-of-the-art by 31.5 Our method can hopefully serve as a strong benchmark and encourage future search on semi-supervised action recognition with Transformer networks.

READ FULL TEXT

page 5

page 7

research
11/25/2021

Learning from Temporal Gradient for Semi-supervised Action Recognition

Semi-supervised video action recognition tends to enable deep neural net...
research
12/14/2021

Temporal Transformer Networks with Self-Supervision for Action Recognition

In recent years, 2D Convolutional Networks-based video action recognitio...
research
09/18/2023

Selective Volume Mixup for Video Action Recognition

The recent advances in Convolutional Neural Networks (CNNs) and Vision T...
research
06/14/2023

EPIC Fields: Marrying 3D Geometry and Video Understanding

Neural rendering is fuelling a unification of learning, 3D geometry and ...
research
03/22/2016

Multi-velocity neural networks for gesture recognition in videos

We present a new action recognition deep neural network which adaptively...
research
08/21/2023

Joint learning of images and videos with a single Vision Transformer

In this study, we propose a method for jointly learning of images and vi...
research
09/01/2022

MAPLE: Masked Pseudo-Labeling autoEncoder for Semi-supervised Point Cloud Action Recognition

Recognizing human actions from point cloud videos has attracted tremendo...

Please sign up or login with your details

Forgot password? Click here to reset