Spatio-Temporal Context for Action Detection

Research in action detection has grown in the recentyears, as it plays a key role in video understanding. Modelling the interactions (either spatial or temporal) between actors and their context has proven to be essential for this task. While recent works use spatial features with aggregated temporal information, this work proposes to use non-aggregated temporal information. This is done by adding an attention based method that leverages spatio-temporal interactions between elements in the scene along the clip.The main contribution of this work is the introduction of two cross attention blocks to effectively model the spatial relations and capture short range temporal interactions.Experiments on the AVA dataset show the advantages of the proposed approach that models spatio-temporal relations between relevant elements in the scene, outperforming other methods that model actor interactions with their context by +0.31 mAP.

READ FULL TEXT
research
07/28/2018

Actor-Centric Relation Network

Current state-of-the-art approaches for spatio-temporal action localizat...
research
01/05/2020

Spatio-Temporal Relation and Attention Learning for Facial Action Unit Detection

Spatio-temporal relations among facial action units (AUs) convey signifi...
research
02/18/2021

An Enhanced Adversarial Network with Combined Latent Features for Spatio-Temporal Facial Affect Estimation in the Wild

Affective Computing has recently attracted the attention of the research...
research
12/30/2018

Actor Conditioned Attention Maps for Video Action Detection

Interactions with surrounding objects and people contain important infor...
research
05/12/2018

HOC-Tree: A Novel Index for efficient Spatio-temporal Range Search

With the rapid development of mobile computing and Web services, a huge ...
research
05/25/2021

ST-HOI: A Spatial-Temporal Baseline for Human-Object Interaction Detection in Videos

Detecting human-object interactions (HOI) is an important step toward a ...
research
04/24/2023

MRSN: Multi-Relation Support Network for Video Action Detection

Action detection is a challenging video understanding task, requiring mo...

Please sign up or login with your details

Forgot password? Click here to reset