Reliable Shot Identification for Complex Event Detection via Visual-Semantic Embedding

10/12/2021
by   Minnan Luo, et al.
0

Multimedia event detection is the task of detecting a specific event of interest in an user-generated video on websites. The most fundamental challenge facing this task lies in the enormously varying quality of the video as well as the high-level semantic abstraction of event inherently. In this paper, we decompose the video into several segments and intuitively model the task of complex event detection as a multiple instance learning problem by representing each video as a "bag" of segments in which each segment is referred to as an instance. Instead of treating the instances equally, we associate each instance with a reliability variable to indicate its importance and then select reliable instances for training. To measure the reliability of the varying instances precisely, we propose a visual-semantic guided loss by exploiting low-level feature from visual information together with instance-event similarity based high-level semantic feature. Motivated by curriculum learning, we introduce a negative elastic-net regularization term to start training the classifier with instances of high reliability and gradually taking the instances with relatively low reliability into consideration. An alternative optimization algorithm is developed to solve the proposed challenging non-convex non-smooth problem. Experimental results on standard datasets, i.e., TRECVID MEDTest 2013 and TRECVID MEDTest 2014, demonstrate the effectiveness and superiority of the proposed method to the baseline algorithms.

READ FULL TEXT

page 1

page 4

page 5

page 6

page 8

page 9

page 10

page 11

research
07/20/2020

MINI-Net: Multiple Instance Ranking Network for Video Highlight Detection

We address the weakly supervised video highlight detection problem for l...
research
03/07/2016

A novel learning-based frame pooling method for Event Detection

Detecting complex events in a large video collection crawled from video ...
research
09/30/2020

Visual Semantic Multimedia Event Model for Complex Event Detection in Video Streams

Multimedia data is highly expressive and has traditionally been very dif...
research
10/10/2015

TagBook: A Semantic Video Representation without Supervision for Event Detection

We consider the problem of event detection in video for scenarios where ...
research
07/04/2018

Video Semantic Salient Instance Segmentation: Benchmark Dataset and Baseline

This paper pushes the envelope on salient regions in a video to decompos...
research
05/13/2019

Modelling Instance-Level Annotator Reliability for Natural Language Labelling Tasks

When constructing models that learn from noisy labels produced by multip...
research
06/14/2019

Cost-sensitive Regularization for Label Confusion-aware Event Detection

In supervised event detection, most of the mislabeling occurs between a ...

Please sign up or login with your details

Forgot password? Click here to reset