PeR-ViS: Person Retrieval in Video Surveillance using Semantic Description

by   Parshwa Shah, et al.

A person is usually characterized by descriptors like age, gender, height, cloth type, pattern, color, etc. Such descriptors are known as attributes and/or soft-biometrics. They link the semantic gap between a person's description and retrieval in video surveillance. Retrieving a specific person with the query of semantic description has an important application in video surveillance. Using computer vision to fully automate the person retrieval task has been gathering interest within the research community. However, the Current, trend mainly focuses on retrieving persons with image-based queries, which have major limitations for practical usage. Instead of using an image query, in this paper, we study the problem of person retrieval in video surveillance with a semantic description. To solve this problem, we develop a deep learning-based cascade filtering approach (PeR-ViS), which uses Mask R-CNN [14] (person detection and instance segmentation) and DenseNet-161 [16] (soft-biometric classification). On the standard person retrieval dataset of SoftBioSearch [6], we achieve 0.566 Average IoU and 0.792 surpassing the current state-of-the-art by a large margin. We hope our simple, reproducible, and effective approach will help ease future research in the domain of person retrieval in video surveillance. The source code and pretrained weights available at


page 1

page 3

page 4

page 7

page 8


Person Retrieval in Surveillance Using Textual Query: A Review

Recent advancement of research in biometrics, computer vision, and natur...

Visual Appearance Based Person Retrieval in Unconstrained Environment Videos

Visual appearance-based person retrieval is a challenging problem in sur...

Human Extraction and Scene Transition utilizing Mask R-CNN

Object detection is a trendy branch of computer vision, especially on hu...

Text-based Person Search in Full Images via Semantic-Driven Proposal Generation

Finding target persons in full scene images with a query of text descrip...

Survey on Deep Learning Techniques for Person Re-Identification Task

Intelligent video-surveillance is currently an active research field in ...

Automatic View-Point Selection for Inter-Operative Endoscopic Surveillance

Esophageal adenocarcinoma arises from Barrett's esophagus, which is the ...

Describe me if you can! Characterized Instance-level Human Parsing

Several computer vision applications such as person search or online fas...

Please sign up or login with your details

Forgot password? Click here to reset