Exploring Part-Informed Visual-Language Learning for Person Re-Identification

08/04/2023
by   Yin Lin, et al.
0

Recently, visual-language learning has shown great potential in enhancing visual-based person re-identification (ReID). Existing visual-language learning-based ReID methods often focus on whole-body scale image-text feature alignment, while neglecting supervisions on fine-grained part features. This choice simplifies the learning process but cannot guarantee within-part feature semantic consistency thus hindering the final performance. Therefore, we propose to enhance fine-grained visual features with part-informed language supervision for ReID tasks. The proposed method, named Part-Informed Visual-language Learning (π-VL), suggests that (i) a human parsing-guided prompt tuning strategy and (ii) a hierarchical fusion-based visual-language alignment paradigm play essential roles in ensuring within-part feature semantic consistency. Specifically, we combine both identity labels and parsing maps to constitute pixel-level text prompts and fuse multi-stage visual features with a light-weight auxiliary head to perform fine-grained image-text alignment. As a plug-and-play and inference-free solution, our π-VL achieves substantial improvements over previous state-of-the-arts on four common-used ReID benchmarks, especially reporting 90.3 for the most challenging MSMT17 database without bells and whistles.

READ FULL TEXT
research
08/05/2018

Improving Deep Visual Representation for Person Re-identification by Global and Local Image-language Association

Person re-identification is an important task that requires learning dis...
research
08/21/2023

Exploring Fine-Grained Representation and Recomposition for Cloth-Changing Person Re-Identification

Cloth-changing person Re-IDentification (Re-ID) is a particularly challe...
research
02/17/2023

Fine-grained Cross-modal Fusion based Refinement for Text-to-Image Synthesis

Text-to-image synthesis refers to generating visual-realistic and semant...
research
05/19/2022

Learning Feature Fusion for Unsupervised Domain Adaptive Person Re-identification

Unsupervised domain adaptive (UDA) person re-identification (ReID) has g...
research
08/17/2023

Fine-grained Text and Image Guided Point Cloud Completion with CLIP Model

This paper focuses on the recently popular task of point cloud completio...
research
08/27/2023

Semantic-aware Consistency Network for Cloth-changing Person Re-Identification

Cloth-changing Person Re-Identification (CC-ReID) is a challenging task ...
research
08/14/2021

Focusing on Persons: Colorizing Old Images Learning from Modern Historical Movies

In industry, there exist plenty of scenarios where old gray photos need ...

Please sign up or login with your details

Forgot password? Click here to reset