Efficient Vision Transformer for Human Pose Estimation via Patch Selection

06/07/2023
by   Kaleab Alemayehu Kinfu, et al.
0

While Convolutional Neural Networks (CNNs) have been widely successful in 2D human pose estimation, Vision Transformers (ViTs) have emerged as a promising alternative to CNNs, boosting state-of-the-art performance. However, the quadratic computational complexity of ViTs has limited their applicability for processing high-resolution images and long videos. To address this challenge, we propose a simple method for reducing ViT's computational complexity based on selecting and processing a small number of most informative patches while disregarding others. We leverage a lightweight pose estimation network to guide the patch selection process, ensuring that the selected patches contain the most important information. Our experimental results on three widely used 2D pose estimation benchmarks, namely COCO, MPII and OCHuman, demonstrate the effectiveness of our proposed methods in significantly improving speed and reducing computational complexity with a slight drop in performance.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
08/10/2016

3D Human Pose Estimation Using Convolutional Neural Networks with 2D Pose Information

While there has been a success in 2D human pose estimation with convolut...
research
05/20/2019

Patch-based 3D Human Pose Refinement

State-of-the-art 3D human pose estimation approaches typically estimate ...
research
04/12/2023

Distilling Token-Pruned Pose Transformer for 2D Human Pose Estimation

Human pose estimation has seen widespread use of transformer models in r...
research
11/23/2019

Simple and Lightweight Human Pose Estimation

Recent research on human pose estimation has achieved significant improv...
research
10/12/2022

Uplift and Upsample: Efficient 3D Human Pose Estimation with Uplifting Transformers

The state-of-the-art for monocular 3D human pose estimation in videos is...
research
01/19/2022

Swin-Pose: Swin Transformer Based Human Pose Estimation

Convolutional neural networks (CNNs) have been widely utilized in many c...
research
05/04/2022

Mobile-URSONet: an Embeddable Neural Network for Onboard Spacecraft Pose Estimation

Spacecraft pose estimation is an essential computer vision application t...

Please sign up or login with your details

Forgot password? Click here to reset