360-Degree Gaze Estimation in the Wild Using Multiple Zoom Scales

by   Ashesh Mishra, et al.

Gaze estimation involves predicting where the person is looking at, given either a single input image or a sequence of images. One challenging task, gaze estimation in the wild, concerns data collected in unconstrained environments with varying camera-person distances, like the Gaze360 dataset. The varying distances result in varying face sizes in the images, which makes it hard for current CNN backbones to estimate the gaze robustly. Inspired by our natural skill to identify the gaze by taking a focused look at the face area, we propose a novel architecture that similarly zooms in on the face area of the image at multiple scales to improve prediction accuracy. Another challenging task, 360-degree gaze estimation (also introduced by the Gaze360 dataset), consists of estimating not only the forward gazes, but also the backward ones. The backward gazes introduce discontinuity in the yaw angle values of the gaze, making the deep learning models affected by some huge loss around the discontinuous points. We propose to convert the angle values by sine-cosine transform to avoid the discontinuity and represent the physical meaning of the yaw angle better. We conduct ablation studies on both ideas, the novel architecture and the transform, to validate their effectiveness. The two ideas allow our proposed model to achieve state-of-the-art performance for both the Gaze360 dataset and the RT-Gene dataset when using single images. Furthermore, we extend the model to a sequential version that systematically zooms in on a given sequence of images. The sequential version again achieves state-of-the-art performance on the Gaze360 dataset, which further demonstrates the usefulness of our proposed ideas.


page 2

page 4

page 6


L2CS-Net: Fine-Grained Gaze Estimation in Unconstrained Environments

Human gaze is a crucial cue used in various applications such as human-r...

Appearance-Based Gaze Estimation in the Wild

Appearance-based gaze estimation is believed to work well in real-world ...

GazeOnce: Real-Time Multi-Person Gaze Estimation

Appearance-based gaze estimation aims to predict the 3D eye gaze directi...

Gaze360: Physically Unconstrained Gaze Estimation in the Wild

Understanding where people are looking is an informative social cue. In ...

GEDDnet: A Network for Gaze Estimation with Dilation and Decomposition

Appearance-based gaze estimation from RGB images provides relatively unc...

DeepWarp: Photorealistic Image Resynthesis for Gaze Manipulation

In this work, we consider the task of generating highly-realistic images...

Eyes are the Windows to the Soul: Predicting the Rating of Text Quality Using Gaze Behaviour

Predicting a reader's rating of text quality is a challenging task that ...

Please sign up or login with your details

Forgot password? Click here to reset