Waymo Open Dataset: Panoramic Video Panoptic Segmentation

by   Jieru Mei, et al.

Panoptic image segmentation is the computer vision task of finding groups of pixels in an image and assigning semantic classes and object instance identifiers to them. Research in image segmentation has become increasingly popular due to its critical applications in robotics and autonomous driving. The research community thereby relies on publicly available benchmark dataset to advance the state-of-the-art in computer vision. Due to the high costs of densely labeling the images, however, there is a shortage of publicly available ground truth labels that are suitable for panoptic segmentation. The high labeling costs also make it challenging to extend existing datasets to the video domain and to multi-camera setups. We therefore present the Waymo Open Dataset: Panoramic Video Panoptic Segmentation Dataset, a large-scale dataset that offers high-quality panoptic segmentation labels for autonomous driving. We generate our dataset using the publicly available Waymo Open Dataset, leveraging the diverse set of camera images. Our labels are consistent over time for video processing and consistent across multiple cameras mounted on the vehicles for full panoramic scene understanding. Specifically, we offer labels for 28 semantic categories and 2,860 temporal sequences that were captured by five cameras mounted on autonomous vehicles driving in three different geographical locations, leading to a total of 100k labeled camera images. To the best of our knowledge, this makes our dataset an order of magnitude larger than existing datasets that offer video panoptic segmentation labels. We further propose a new benchmark for Panoramic Video Panoptic Segmentation and establish a number of strong baselines based on the DeepLab family of models. We will make the benchmark and the code publicly available. Find the dataset at https://waymo.com/open.


page 2

page 5

page 7

page 10

page 12


Highway Driving Dataset for Semantic Video Segmentation

Scene understanding is an essential technique in semantic segmentation. ...

Immunofluorescence Capillary Imaging Segmentation: Cases Study

Nonunion is one of the challenges faced by orthopedics clinics for the t...

Context-Aware 3D Object Localization from Single Calibrated Images: A Study of Basketballs

Accurately localizing objects in three dimensions (3D) is crucial for va...

EasyPortrait - Face Parsing and Portrait Segmentation Dataset

Recently, due to COVID-19 and the growing demand for remote work, video ...

HandSeg: A Dataset for Hand Segmentation from Depth Images

We introduce a large-scale RGBD hand segmentation dataset, with detailed...

DeepSportradar-v1: Computer Vision Dataset for Sports Understanding with High Quality Annotations

With the recent development of Deep Learning applied to Computer Vision,...

FreeLabel: A Publicly Available Annotation Tool based on Freehand Traces

Large-scale annotation of image segmentation datasets is often prohibiti...

Please sign up or login with your details

Forgot password? Click here to reset