FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents

05/27/2019
by   Guillaume Jaume, et al.
0

In this paper, we present a new dataset for Form Understanding in Noisy Scanned Documents (FUNSD). Form Understanding (FoUn) aims at extracting and structuring the textual content of forms. The dataset comprises 200 fully annotated real scanned forms. The documents are noisy and exhibit large variabilities in their representation making FoUn a challenging task. The proposed dataset can be used for various tasks including text detection, optical character recognition (OCR), spatial layout analysis and entity labeling/linking. To the best of our knowledge this is the first publicly available dataset with comprehensive annotations addressing the FoUn task. We also present a set of baselines and introduce metrics to evaluate performance on the FUNSD dataset. The FUNSD dataset can be downloaded at https://guillaumejaume.github. io/FUNSD/.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
05/17/2022

Detection Masking for Improved OCR on Noisy Documents

Optical Character Recognition (OCR), the task of extracting textual info...
research
08/25/2023

Nougat: Neural Optical Understanding for Academic Documents

Scientific knowledge is predominantly stored in books and scientific jou...
research
04/26/2022

Symlink: A New Dataset for Scientific Symbol-Description Linking

Mathematical symbols and descriptions appear in various forms across doc...
research
06/02/2021

End-to-End Hierarchical Relation Extraction for Generic Form Understanding

Form understanding is a challenging problem which aims to recognize sema...
research
03/22/2019

Line-items and table understanding in structured documents

Table detection and extraction has been studied in the context of docume...
research
11/07/2016

Presenting a New Dataset for the Timeline Generation Problem

The timeline generation task summarises an entity's biography by selecti...
research
12/14/2021

Text Classification Models for Form Entity Linking

Forms are a widespread type of template-based document used in a great v...

Please sign up or login with your details

Forgot password? Click here to reset