Instance Segmentation for Chinese Character Stroke Extraction, Datasets and Benchmarks

by   Lizhao Liu, et al.

Stroke is the basic element of Chinese character and stroke extraction has been an important and long-standing endeavor. Existing stroke extraction methods are often handcrafted and highly depend on domain expertise due to the limited training data. Moreover, there are no standardized benchmarks to provide a fair comparison between different stroke extraction methods, which, we believe, is a major impediment to the development of Chinese character stroke understanding and related tasks. In this work, we present the first public available Chinese Character Stroke Extraction (CCSE) benchmark, with two new large-scale datasets: Kaiti CCSE (CCSE-Kai) and Handwritten CCSE (CCSE-HW). With the large-scale datasets, we hope to leverage the representation power of deep models such as CNNs to solve the stroke extraction task, which, however, remains an open question. To this end, we turn the stroke extraction problem into a stroke instance segmentation problem. Using the proposed datasets to train a stroke instance segmentation model, we surpass previous methods by a large margin. Moreover, the models trained with the proposed datasets benefit the downstream font generation and handwritten aesthetic assessment tasks. We hope these benchmark results can facilitate further research. The source code and datasets are publicly available at:


page 1

page 2

page 3

page 4


AutoKary2022: A Large-Scale Densely Annotated Dateset for Chromosome Instance Segmentation

Automated chromosome instance segmentation from metaphase cell microscop...

LineFormer: Rethinking Line Chart Data Extraction as Instance Segmentation

Data extraction from line-chart images is an essential component of the ...

Advanced Deep Networks for 3D Mitochondria Instance Segmentation

Mitochondria instance segmentation from electron microscopy (EM) images ...

HRCenterNet: An Anchorless Approach to Chinese Character Segmentation in Historical Documents

The information provided by historical documents has always been indispe...

Instance Segmentation Based Graph Extraction for Handwritten Circuit Diagram Images

Handwritten circuit diagrams from educational scenarios or historic sour...

A Large-Scale Dataset for Benchmarking Elevator Button Segmentation and Character Recognition

Human activities are hugely restricted by COVID-19, recently. Robots tha...

DuReader_retrieval: A Large-scale Chinese Benchmark for Passage Retrieval from Web Search Engine

In this paper, we present DuReader_retrieval, a large-scale Chinese data...

Please sign up or login with your details

Forgot password? Click here to reset