I3rab: A New Arabic Dependency Treebank Based on Arabic Grammatical Theory

07/11/2020
by   Dana Halabi, et al.
0

Treebanks are valuable linguistic resources that include the syntactic structure of a language sentence in addition to POS-tags and morphological features. They are mainly utilized in modeling statistical parsers. Although the statistical natural language parser has recently become more accurate for languages such as English, those for the Arabic language still have low accuracy. The purpose of this paper is to construct a new Arabic dependency treebank based on the traditional Arabic grammatical theory and the characteristics of the Arabic language, to investigate their effects on the accuracy of statistical parsers. The proposed Arabic dependency treebank, called I3rab, contrasts with existing Arabic dependency treebanks in two main concepts. The first concept is the approach of determining the main word of the sentence, and the second concept is the representation of the joined and covert pronouns. To evaluate I3rab, we compared its performance against a subset of Prague Arabic Dependency Treebank that shares a comparable level of details. The conducted experiments show that the percentage improvement reached up to 7.5

READ FULL TEXT

page 13

page 17

page 29

research
01/20/2021

Exploratory Arabic Offensive Language Dataset Analysis

This paper adding more insights towards resources and datasets used in A...
research
10/25/2015

Statistical Parsing by Machine Learning from a Classical Arabic Treebank

Research into statistical parsing for English has enjoyed over a decade ...
research
10/28/2015

CBAS: context based arabic stemmer

Arabic morphology encapsulates many valuable features such as word root....
research
01/29/2019

An Arabic Dependency Treebank in the Travel Domain

In this paper we present a dependency treebank of travel domain sentence...
research
03/07/2021

Automatic Difficulty Classification of Arabic Sentences

In this paper, we present a Modern Standard Arabic (MSA) Sentence diffic...
research
07/13/2018

Image Classification for Arabic: Assessing the Accuracy of Direct English to Arabic Translations

Image classification is an ongoing research challenge. Most of the avail...
research
07/11/2013

Genetic approach for arabic part of speech tagging

With the growing number of textual resources available, the ability to u...

Please sign up or login with your details

Forgot password? Click here to reset