גהאד אלצאנע

אקדמי בכיר

VML-HD

The historical Arabic documents dataset for recognition systems

Majeed Kassis, Alaa Abdalhaleem, Ahmad Droby, Reem Alaasam, Jihad El-Sana

In this paper we present a new database with handwritten Arabic script. It is based on five books written by different writers from the years 1088-1451. We took 680 pages from these five books, and fully annotated them on the sub-word level. For each page we manually applied bounding boxes on the different sub-words and annotated the sequence of characters. It consists of 121,636 sub-word appearances consisted of 244,553 characters out of a vocabulary of 1,731 forms of sub-words. The database is described in detail and is designed for training and testing recognition systems for handwritten Arabic sub-words. This database is available for the purpose of research, and we encourage researchers to develop and test new methods using our database.

שפת פרסום אנגלית
דפים 11-14
סטטוס פרסום פורסם - 13.10.2017
8067751

ASJC Scopus subject areas

Computer Vision and Pattern Recognition
Linguistics and Language
Computer Science Applications
גישה למסמך
10.1109/ASAR.2017.8067751
קבצים וקישורים אחרים
Link to publication in Scopus