JIHAD EL SANA

Senior Academic

Comprehensive synthetic Arabic database for on/off-line script recognition research

Raid M. Saabni, Jihad A. El-Sana

Developing and maintaining large comprehensive databases for script recognition that include different shapes for each word in the lexicon is expensive and difficult. In this paper, we present an efficient system that automatically generates prototypes for each word in a lexicon using multiple appearances of each letter. Large sets of different shapes are created for each letter in each position. These sets are then used to generate valid shapes for each word-part. The number of valid permutations for each word is large and prohibits practical training and searching for various tasks, such as script recognition and word spotting. We apply dimensionality reduction and clustering techniques to maintain compact representation of these databases, without affecting their ability to represent the wide variety of handwriting styles. In addition, a database for off-line script recognition is generated from the on-line strokes using a standard dilation technique, while making special efforts to resemble pen's path. We also examined and used several layout techniques for producing words from the generated word-parts. Our experimental results show that the proposed system can automatically generate large databases, whose quality is at least as good as the manually generated ones.

Publication language English
Pages 285-294
Journal International Journal on Document Analysis and Recognition
Volume 16
Issue number 3
Publication status Published - 01.09.2013

Keywords

Arabic
Database
Kmeans
PCA
Recognition
Script
Synthetic

ASJC Scopus subject areas

Software
Computer Vision and Pattern Recognition
Computer Science Applications

Sustainable Development Goals

SDG 3 - Good Health and Well-being
SDG 4 - Quality Education
Access to Document
10.1007/s10032-012-0189-5
Other files and links
Link to publication in Scopus