גהאד אלצאנע

אקדמי בכיר

The pinkas dataset

Berat Kurar Barakat, Jihad El-Sana, Irina Rabaev

In historical document image processing, datasets account for a significant part of any research, and are crucial for the diversity and abundance of experimental results, which contribute to the development of new algorithms to meet the new challenge. Moreover, they are very important for benchmarking processing algorithms. Numerous publicly available document image datasets of different languages have been emerged. However, current segmentation and recognition performances are nearly saturated with respect to the present publicly available datasets. As such, collecting and labelling historical document images is a burden on historical document image processing researchers. This paper introduces a public historical document image dataset, Pinkas dataset, with new challenges to open room for improvement and identify strengths and weaknesses of available processing algorithms. It is the first dataset in medieval handwritten Hebrew and fully labeled at word, line and page level by an expert of historical Hebrew manuscripts. Pinkas dataset contributes to the diversity of benchmarking standards. In this paper we present meta features of Pinkas dataset and apply recent word spotting algorithms to analyze the room for improvement in terms of performance.

שפת פרסום אנגלית
דפים 732-737
סטטוס פרסום פורסם - 01.09.2019
8978129

Keywords

Handwritten dataset
Handwritten hebrew dataset
Historical document image analysis

ASJC Scopus subject areas

Computer Vision and Pattern Recognition
גישה למסמך
10.1109/ICDAR.2019.00122
קבצים וקישורים אחרים
Link to publication in Scopus