Mark Last

Senior Academic

Clustering-based classification of document streams with active learning

Mark Last, Maxim Stoliar, Menahem Friedman

Automated categorization of textual information is becoming an increasingly important task in the digital world. However, most classification algorithms build upon manual labeling of text documents, which is a time-consuming and costly process. In this paper, we present a novel methodology for clustering-based classification of stationary document streams using active learning. The proposed active learning clusteringbased classification algorithm (ACCA) obtains a continuous stream of unlabeled documents. The arriving documents are clustered incrementally so that each incoming document is inserted into an existing cluster or used to start a new cluster of its own. The number of possible clusters is unlimited. From time to time, an expert is called to label several clusters for the classification mechanism. With arrival of more documents, the expert can be called less frequently, since most of the incoming documents will eventually belong to existing labeled clusters. Our algorithm is aimed at finding the fastest way of reaching the point where most arriving documents can be classified automatically without the experts assistance. The evaluation experiments on two benchmark corpora show that active learning and clustering can increase the percentage of automatically and accurately categorized documents over time.

Publication language English
Pages 92-117
Publication status Published - 11.01.2018

ASJC Scopus subject areas

General Computer Science
Other files and links
Link to publication in Scopus