מרק לסט

אקדמי בכיר

Clustering of web documents using graph representations

Adam Schenker, Horst Bunke, Mark Last, Abraham Kandel

In this paper we describe a clustering method that allows the use of graph-based representations of data instead of traditional vector-based representations. Using this new method we conduct content-based clustering of two web document collections. Clustering of web documents is performed to organize the documents with little or no human intervention. Benefits of clustering include easier browsing and improved retrieval speed. In order to measure the performance of our graph-matching approach, we compare it to the popular vector-based k-means method. We perform experiments using different graph distance measures as well as various document representations that utilize graphs. The results with the k-means clustering algorithm show that the graph-based approach can outperform traditional vector-based methods.

שפת פרסום אנגלית
דפים 247-265
סטטוס פרסום פורסם - 19.04.2007

Keywords

Graph distance
Graph representations
k-Means

ASJC Scopus subject areas

Artificial Intelligence
גישה למסמך
10.1007/978-3-540-68020-8_10
קבצים וקישורים אחרים
Link to publication in Scopus