מרק לסט

אקדמי בכיר

A graph-based framework for web document mining

Adam Schenker, Horst Bunke, Mark Last, Abraham Kandel

In this paper we describe methods of performing data mining on web documents, where the web document content is represented by graphs. We show how traditional clustering and classification methods, which usually operate on vector representations of data, can be extended to work with graph-based data. Specifically, we give graphtheoretic extensions of the k-Nearest Neighbors classification algorithm and the k-means clustering algorithm that process graphs, and show how the retention of structural information can lead to improved performance over the case of the vector model approach. We introduce several different types of web document representations that utilize graphs and compare their performance for clustering and classification.

שפת פרסום אנגלית
דפים 401-412
סטטוס פרסום פורסם - 01.01.2004

ASJC Scopus subject areas

Theoretical Computer Science
General Computer Science
גישה למסמך
10.1007/978-3-540-28640-0_38
קבצים וקישורים אחרים
Link to publication in Scopus