Armin Shmilovici Leib

Senior Academic

On dimensionality reduction of high dimensional data sets

High dimensional databases are demanding in terms of the computational power required for their processing. Dimensionality reduction can effectively reduce the costs of various operations (e.g. classification), This research presents an explanation why dimensionality reduction is often possible with minimum information loss. Three kinds of greedy dimensionality reduction techniques are presented: Information Gain (Entropy), Polytomous Logistic Regression and random removal of attributes. An empirical comparison of the effect of the above methods on 10 benchmark data-sets revealed that a relatively simple logistic regression method provided mostly the best results.
Publication language English
Pages 233-238
Volume 76
Publication status Published - 2002

Keywords

Data mining
Dimensionality reduction
Logistic regression
Other files and links
View record in Web of Science