
Armin Shmilovici Leib
Senior Academic
On dimensionality reduction of high dimensional data sets
High dimensional databases are demanding in terms of the computational power required for their processing. Dimensionality reduction can effectively reduce the costs of various operations (e.g. classification), This research presents an explanation why dimensionality reduction is often possible with minimum information loss. Three kinds of greedy dimensionality reduction techniques are presented: Information Gain (Entropy), Polytomous Logistic Regression and random removal of attributes. An empirical comparison of the effect of the above methods on 10 benchmark data-sets revealed that a relatively simple logistic regression method provided mostly the best results.
| Publication language | English |
| Pages | 233-238 |
| Volume | 76 |
| Publication status | Published - 2002 |
Keywords
Data mining
Dimensionality reduction
Logistic regression