
ארמין שמילוביץ
אקדמי בכיר
On dimensionality reduction of high dimensional data sets
High dimensional databases are demanding in terms of the computational power required for their processing. Dimensionality reduction can effectively reduce the costs of various operations (e.g. classification), This research presents an explanation why dimensionality reduction is often possible with minimum information loss. Three kinds of greedy dimensionality reduction techniques are presented: Information Gain (Entropy), Polytomous Logistic Regression and random removal of attributes. An empirical comparison of the effect of the above methods on 10 benchmark data-sets revealed that a relatively simple logistic regression method provided mostly the best results.
| שפת פרסום | אנגלית |
| דפים | 233-238 |
| כרך | 76 |
| סטטוס פרסום | פורסם - 2002 |
Keywords
Data mining
Dimensionality reduction
Logistic regression