Combining Multiple Feature Selection Methods and Deep Learning for High-dimensional Data
File(s)9_1_27_45_mldm.pdf (234.19 KB)
Published version
Author(s)
Mares, MA
Wang, S
Guo, Y
Type
Journal Article
Abstract
Feature or variable selection when the number of features is relatively
large to the number of samples or n << p is a challenge in many machine learning
applications. A large number of statistical methods have been developed to address
this challenge. Each method uses different statistical assumptions about the
shape of the regression function relating the predicted variable to the predictors.
In this paper we propose an alternative: combining results from different feature
selection methods relying on disjoint assumptions about the regression function.
We show that our method will lead to better sensitivity than using different methods
individually, on synthetic datasets and datasets from the UCI machine learning
repository. Our empirical studies on data with n << p show that the accuracy
obtained when training deep neural networks with variables selected using our
method is at least as good as the accuracy obtained when not selecting variables
in advance. Our first conclusion is that the feature selection results are improved
by enlarging the body of limiting assumptions about the function relating the predicted
variable to the predictors. Our second conclusion is that, feature selection
can improve accuracy in deep learning at least on data with n << p.
large to the number of samples or n << p is a challenge in many machine learning
applications. A large number of statistical methods have been developed to address
this challenge. Each method uses different statistical assumptions about the
shape of the regression function relating the predicted variable to the predictors.
In this paper we propose an alternative: combining results from different feature
selection methods relying on disjoint assumptions about the regression function.
We show that our method will lead to better sensitivity than using different methods
individually, on synthetic datasets and datasets from the UCI machine learning
repository. Our empirical studies on data with n << p show that the accuracy
obtained when training deep neural networks with variables selected using our
method is at least as good as the accuracy obtained when not selecting variables
in advance. Our first conclusion is that the feature selection results are improved
by enlarging the body of limiting assumptions about the function relating the predicted
variable to the predictors. Our second conclusion is that, feature selection
can improve accuracy in deep learning at least on data with n << p.
Date Issued
2016-07-01
Date Acceptance
2016-05-09
Citation
Transactions on Machine Learning and Data Mining, 2016, 9 (1), pp.27-45
ISSN
1865-6781
Publisher
ibai Publishing
Start Page
27
End Page
45
Journal / Book Title
Transactions on Machine Learning and Data Mining
Volume
9
Issue
1
Copyright Statement
© ibai-publishing
Sponsor
Commission of the European Communities
Grant Number
FP7 - 115446
Subjects
08 Information And Computing Sciences
Publication Status
Published