Classification of high-dimensional data with spiked covariance matrix structure

10/05/2021
by   Yin-Jen Chen, et al.
0

We study the classification problem for high-dimensional data with n observations on p features where the p × p covariance matrix Σ exhibits a spiked eigenvalues structure and the vector ζ, given by the difference between the whitened mean vectors, is sparse with sparsity at most s. We propose an adaptive classifier (adaptive with respect to the sparsity s) that first performs dimension reduction on the feature vectors prior to classification in the dimensionally reduced space, i.e., the classifier whitened the data, then screen the features by keeping only those corresponding to the s largest coordinates of ζ and finally apply Fisher linear discriminant on the selected features. Leveraging recent results on entrywise matrix perturbation bounds for covariance matrices, we show that the resulting classifier is Bayes optimal whenever n →∞ and s √(n^-1ln p)→ 0. Experimental results on real and synthetic data sets indicate that the proposed classifier is competitive with existing state-of-the-art methods while also selecting a smaller number of features.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
02/03/2016

High-Dimensional Regularized Discriminant Analysis

Regularized discriminant analysis (RDA), proposed by Friedman (1989), is...
research
01/21/2013

Supervised Classification Using Sparse Fisher's LDA

It is well known that in a supervised classification setting when the nu...
research
10/30/2017

Distance-based classifier by data transformation for high-dimension, strongly spiked eigenvalue models

We consider classifiers for high-dimensional data under the strongly spi...
research
11/28/2022

High dimensional discriminant rules with shrinkage estimators of covariance matrix and mean vector

Linear discriminant analysis is a typical method used in the case of lar...
research
06/11/2020

Improved Design of Quadratic Discriminant Analysis Classifier in Unbalanced Settings

The use of quadratic discriminant analysis (QDA) or its regularized vers...
research
04/17/2020

Asymptotic Analysis of an Ensemble of Randomly Projected Linear Discriminants

Datasets from the fields of bioinformatics, chemometrics, and face recog...
research
12/14/2017

Fast robust correlation for high dimensional data

The product moment covariance is a cornerstone of multivariate data anal...

Please sign up or login with your details

Forgot password? Click here to reset