Generalized Permutation Framework for Testing Model Variable Significance

05/28/2021
by   Yue Wu, et al.
0

A common problem in machine learning is determining if a variable significantly contributes to a model's prediction performance. This problem is aggravated for datasets, such as gene expression datasets, that suffer the worst case of dimensionality: a low number of observations along with a high number of possible explanatory variables. In such scenarios, traditional methods for testing variable statistical significance or constructing variable confidence intervals do not apply. To address these problems, we developed a novel permutation framework for testing the significance of variables in supervised models. Our permutation framework has three main advantages. First, it is non-parametric and does not rely on distributional assumptions or asymptotic results. Second, it not only ranks model variables in terms of relative importance, but also tests for statistical significance of each variable. Third, it can test for the significance of the interaction between model variables. We applied this permutation framework to multi-class classification of the Iris flower dataset and of brain regions in RNA expression data, and using this framework showed variable-level statistical significance and interactions.

READ FULL TEXT
research
03/01/2016

Kernel-based Tests for Joint Independence

We investigate the problem of testing whether d random variables, which ...
research
09/14/2023

Statistically Valid Variable Importance Assessment through Conditional Permutations

Variable importance assessment has become a crucial step in machine-lear...
research
05/23/2019

Computationally Efficient Feature Significance and Importance for Machine Learning Models

We develop a simple and computationally efficient significance test for ...
research
02/02/2023

Hypothesis Testing and Machine Learning: Interpreting Variable Effects in Deep Artificial Neural Networks using Cohen's f2

Deep artificial neural networks show high predictive performance in many...
research
06/01/2023

Interaction Measures, Partition Lattices and Kernel Tests for High-Order Interactions

Models that rely solely on pairwise relationships often fail to capture ...
research
04/28/2022

Generalized permutation tests

Permutation tests are an immensely popular statistical tool, used for te...
research
06/30/2016

A Permutation-based Model for Crowd Labeling: Optimal Estimation and Robustness

The aggregation and denoising of crowd labeled data is a task that has g...

Please sign up or login with your details

Forgot password? Click here to reset