PCA-Based Missing Information Imputation for Real-Time Crash Likelihood Prediction Under Imbalanced Data

02/11/2018
by   Jintao Ke, et al.
0

The real-time crash likelihood prediction has been an important research topic. Various classifiers, such as support vector machine (SVM) and tree-based boosting algorithms, have been proposed in traffic safety studies. However, few research focuses on the missing data imputation in real-time crash likelihood prediction, although missing values are commonly observed due to breakdown of sensors or external interference. Besides, classifying imbalanced data is also a difficult problem in real-time crash likelihood prediction, since it is hard to distinguish crash-prone cases from non-crash cases which compose the majority of the observed samples. In this paper, principal component analysis (PCA) based approaches, including LS-PCA, PPCA, and VBPCA, are employed for imputing missing values, while two kinds of solutions are developed to solve the problem in imbalanced data. The results show that PPCA and VBPCA not only outperform LS-PCA and other imputation methods (including mean imputation and k-means clustering imputation), in terms of the root mean square error (RMSE), but also help the classifiers achieve better predictive performance. The two solutions, i.e., cost-sensitive learning and synthetic minority oversampling technique (SMOTE), help improve the sensitivity by adjusting the classifiers to pay more attention to the minority class.

READ FULL TEXT

page 23

page 28

research
05/30/2022

Principle Components Analysis based frameworks for efficient missing data imputation algorithms

Missing data is a commonly occurring problem in practice, and imputation...
research
05/10/2023

Blockwise Principal Component Analysis for monotone missing data imputation and dimensionality reduction

Monotone missing data is a common problem in data analysis. However, imp...
research
04/11/2020

Spatial Matrix Completion for Spatially-Misaligned and High-Dimensional Air Pollution Data

In health-pollution cohort studies, accurate predictions of pollutant co...
research
11/14/2022

Machine Learning Performance Analysis to Predict Stroke Based on Imbalanced Medical Dataset

Cerebral stroke, the second most substantial cause of death universally,...
research
02/02/2023

Conditional expectation for missing data imputation

Missing data is common in datasets retrieved in various areas, such as m...
research
03/21/2015

Fast Imbalanced Classification of Healthcare Data with Missing Values

In medical domain, data features often contain missing values. This can ...
research
10/11/2022

Combining datasets to increase the number of samples and improve model fitting

For many use cases, combining information from different datasets can be...

Please sign up or login with your details

Forgot password? Click here to reset