Genetic-algorithm-based Variable Selection for building a Pavement Deterioration Prediction Modeling
Predicting pavement deterioration is an essential component of pavement management systems. However, traffic and environmental variables commonly used in prediction models often exhibit high levels of multicollinearity, which may affect the variable selection and model performance. This study investigated the multicollinearity structure of pavement deterioration variables and compared model-specific variable selection characteristics using a genetic algorithm (GA)-based wrapper approach. A dataset consisting of 20,253 observations collected from 30 major arterial roads in Daejeon Metropolitan City was analyzed. Multicollinearity was diagnosed using the variance inflation factor (VIF), and GA-based variable selection was independently applied to multiple linear regression (MLR), random forest (RF), and XGBoost models. Seven of the eight explanatory variables exhibited VIF values greater than 10, indicating substantial multicollinearity. After variable selection, the maximum VIF decreased substantially in both MLR and RF, whereas several highly correlated variables remained in XGBoost. In addition, the selected variable subsets differed across the model types. RF and XGBoost showed improved predictive performances after variable selection, whereas MLR exhibited little performance change despite the reduction in multicollinearity. These findings suggest that the effects of variable selection under multicollinear conditions vary according to the model structure and highlight the importance of model-specific variable selection strategies in pavement deterioration prediction.