•  
  •  
 

Malaysian Journal of Computing (MJoC)

Corresponding Author

Norshahida Shaadan ([email protected])

Abstract

Carbon monoxide (CO) pollution is a major environmental and public health concern because it is colourless and odourless and is produced mainly by incomplete combustion from vehicles, industries, and household sources. Accurate prediction of CO levels and identification of key contributing factors can support early warning and targeted mitigation strategies. This study compared eight machine-learning regression models—Support Vector Regression (SVR), Random Forest Regression (RFR), Decision Tree Regression (DTR), K-Nearest Neighbours (KNN), CatBoost, XGBoost, LightGBM, and Neural Network Regression (NNR)—using daily air-quality data from the Department of Environment Malaysia for 2018–2023. The dataset comprised 42 input features representing air pollutants, meteorological variables, and volatile organic compounds. Data cleaning, preprocessing, mutual-information-based feature selection, model training, and performance evaluation were conducted using the top 30, 20, and 10 features. CatBoost achieved the best performance across all feature-set sizes, with RMSE values of 0.520, 0.552, and 0.622 and corresponding R² values of 0.750, 0.718, and 0.642, respectively. SVR showed the weakest overall performance. The ten highest-ranked contributors were NO2, NOx, PM2.5, PM10, NO, OP_I_Butane, OP_I_Pentane, OP_N_Butane, OP_2_Methylpentane, and OP_Propane. These findings demonstrate the value of ensemble machine-learning methods for CO prediction and provide evidence on the environmental factors associated with CO variability in Shah Alam.

Publication Date

10-1-2026

Volume

11

Issue

2

Recommendation of Reviewers

yes

Share

COinS