Optimizing bank customer churn prediction with machine learning models: a comparative study
Abstract
This study utilizes five machine learning algorithms to classify and predict churn behavior based on financial and demographic characteristics from a dataset of 10,000 customers of a confidential multinational bank. This dataset is publicly accessible on Kaggle. The models in this study are random forest (RF), decision tree (DT), light gradient boosting machine (LightGBM), categorical boosting (CatBoost), and multi-layer perceptron (MLP). After initial data preprocessing, the imbalanced data is addressed by the synthetic minority over-sampling technique (SMOTE). Based on the model performance evaluation, LightGBM and CatBoost are the best-performing models in terms of accuracy, F1-score, and average training time. Moreover, the correlation analysis, including chi-square, variance analysis, and pearson correlation, highlights eight critical factors influencing customer churn, including CreditScore, Geography, Gender, Age, Balance, NumOfProducts, IsActiveMember, and Exited. Based on the experimental results, LightGBM and CatBoost are recommended due to superior model performance and processing times, although these models demand significant computational resources based on CPU and RAM usage. The study suggests a new research direction involving the combination of LightGBM and MLP to better capture non-linear relationships and improve prediction performance.
Keywords
Boosting algorithms; Customer churn; Data imbalance; Data mining; Machine learning; Model optimization
Full Text:
PDFDOI: https://doi.org/10.11591/eei.v15i4.10732
Refbacks
- There are currently no refbacks.

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Bulletin of Electrical Engineering and Informatics (BEEI)
ISSN: 2089-3191
,
e-ISSN: 2302-9285
This journal is published by the
Institute of Advanced Engineering and Science (IAES)
.