Stability of SHAP-Based Feature Importance Ranking under Class Imbalance, Feature Correlation, and Feature Dimensionality

Abstract View: 32,
Download: 10
Download: 0

Authors

DOI:

https://doi.org/10.32665/statkom.v5i1.6455

Keywords:

Feature Importance, LightGBM, SHAP, SRA, XGBoost

Abstract

Background: Gradient boosting machine learning models, such as XGBoost and LightGBM, are widely used because of their high predictive performance. However, their complexity requires interpretation methods. SHAP is often used to explain feature contributions, but the stability of its interpretations may be affected by data characteristics and the model used.

Objective: This study evaluates the stability of SHAP-based feature importance rankings in XGBoost and LightGBM under various data conditions.

Methods: The study used controlled simulation and empirical BPJS Kesehatan claims data. In the simulation, 64 datasets were generated from combinations of minority class proportion, feature correlation, and number of features. The models were fitted repeatedly, and ranking stability was evaluated using Sequential Rank Agreement (SRA), where smaller values indicate more stable rankings.

Results: Higher feature correlation and extreme class imbalance reduced feature ranking stability. In the BPJS Kesehatan data, imbalance handling methods improved both model performance and interpretation stability. LightGBM with ADASYN produced a smaller SRA value (0.5350) than the model without imbalance handling (2.2534).

Conclusion: Feature correlation and class imbalance play an important role in determining the stability of SHAP interpretations. SMOTE and ADASYN improve predictive performance and increase the stability of feature importance rankings.

References

Alshboul, O., Almasabha, G., Shehadeh, A., & Al-Shboul, K. (2024). A comparative study of LightGBM, XGBoost, and GEP models in shear strength management of SFRC-SBWS. Structures, 61, 106009. https://doi.org/10.1016/j.istruc.2024.106009

Angulo, E., Romero, F. P., & López-Gómez, J. A. (2021). A comparison of different soft-computing techniques for the evaluation of handball goalkeepers. Soft Computing, 26(6), 3045–3058. https://doi.org/10.1007/s00500-021-06440-7

Asif, D., Arif, M. S., & Mukheimer, A. (2025). A data-driven approach with explainable artificial intelligence for customer churn prediction in the telecommunications industry. Results in Engineering, 26, 104629. https://doi.org/10.1016/j.rineng.2025.104629

Aswi, A., Rahardiantoro, S., Kurnia, A., Sartono, B., Handayani, D., & Nurwan, N. (2025). Bayesian spatio-temporal conditional autoregressive localized modeling techniques for socioeconomic factors and stunting in Indonesia. MethodsX, 15, 103464. https://doi.org/10.1016/j.mex.2025.103464

Ballegeer, M., Bogaert, M., & Benoit, D. F. (2025). Evaluating the stability of model explanations in instance-dependent cost-sensitive credit scoring. European Journal of Operational Research, 326(3), 630–640. https://doi.org/10.1016/j.ejor.2025.05.039

Bracke, P., Datta, A., Jung, C., & Sen, S. (2019). Machine Learning Explainability in Finance: An application to default Risk analysis. SSRN Electronic Journal. https://doi.org/10.2139/ssrn.3435104

Bussmann, N., Giudici, P., Marinelli, D., & Papenbrock, J. (2020). Explainable machine learning in credit risk management. Computational Economics, 57(1), 203–216. https://doi.org/10.1007/s10614-020-10042-0

Chen, T., & Guestrin, C. (2016). XGBoost: A Scalable Tree Boosting System. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 785–794. https://doi.org/10.1145/2939672.2939785

Dharmawan, H., Sartono, B., Kurnia, A., Hadi, A. F., & Ramadhani, E. (2022). A study of machine learning algorithms to measure the feature importance in class-imbalance data of food insecurity cases in Indonesia. Communications in Mathematical Biology and Neuroscience. https://doi.org/10.28919/cmbn/7636

Dietz, L. W., Sertkan, M., Myftija, S., Palage, S. T., Neidhardt, J., & Wörndl, W. (2022). A Comparative Study of Data-Driven Models for Travel Destination Characterization. Frontiers in Big Data, 5, 829939. https://doi.org/10.3389/fdata.2022.829939

Ekstrøm, C. T., Gerds, T. A., & Jensen, A. K. (2018). Sequential rank agreement methods for comparison of ranked lists. Biostatistics, 20(4), 582–598. https://doi.org/10.1093/biostatistics/kxy017

Ginasputri, H. N. A., Saputri, A. D., Wahid, S. N., Ningrum, A. F., & Haris, M. A. (2025). Analisis sentimen ulasan pengguna BRIMO terhadap pembaruan fitur aplikasi menggunakan Naive Bayes dengan seleksi fitur Chi-Square. Jurnal Statistika Dan Komputasi, 4(2), 56–67. https://doi.org/10.32665/statkom.v4i2.5194

Imani, M., Maroosi, A., Shojaei, S., Heidari, K., Hoseinzadeh, S. M., Daneshi, N., Saber, Z., Sajadi, N., & Mohammadzadeh, M. (2026). Predicting cardiovascular diseases using imbalanced data: An XGBoost-based analysis of the 2022 BRFSS dataset. American Heart Journal Plus Cardiology Research and Practice, 63, 100719. https://doi.org/10.1016/j.ahjo.2026.100719

Liu, W., Fan, H., & Xia, M. (2021). Credit scoring based on tree-enhanced gradient boosting decision trees. Expert Systems With Applications, 189, 116034. https://doi.org/10.1016/j.eswa.2021.116034

Lundberg, S., & Lee, S.-I. (2017). A Unified Approach to Interpreting Model Predictions. Advances in Neural Information Processing Systems. https://neurips.cc/Conferences/2017/Schedule?showEvent=9253

Mohammed, A. M. A., Husain, O., Abdulkareem, M., Yunus, N. Z. M., Jamaludin, N., Mutaz, E., Elshafie, H., & Hamdan, M. (2024). Explainable Artificial Intelligence for predicting the compressive strength of soil and ground granulated blast furnace slag mixtures. Results in Engineering, 25, 103637. https://doi.org/10.1016/j.rineng.2024.103637

Molnar, Christoph. (2025). Interpretable machine learning : a guide for making black box models explainable (3rd ed.). Christoph Molnar. https://christophm.github.io/interpretable-ml-book

Noviandy, T. R., Maulana, A., Irvanizam, I., Idroes, G. M., Maulydia, N. B., Tallei, T. E., Subianto, M., & Idroes, R. (2025). Interpretable machine learning approach to predict Hepatitis C virus NS5B inhibitor activity using voting-based LightGBM and SHAP. Intelligent Systems With Applications, 25, 200481. https://doi.org/10.1016/j.iswa.2025.200481

Óskarsdóttir, M., & Bravo, C. (2021). Multilayer network analysis for improved credit risk prediction. Omega, 105, 102520. https://doi.org/10.1016/j.omega.2021.102520

Rahardiantoro, S., Juhanda, A. R. N., Kurnia, A., Aswi, A., Sartono, B., Handayani, D., Soleh, A. M., Yanti, Y., & Cramb, S. (2024). Spatio-temporal modeling to identify factors associated with stunting in Indonesia using a Modified Generalized Lasso. Spatial and Spatio-temporal Epidemiology, 51, 100694. https://doi.org/10.1016/j.sste.2024.100694

Salih, A. M., Raisi‐Estabragh, Z., Galazzo, I. B., Radeva, P., Petersen, S. E., Lekadir, K., & Menegaz, G. (2024). A perspective on explainable artificial intelligence methods: SHAP and LIME. Advanced Intelligent Systems, 7(1). https://doi.org/10.1002/aisy.202400304

Sartono, B., Elenaputri, T. S., Angraini, Y., & Dito, G. A. (2025). Long Short‐Term Memory‐Based prediction of Indonesian composite stock index returns for early identification of market crises. Applied Computational Intelligence and Soft Computing, 2025(1). https://doi.org/10.1155/acis/6174081

Sejling, C., Jensen, A. K., Zhang, J., Loft, S., Andersen, Z. J., Brandt, J., Stayner, L. T., Pedersen, M., & Budtz‐Jørgensen, E. (2025). Novel approach for hierarchical family selection of an ambient air pollutant mixture with application to childhood asthma. Environmetrics, 36(5). https://doi.org/10.1002/env.70020

Shapley, L. S. (1953). 17. A value for N-Person games. In Princeton University Press eBooks (pp. 307–318). https://doi.org/10.1515/9781400881970-018

Zhang, J., & Zhao, Z. (2025). Corporate ESG rating prediction based on XGBoost-SHAP interpretable machine learning model. Expert Systems With Applications, 295, 128809. https://doi.org/10.1016/j.eswa.2025.128809

Zhao, C., Yan, Z., Sun, X., & Wu, M. (2024). Enhancing aspect category detection in imbalanced online reviews: An integrated approach using Select-SMOTE and LightGBM. International Journal of Intelligent Networks, 5, 364–372. https://doi.org/10.1016/j.ijin.2024.10.002

Zhu, J., Pu, S., He, J., Su, D., Cai, W., Xu, X., & Liu, H. (2024). Processing imbalanced medical data at the data level with assisted-reproduction data as an example. BioData Mining, 17(1), 29. https://doi.org/10.1186/s13040-024-00384-y

Published

2026-06-30
Abstract View: 32, PDF Download: 10 SIMILARITY INDEX Download: 0