An interpretable hybrid ensemble model for early academic risk detection with shap-based explainability
DOI:
https://doi.org/10.52465/joscex.v7i3.50Keywords:
Academic risk, Educational data mining, Ensemble learning, Interpretable machine learning , PredictionAbstract
Student dropout remains a major challenge in educational institutions, affecting both academic achievement and institutional performance. Early identification of at-risk students is essential to support timely intervention and improve retention rates. This study proposes an interpretable hybrid ensemble model for early academic risk detection by integrating stacking ensemble learning, Bayesian optimization, and SHAP-based interpretability. The study used a public dataset from Kaggle containing 10,000 student records with demographic, behavioral, and academic performance attributes. Exploratory analysis indicated that at-risk students generally had lower GPA and attendance rates, while higher stress levels were associated with increased dropout risk. Due to the moderately imbalanced class distribution in the dataset, SMOTE was applied only to the training data after train–test splitting to improve minority class representation while avoiding data leakage during evaluation. Experimental results demonstrate that the proposed hybrid model outperformed individual baseline classifiers and standard stacking ensemble methods, achieving an accuracy of 0.93 ± 0.01, precision of 0.92 ± 0.01, recall of 0.91 ± 0.01, and F1-score of 0.91 ± 0.01 under stratified 5-fold cross-validation. The integration of Bayesian Optimization improved model stability and classification consistency, while SHAP-based analysis identified GPA, CGPA, attendance rate, study hours, and stress index as the most influential contributors to prediction outcomes.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Soft Computing Exploration

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
