International Journal of Scientific Engineering and Research (IJSER)
Call for Papers | Fully Refereed | Open Access | Double Blind Peer Reviewed | ISSN: 2347-3878


Downloads: 1

India | Computer Science and Engineering | Volume 14 Issue 8, August 2026 | Pages: 55 - 65


An Interpretable Machine Learning Framework with Cross-Model Feature Stability Analysis for Heart Disease Prediction

Rekha P, Dr. Malliga P

Abstract: Machine learning models for heart disease prediction are typically compared on discrimination alone, and the feature importances of a single selected model are then reported as though they described the data rather than the model. This paper develops an interpretable framework in which cross-model agreement is treated as a measurable quantity rather than an assumption. Logistic Regression, Random Forest, and Extreme Gradient Boosting (XGBoost) are trained on the Cleveland heart disease dataset, and SHapley Additive exPlanations (SHAP) are computed for every model at both the cohort and the individual-patient level. A Feature Stability Score (FSS) is used to combine the three attribution profiles into a single ranking. We show analytically and empirically that the variance-based form of this score, FSS = 1 - σ2, is not scale-invariant and is therefore confounded with importance itself: across the thirteen predictors its rank correlation with mean importance is ρ = -0.973, so the least important feature is scored as the most stable, and the resulting composite ranking correlates with raw importance at ρ = +0.995 - the stability term changes essentially nothing. We propose a scale-invariant reformulation based on the coefficient of variation of per-model normalised attributions, under which stability becomes an independent signal (ρ = +0.703) and the ranking changes materially: oldpeak rises from seventh to fourth while sex, chol, and age fall, the latter because Logistic Regression and XGBoost disagree about it by a factor of thirty. Random Forest attains the best discrimination (accuracy 0.8361, ROC-AUC 0.9188). At the individual level the three models disagree sharply on a representative patient, returning probabilities of 0.4987, 0.6224, and 0.1995 and splitting on the resulting decision, while the patient-level feature ranking diverges from the cohort ranking (ρ = 0.698, with eleven of thirteen positions changed). These results indicate that a single-model explanation of a heart disease classifier is not sufficient evidence about the underlying predictors.

Keywords: Heart disease prediction, explainable artificial intelligence, SHAP, feature stability, cross-model agreement, model disagreement, clinical decision support


View Article PDF


Rate This Article


Top