An Explainable AI-Driven Reliability Engineering Framework for Predictive Failure Management and Self-Healing Operations in Hyperscale GPU Clusters
Main Article Content
Abstract
The research utilizes explainable AI-based framework of reliability engineering for predictive failure management and self-healing operation to be developed and evaluated in hyper scale GPU clusters. The telemetry data produced were synthetic to model hardware, thermal, fabric, workload and maintenance condition. Random Forest, XGBoost and Random Survival Forest were used with XGBoost performing at an accuracy of 88.12% and an ROC-AUC of 92.13%. SHAP analysis and interpretability, a predictive compute readiness helped with risk-aware intervention. This improved consistency, resilience, efficiency, transparency and operational control is seen in a decrease in interruptions from 258 on the long tail of service metrics to 111, a decrease in MTTR from 5.5 hours to 1.5 hours and an increase in goodput to 99.42% of the long tail.