SLO-Driven Self-Healing Systems for Healthcare IT: Control Loops, Error Budgets, and Observability-Guided Reliability
Main Article Content
Abstract
Healthcare IT systems operate under strict availability requirements and support complex, interdependent workflows such as prescription fulfillment and claims adjudication. While modern observability platforms improve detection and diagnosis of system issues, they are primarily used for monitoring rather than control, leaving reliability objectives weakly enforced at the workflow level.
This paper presents a conceptual engineering framework that treats service level objectives (SLOs) and error budgets as runtime control inputs within a feedback-driven, self-healing architecture. Observability signals are continuously evaluated against these inputs to detect deviations, and bounded automated actions are applied to steer the system toward acceptable operating conditions while limiting unintended impact. The framework introduces a multi-layer SLO model spanning infrastructure, service, and workflow levels, and defines error budget burn rate as a primary control signal for proportional response. It also outlines control loop design patterns and stability mechanisms, including damping, cooldown intervals, and scoped actions. Rather than guaranteeing absolute reliability, the approach enables bounded and measurable reliability enforcement aligned with end-to-end workflow outcomes, providing a structured path from observability to actionable reliability engineering in healthcare IT systems.