A Recovery-Aware Optimization Framework for Petabyte-Scale Distributed Data Processing: Balancing Performance, Reliability, and Infrastructure Cost

Main Article Content

Chalapathi Koneni, Ashok Kumar

Abstract

A batch processing performance problem is not a single objective since throughput, recovery behavior and cost of infrastructure are highly interrelated at a petabyte scale. Over-parallelism can enhance performance but also raise the shuffling overhead, sensitivity to failures and operational cost, and conservative choices can decrease performance and serve to slow down data delivery. There is a major challenge as petabyte scale distributed processing systems have a rising failure probability, recovery cost and infrastructure cost. In this study, a new recovery-conscious optimization model has been suggested, which considers the decision made in checkpointing and also materialization, depending on workload features, failure requirements, re-computation cost, and system efficiency. A compatible model workload scenario that comprises 500,000 petabytes of workload cases is constructed in order to broaden synthetic data to simulate diverse distributed settings. The four recovery strategies No Checkpoint, Periodic Checkpoint, Adaptive Checkpoint and Selective Materialization have been compared by utilizing a cost-based optimization model and Pareto analysis. Findings indicate that Adaptive Checkpointing offers better recovery performance through proportional reliability, throughput and cost of operation.

Article Details

Section
Articles