Budget-Aware Calibrated Cascades for E-Commerce Customer Assistants: A Cost–Quality Routing Benchmark

Main Article Content

Veera Ravindra Divi

Abstract

Large language models (LLMs) are increasingly being deployed at the forefront of e-commerce customer support, where retailers may handle millions of daily interactions covering order tracking, return policies, product inquiries, and complex customer complaints. While using advanced models for every interaction ensures high accuracy, it becomes economically impractical at such scale; conversely, relying solely on smaller models reduces costs but often fails to address high-value or complex queries, leading to increased returns, customer churn, and escalation expenses. Traditional cost-saving approaches, such as model cascades that rely on a model’s self-reported confidence to decide escalation, were originally designed for open-domain question answering and exhibit key limitations in e-commerce contexts. Specifically, model confidence is often poorly calibrated and subject to drift, and a single accuracy target does not account for the varying business impact of incorrect responses across different query types. To address these challenges, this work introduces Budget-Aware Calibrated Cascading (BACC), a routing strategy that bases escalation decisions on calibrated correctness estimates weighted by the business cost of potential errors, while adhering to per-conversation cost constraints and latency service-level objectives (SLOs). Additionally, a reproducible simulation framework is presented, modeling realistic e-commerce query distributions, asymmetric error costs, tiered model pricing based on publicly available 2023 data, and multiple traffic-shift scenarios. Results from these simulations demonstrate that BACC achieves a more efficient cost–quality trade-off compared to traditional raw-confidence cascades like FrugalGPT, and maintains stable performance and calibration even under shifting traffic conditions. All findings are derived from transparent, seeded simulations and are intended to illustrate the methodology rather than represent outcomes from a live deployment.

Article Details

Section
Articles