A system handles 100K requests/day. Baseline frontier cost is $3,000/day. With 70% routed to small models ($77/day), 30% on frontier ($900/day), and classifier overhead ($5/day), what is the daily cost and percentage savings?
- About $1,500/day with 50% savings, because the routing split averages the frontier and small model costs equally across the full volume of production traffic
- About $500/day with 83% savings, because semantic caching handles an additional 50% of the small-model traffic at zero incremental LLM cost on top of the routing savings
- About $982/day with 67% savings, computed as $77 small-model cost plus $900 frontier cost plus $5 classifier overhead, confirming the 40-70% savings range reported in production deployments
- About $2,100/day with 30% savings, because the 30% of complex queries on frontier models dominate the cost and limit the achievable savings regardless of how much traffic is routed cheaply, based on empirical data from compound systems in production
Why
The calculation follows the routing cost formula directly. Small model daily cost for 70,000 requests at approximately $1/M input and $2/M output with 500 input and 300 output tokens per request equals roughly $77. Frontier model daily cost for 30,000 requests at $15/M input and $75/M output equals roughly $900. Classifier overhead for 100,000 requests at approximately 200 tokens each at $0.25/M equals roughly $5. Total daily cost is $77 plus $900 plus $5 equals $982. Savings from the $3,000 baseline are $2,018, or approximately 67%. This aligns with the 40-70% savings range reported in production deployments. The $1,500 estimate incorrectly averages costs. The $2,100 estimate overestimates frontier dominance. The $500 estimate incorrectly adds caching benefits not specified in the scenario. The classifier costs less than 1% of the total budget but controls where the other 99% goes. The break-even point for adding routing infrastructure is approximately $5,000 monthly inference spend.