You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
This overstates frontier API cost because it ignores cached token discounts. Cloud APIs (OpenAI, Anthropic, Google) charge ~50% less for recently-seen input tokens. Without modeling this, the self-hosted savings shown on the routing page are inflated.
Cached input tokens (~50% discount, based on cache hit rate)
Output tokens
Breakeven utilization curve (self-hosted GPUs aren't always at 100%)
Add pricePerMCached to FrontierModel interface — currently only has pricePerMInput and pricePerMOutput. Needs a third field for cached input token pricing.
Add cache hit rate input per tier — the routing page lets you configure tokens in/out per tier, but has no cache hit rate slider. Different tiers will have different cache rates (e.g., RAG/doc processing = high cache, conversational = low).
Add utilization modeling — current self-hosted calc assumes GPUs run 24/7 at full capacity. In reality, utilization varies. The standalone calculator models the breakeven utilization % where self-hosting becomes cheaper.
Problem
The routing economics page (
lib/routing/calc.ts) uses a naive cost calculation for frontier API pricing:This overstates frontier API cost because it ignores cached token discounts. Cloud APIs (OpenAI, Anthropic, Google) charge ~50% less for recently-seen input tokens. Without modeling this, the self-hosted savings shown on the routing page are inflated.
What needs to change
Use the shared cost engine being built in Cost comparison: self-hosted on-prem vs cloud API cost per million tokens #10. That engine implements the blended cost model:
Add
pricePerMCachedtoFrontierModelinterface — currently only haspricePerMInputandpricePerMOutput. Needs a third field for cached input token pricing.Add cache hit rate input per tier — the routing page lets you configure tokens in/out per tier, but has no cache hit rate slider. Different tiers will have different cache rates (e.g., RAG/doc processing = high cache, conversational = low).
Add utilization modeling — current self-hosted calc assumes GPUs run 24/7 at full capacity. In reality, utilization varies. The standalone calculator models the breakeven utilization % where self-hosting becomes cheaper.
Context
Current routing calc location
lib/routing/calc.ts—computeFrontierCost()andcomputeSelfHostedCost()lib/pricing/frontier-models.ts— model pricing catalog (needspricePerMCachedfield)Blocked by