Outcome
After sizing a model, users should seamlessly answer three cost questions without leaving ConfigIQ:
- Should I self-host or pay per-token? — Compare self-hosted cost/M tokens vs frontier API pricing, with cached token blending and breakeven utilization
- What does my self-hosted infra cost? — Full TCO breakdown (GPU, power, networking, ops, storage, software) for cloud IaaS vs on-prem
- How much more do I save with semantic routing? — Route queries to right-sized models based on complexity (e.g., simple → Llama 8B, complex → Llama 70B) and show the blended cost savings vs sending everything to one large model or a frontier API
Note: "Routing" here means model selection routing — classifying queries by complexity and sending each tier to an appropriately sized OSS model. This is NOT llm-d intelligent scheduling (which routes requests to replicas with warm KV cache for the same model). Model selection routing is a cost optimization; llm-d scheduling is an infrastructure optimization.
Flow
Sizing (AIC)
→ "You need 8x H100 for Llama 70B"
↓
View 1: Should I self-host?
→ "At 45%+ utilization you save $X vs OpenAI"
↓
View 2: What's the infra cost?
→ "$82K/mo cloud IaaS, $54K/mo on-prem"
↓
View 3: Can routing save more?
→ "Route 60% simple queries to 8B model, save additional 40%"
Use case presets for workload shape
Workload shape (token ratio, cache hit rate, concurrency) varies dramatically by use case. Instead of users guessing these values, the cost views should offer use case presets that auto-populate the workload parameters:
| Use case |
Input/Output ratio |
Cache hit rate |
Effect on cost |
| RAG / doc processing |
High input, low output |
High (repeated docs) |
Cloud API cheaper (cached input discounts) |
| Chatbot / conversational |
Moderate both |
Low (unique conversations) |
Self-hosted wins faster |
| Agentic workflows |
Low input, high output |
Low |
Output-heavy = expensive on APIs, self-hosted wins |
| Code generation |
High input (context), high output |
Moderate |
Depends on volume |
| Summarization |
Very high input, low output |
High |
Cloud caching helps a lot |
This matters because:
- Current model assumes a generic 30/70 input/output ratio for all use cases (flagged by team)
- Agentic workflows flip the ratio entirely — output-heavy
- Cache hit rate changes the blended API cost significantly
- Cloud APIs expire cache in 5min–1hr; on-prem (llm-d KV-cache-aware scheduling) can tune cache per use case — a self-hosting advantage not captured today
The routing page's tier structure (simple/standard/complex) is the right skeleton — but tiers should be driven by use case presets rather than generic complexity buckets.
Current state
| View |
Status |
Gap |
| View 1: Self-host vs API |
Standalone only (separate calculator) |
Not integrated into ConfigIQ |
| View 2: Infra TCO |
Built (cluster cost page) |
Hardcoded pricing, no sizing integration |
| View 3: Routing savings |
Built (routing economics page) |
Naive cost calc, stale model pricing |
Shared cost engine
All three views should use a shared cost engine that implements:
- Cached token blending (input, cached input, output)
- Breakeven utilization modeling
- Per-use-case token ratios via presets (not one-size-fits-all)
- Cache hit rate modeling with on-prem advantage (longer cache TTL, KV-cache-aware scheduling)
- Auto-population from sizing output (GPU type + count)
Reference: standalone self-hosted LLM cost calculator — https://openshift-psap.github.io/self-hosted-llm-cost-calculator/
Child issues
Outcome
After sizing a model, users should seamlessly answer three cost questions without leaving ConfigIQ:
Flow
Use case presets for workload shape
Workload shape (token ratio, cache hit rate, concurrency) varies dramatically by use case. Instead of users guessing these values, the cost views should offer use case presets that auto-populate the workload parameters:
This matters because:
The routing page's tier structure (simple/standard/complex) is the right skeleton — but tiers should be driven by use case presets rather than generic complexity buckets.
Current state
Shared cost engine
All three views should use a shared cost engine that implements:
Reference: standalone self-hosted LLM cost calculator — https://openshift-psap.github.io/self-hosted-llm-cost-calculator/
Child issues