Skip to content

Cost experience: unified cost views for sizing-to-economics flow #13

Description

@aditisaluja5

Outcome

After sizing a model, users should seamlessly answer three cost questions without leaving ConfigIQ:

  1. Should I self-host or pay per-token? — Compare self-hosted cost/M tokens vs frontier API pricing, with cached token blending and breakeven utilization
  2. What does my self-hosted infra cost? — Full TCO breakdown (GPU, power, networking, ops, storage, software) for cloud IaaS vs on-prem
  3. How much more do I save with semantic routing? — Route queries to right-sized models based on complexity (e.g., simple → Llama 8B, complex → Llama 70B) and show the blended cost savings vs sending everything to one large model or a frontier API

Note: "Routing" here means model selection routing — classifying queries by complexity and sending each tier to an appropriately sized OSS model. This is NOT llm-d intelligent scheduling (which routes requests to replicas with warm KV cache for the same model). Model selection routing is a cost optimization; llm-d scheduling is an infrastructure optimization.

Flow

Sizing (AIC)
  → "You need 8x H100 for Llama 70B"
     ↓
View 1: Should I self-host?
  → "At 45%+ utilization you save $X vs OpenAI"
     ↓
View 2: What's the infra cost?
  → "$82K/mo cloud IaaS, $54K/mo on-prem"
     ↓
View 3: Can routing save more?
  → "Route 60% simple queries to 8B model, save additional 40%"

Use case presets for workload shape

Workload shape (token ratio, cache hit rate, concurrency) varies dramatically by use case. Instead of users guessing these values, the cost views should offer use case presets that auto-populate the workload parameters:

Use case Input/Output ratio Cache hit rate Effect on cost
RAG / doc processing High input, low output High (repeated docs) Cloud API cheaper (cached input discounts)
Chatbot / conversational Moderate both Low (unique conversations) Self-hosted wins faster
Agentic workflows Low input, high output Low Output-heavy = expensive on APIs, self-hosted wins
Code generation High input (context), high output Moderate Depends on volume
Summarization Very high input, low output High Cloud caching helps a lot

This matters because:

  • Current model assumes a generic 30/70 input/output ratio for all use cases (flagged by team)
  • Agentic workflows flip the ratio entirely — output-heavy
  • Cache hit rate changes the blended API cost significantly
  • Cloud APIs expire cache in 5min–1hr; on-prem (llm-d KV-cache-aware scheduling) can tune cache per use case — a self-hosting advantage not captured today

The routing page's tier structure (simple/standard/complex) is the right skeleton — but tiers should be driven by use case presets rather than generic complexity buckets.

Current state

View Status Gap
View 1: Self-host vs API Standalone only (separate calculator) Not integrated into ConfigIQ
View 2: Infra TCO Built (cluster cost page) Hardcoded pricing, no sizing integration
View 3: Routing savings Built (routing economics page) Naive cost calc, stale model pricing

Shared cost engine

All three views should use a shared cost engine that implements:

  • Cached token blending (input, cached input, output)
  • Breakeven utilization modeling
  • Per-use-case token ratios via presets (not one-size-fits-all)
  • Cache hit rate modeling with on-prem advantage (longer cache TTL, KV-cache-aware scheduling)
  • Auto-population from sizing output (GPU type + count)

Reference: standalone self-hosted LLM cost calculator — https://openshift-psap.github.io/self-hosted-llm-cost-calculator/

Child issues

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    cost-modelingCost estimation and comparison features

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions