Skip to content

Batch sizing: get recommendations for a list of models in one run #67

Description

@aditisaluja5

Today ConfigIQ sizes one model at a time. Sizing a set of models (e.g. for an architecture doc or a customer deployment plan) means running the flow once per model and manually collecting the results into a table.

Request: Let a user submit a list of models and get sizing back for all of them in one run.

Proposed shape:

  • Input a list of models (paste HuggingFace IDs, or multi-select from the model picker), plus one shared workload profile (GPU, ISL, OSL, TTFT/concurrency target, precision).
  • ConfigIQ runs the recommend flow for each model and returns a single results table: model, GPUs/replica (TP), replicas for target, memory/replica, TTFT, TPOT, throughput.
  • Export the table to CSV.
  • Handle unsupported models gracefully in the output (mark "not supported in AIC" rather than failing the whole batch).

Why: Real sizing tasks are almost always multi-model (a customer runs a fleet, an arch doc lists 10+ models). One-at-a-time is slow and error-prone when assembling comparisons.

Notes:

  • Per-model workload overrides could come later; a shared workload profile covers the common case.
  • Pairs well with the existing CSV export on the performance page.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    kind/featureCategorizes issue or PR as related to a new feature.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions